Underwater target detection method fusing lightweight multi-scale and flexible attention
Through the multi-scale feature extraction module and soft channel attention module combined with the bottleneck structure, the problem of feature loss and large model scale in underwater target detection is solved, and lightweight and efficient underwater target detection is achieved.
Patent Information
- Application Number
- CN202510796659.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art has a multi-scale feature fusion method in underwater target detection that damages features, resulting in feature loss, and high requirements for large-scale model calculations, limiting real-time and hardware requirements.
The multi-scale feature extraction module and soft channel attention module are used to combine the bottleneck structure to fusion feature through multi-scale convolution kernel and Softmax function to avoid traditional pooling and upsampling, reduce the model scale and retain feature information.
It realizes lightweight underwater target detection, taking into account local details and global information, improving detection accuracy and speed, and reducing computing needs.
Smart Images

Figure CN120339819A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of underwater side-scan sonar image processing, and in particular to an underwater target detection method that combines lightweight multi-scale and soft communication attention. Background Art
[0002] Side-scan sonar can clearly image the seabed at a relatively long distance. It is the main tool for humans to conduct seabed exploration and plays an important role in tasks such as seabed mapping, resource exploration, shipwreck salvage, and seabed oil and gas pipeline detection. In recent years, with the improvement of computer computing power, object detection methods based on deep learning have been applied in more and more fields. They overcome the dependence on manual feature extraction of traditional object detection methods and have strong adaptability to different backgrounds.
[0003] To enhance the network's adaptability to different targets and help the network better and faster focus on the target areas of concern, multi-scale feature extraction and attention mechanisms are often embedded in existing models. Dalian University of Technology uses an upsampling method in CN112257810A to splice deep features after upsampling with shallow features to achieve multi-scale feature fusion. The Dalian Naval Academy of the Chinese People's Liberation Army discloses an improved network based on the YOLOv8 model in CN118736187A for detecting targets in sonar images, which embeds an EMA attention module. Sichuan University designs a multi-scale feature fusion module using a method of multiple pooling splicing in CN119646484A to enhance the network's adaptability to multi-scale targets. However, there are still some problems in the research of existing technologies. The mainstream multi-scale feature fusion methods either splice shallow features with deep features after multiple poolings, or splice deep features with shallow features after upsampling. Although this method enables the final detection layer to take into account both shallow and deep features, the traditional pooling and upsampling methods themselves will damage the extracted features, and this impact is more obvious when processing low-resolution images (such as sonar images). The application of basic pooling layers is also common in various attention mechanisms, which easily leads to the loss of some features. In addition, in industrial applications, the algorithm usually requires good real-time performance, and the performance of deep learning algorithms is often positively correlated with the scale of trainable parameters they contain. Some top-performing models usually have a large scale, which largely makes up for the impact of local feature loss. However, this type of model has high requirements for computer hardware and slow running speed, which to a certain extent limits the application of deep learning algorithms in the field of seabed exploration. Summary of the Invention
[0004] To address the problems of the above-mentioned existing technologies, the present invention proposes an underwater target detection method that integrates lightweight multi-scale and soft channel attention. A self-developed multi-scale feature extraction module is used to provide a rich receptive field for the neural network model; a soft channel attention module is used to enable the neural network model to easily focus on the important channels of the feature map; a fusion bottleneck structure is used to compress the feature channels in the middle and later parts of the neural network model to reduce the scale of the neural network model; at the same time, direct use of upsampling, downsampling, or pooling methods is avoided, and instead, multi-scale convolutional kernels and the Softmax function are used to reduce the dimension and fuse features, avoiding damage to the extracted features.
[0005] To achieve the above objectives, the present invention proposes an underwater target detection method that integrates lightweight multi-scale and soft channel attention, including the following steps: S1. Establish a neural network model Build a basic network model including an integrated convolutional layer module, a C2f layer module, and a target detection head module; Based on the basic network model, embed a multi-scale feature extraction module and a channel soft attention module to construct an improved neural network model. The multi-scale feature extraction module provides a rich receptive field for the improved neural network model, and the channel soft attention module is used to reduce the difficulty for the improved neural network model to focus on the important channels of the feature map, making it easy to identify each pixel point on the feature map; S2. Prepare the dataset Adopt the SCTD seafloor side-scan sonar dataset and divide it into a training set and a test set; S3. Train the improved neural network model Use the training set to train the improved neural network model to obtain an underwater target detection model; S4. Test the underwater target detection model Use the test set to test the underwater target detection model and evaluate the performance of the model from two aspects: visual and objective evaluation metrics.
[0006] Furthermore, the SCTD seafloor side-scan sonar dataset includes three types of target data: ships, airplanes, and dummies. The format of the SCTD seafloor side-scan sonar dataset is converted from xml format tags to txt format.
[0007] Preferably, the present invention uses the Softmax function to improve the channel soft attention module to form a spatial fusion soft channel attention module (SSCA). The spatial fusion soft channel attention module uses the Softmax function to assign a weight to each pixel point on each feature map. The weights are added to obtain a pooling result. The pooled value contains the information of all pixel points of the original feature map of this channel and adds a non-linear component to the model. The pooling result passes through a multi-layer perceptron and then passes through SiLU activation to obtain the final weight of each channel. The original feature map is multiplied by the final weight to obtain a new feature map. The mathematical expression of the Softmax function is: ; where , is the pixel value of a certain point on a single channel, and N is the number of pixel points of the single-channel feature image.
[0008] Then the mathematical description of the spatial information fusion pooling of the m-th channel of the feature map is: .
[0009] Preferably, the improved neural network model is composed of a conventional integrated convolutional layer module, a C2f layer module, a multi-scale extraction and soft channel attention layer module, a multi-scale extraction and soft channel attention layer module based on the bottleneck structure, and an object detection detection head module.
[0010] Preferably, the structures in the object detection model YOLO are adopted for the conventional integrated convolutional layer module, the C2f layer module, and the object detection detection head module. The conventional integrated convolutional layer module integrates two-dimensional convolution, batch normalization, and the SiLU activation function. The C2f layer module uses a convolutional layer to adjust the channels of the feature map, divides the channels of the adjusted feature map into two parts, one part is directly output, and the other part is output after passing through multiple bottleneck structures, and then the feature map is concatenated in the channel dimension through a concatenation layer, and finally the number of channels is restored to the number of channels before the Split layer. The object detection detection head has two branches, each branch is composed of several consecutive convolutional layers, and the output after the feature map is processed by the object detection detection head is used to calculate the loss function.
[0011] Preferably, the calculation of the loss function includes calculating the classification loss, the complete intersection over union loss, and the distribution focal loss.
[0012] Preferably, the multi-scale extraction and soft channel attention layer module includes three downsampling convolutional branches. The kernel sizes (k) of the downsampling convolutional branches are 3, 5, and 7 respectively, the stride (s) is 2 for all, and the padding (p) sizes are 1, 2, and 3 respectively. The downsampling convolutional branches use convolutional kernels of different scales to extract features from the feature map. The three downsampling convolutional branches have different receptive fields, which are 9, 25, and 49 respectively. After subsequent splicing and fusion, the local details and global information of the image can be taken into account.
[0013] Preferably, the multi-scale extraction and soft channel attention layer module based on the bottleneck structure first uses ConvD to reduce the dimension of the channels of the input feature map, then performs multi-scale extraction, and then uses ConvU to increase the dimension of the channels of the feature map to limit the scale of the neural network model when there are too many channels in the feature map.
[0014] Furthermore, the multi-scale extraction and soft channel attention layer module based on the bottleneck structure reserves c-down and c-up parameters to facilitate the control of the reduction and increase multiples of the channels. The number of channels of the feature map before and after being processed by ConvD are respectively: output channels = input channels × c-down, where 0 < c-down ≤ 1; the number of channels of the feature map before and after being processed by ConvU are respectively: output channels = input channels × c-up, where c-up ≥ 1; where, c-down is the channel reduction multiple, and c-up is the channel increase multiple.
[0015] Compared with the prior art, the following beneficial effects are achieved: 1. The multi-scale feature extraction module proposed by the present invention includes three specially designed downsampling convolutional branches. The three downsampling convolutional branches use convolutional kernels of different scales to extract features from the feature map. The three branches have different receptive fields. After subsequent splicing and fusion, the local details and global information of the image can be taken into account; all feature maps are spliced in the channel dimension through the Concat layer, and then the multi-scale feature extraction module assigns weights to each channel according to the importance degree, so as to facilitate the subsequent compression of the channels by the C2f layer, discard redundant information while limiting the scale of the neural network model.
[0016] 2. Different from the classical channel attention mechanism that performs max-pooling and average-pooling on each channel of the feature map (max-pooling only retains the pixel point with the largest pixel value on each channel, while average-pooling retains the mean value of all pixel points on each channel, both of which will cause a certain degree of feature loss), the present invention improves the channel soft attention module using the Softmax function to form a spatially fused soft channel attention module (SSCA). The spatially fused soft channel attention module uses the Softmax function to assign a weight to each pixel point on each feature map, and the weights are added to obtain the pooling result. The pooled value contains the information of all pixel points of the original feature map of this channel and can add a non-linear component to the model; the pooling result passes through a multi-layer perceptron and then passes through SiLU activation to obtain the final weight of each channel, and the final weight is multiplied by the original feature map to obtain a new feature map, avoiding damage to the extracted features.
[0017] 3. The multi-scale extraction and soft channel attention module layer based on the bottleneck structure proposed by the present invention first uses ConvD to reduce the dimension of the input channels, then performs multi-scale extraction, and then uses ConvU to increase the dimension of the feature map channels to limit the scale of the neural network model in the case of too many feature channels. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a schematic flow chart of the present invention; Figure 2 is a schematic diagram of the overall structure of the improved neural network model of the present invention; Figure 3 is a schematic diagram of the internal structures of Conv, C2f, Detect-Head, and Bottelneck adopted from the YOLOv8 model; Figure 4 is a schematic diagram of the internal structure of the spatially fused soft channel attention module; Figure 5 is a schematic diagram of the internal structure of the multi-scale soft channel attention feature extraction module; Figure 6 is a schematic diagram of the internal structure of the multi-scale extraction and soft channel attention layer based on the bottleneck structure; Figure 7 is the convergence curve of the loss function during the training process of the present invention; Figure 8 (a) is a PR curve graph; FIG. 8(b) is an IoU intersection ratio graph; Figure 9 is the effect diagram of the object detection result. DETAILED DESCRIPTION OF THE INVENTION
[0019] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention. Specific Embodiment 1 As shown in the Figure 1 accompanying drawings, an underwater target detection method integrating lightweight multi-scale and soft channel attention includes the following steps: S1. Establish a neural network model Build a basic network model including an integrated convolutional layer module, a C2f layer module, and a target detection head module; Based on the basic network model, embed a multi-scale feature extraction module and a channel soft attention module to construct an improved neural network model. The multi-scale feature extraction module provides a rich receptive field for the basic network model, and the channel soft attention module is used to reduce the difficulty of the basic network model focusing on the important channels of the feature map, making it easy to identify each pixel point on the feature map; S2. Dataset preparation Use the publicly available SCTD seafloor side-scan sonar dataset, which contains three types of targets, namely ships, airplanes, and dummies, and divide the training set and the test set. For the convenience of training, convert the original xml format labels into txt format; S3. Train the improved neural network model; Use the training set generated in S2 to train the improved neural network model to obtain an underwater target detection model; S4. Test the underwater target detection model; Use the test set generated in S2 to test the underwater target detection model obtained in S3, and evaluate the performance of the model from two aspects: visual and objective evaluation indicators; In a preferred embodiment, the improved neural network model consists of 1 conventional integrated convolutional layer (Conv), 4 C2f layers, 1 multi-scale extraction and soft channel attention layer (Multi scale extraction and soft channel attention layer, MSSCA), 3 multi-scale extraction and soft channel attention layers based on the bottleneck structure (Multi scale extraction and soft channel attention layer based on bottleneck structure, B-MSSCA), and a target detection detection head (Detect-Head). Among them, the Conv layer, C2f layer, and Detect-Head follow the structures in the mainstream target detection model YOLO, and their internal structures are shown in (a), (b), and (c) of Figure 3 respectively; the Conv layer integrates two-dimensional convolution (Conv2d), batch normalization (BatchNorm2d), and SiLU activation function; the C2f layer first adjusts the channels of the feature map using a convolutional layer with a kernel size of 1×1, then the Split layer divides the channels of the feature map into two parts, one part is directly output, and the other part is output after passing through multiple bottleneck structures (Bottleneck), then the feature map is concatenated in the channel dimension through the Concat layer, and finally the number of channels is restored to the number of channels before the Split layer using a 1×1 convolution; the Detect-Head has two branches, each branch consists of several consecutive convolutional layers, and the output after the feature map is processed by the Detect-Head is used to calculate the loss functions, namely classification (Classification, CLS) loss, complete intersection over union loss (Complete Intersection over Union Loss, CIoU) loss, and distribution focal loss (Distribution Focal Loss, DFL).
[0021] In a preferred embodiment, in step 1 of the present invention, the existing channel attention module is improved, and a soft channel attention module (Soft Channel Attention with Softmax, SSCA) that uses the Softmax function for spatial fusion is proposed. The internal structure of the soft channel attention module is shown in Figure 4. Different from the classical channel attention mechanism, which performs max pooling and average pooling on each channel of the feature map (max pooling only retains the pixel point with the largest pixel value on each channel, while average pooling retains the mean value of all pixel points on each channel, both of which will cause a certain degree of feature loss), the soft channel attention module uses the Softmax function to assign a weight to each pixel point on each feature map, and then adds them up to obtain the pooling result. The pooled value contains the information of all pixel points of the original feature map of this channel (i.e., spatial information fusion pooling), and can add a non-linear component to the model. Then, it passes through a multilayer perceptron (MLP), and after SiLU activation, the final weight of each channel is obtained and multiplied by the original feature map to obtain a new feature map; among them, the mathematical expression of the Softmax function is: ; , is the pixel value of a certain point on a single channel, and N is the number of pixel points of the single-channel feature image is..; then the mathematical description of the spatial information fusion pooling of the m-th channel of the feature map is: .
[0022] As shown in the attached Figure 5 figure, in a preferred embodiment, the multi-scale soft channel attention feature extraction module (MSSCA) designed by the present invention includes three specially designed downsampling convolutional branches. The convolutional kernel sizes (k) are 3, 5, and 7 respectively, the stride (s) is 2 for all, and the edge padding (p) sizes are 1, 2, and 3 respectively. Feature extraction is performed on the feature map using convolutional kernels of different scales. The downsampling convolutional branches have different receptive fields, which are 9, 25, and 49 respectively. After subsequent splicing and fusion, the local details and global information of the image can be taken into account. The multi-scale soft channel attention feature extraction module splices all feature maps in the channel dimension through a Concat layer, and then passes through the soft channel attention module to assign weights to each channel according to the importance, so as to facilitate subsequent compression of the channels using the C2f layer, discard redundant information while restricting the network scale.
[0023] As shown in the attached Figure 6As shown, in a preferred embodiment, the multi-scale extraction and soft channel attention module based on bottleneck structure (B-MSSCA) designed by the present invention combines the bottleneck structure on the basis of the multi-scale soft channel attention feature extraction module. First, ConvD is used to reduce the dimension of the input channels, then multi-scale extraction is performed, and then ConvU is used to increase the dimension of the feature map channels, which is used to limit the network scale when there are too many feature channels. The parameters c-down and c-up are reserved in the multi-scale extraction and soft channel attention module based on the bottleneck structure to facilitate the control of the reduction and increase multiples of the channels. The number of channels of the feature map before and after ConvD is: output channel number = input channel number × c-down, where 0 < c-down ≤ 1. The number of channels of the feature map before and after ConvU is: output channel number = input channel number × c-up, where c-up ≥ 1; c-down is the channel reduction multiple, and c-up is the channel increase multiple. Specific Embodiment 2 Combined with the attached Figures 1-9 As shown, this embodiment uses the publicly available SCTD dataset to test the underwater target detection model proposed by the present invention and compares it with the relatively advanced YOLOv11 series models. The comparison results are shown in Table 1.
[0025] An underwater target detection method integrating lightweight multi-scale and soft channel attention includes the following steps: Step 1: First, build an improved neural network model, and the structure of the improved neural network model is shown in Figure 2; Step 2: Convert the SCTD dataset, convert the original xml format tags to txt format, and divide the training set and the test set, where there are 329 images in the training set and 28 images in the test set; Step 3: Train the improved neural network model to obtain an underwater target detection model. The input image size in the model is set to 3×640×640, the learning rate is set to 0.01, the number of training rounds is set to 200 rounds, and the BatchSize is set to 4; the convergence of the loss function of the improved neural network model on the training set during the training process is shown in Figure 7 of the attached drawings; Step 4: Test the underwater target detection model, and some detection results are shown in Figure 9 of the attached drawings; Step 5: Perform quantitative analysis on the detection results.
[0026] The present invention objectively and quantitatively evaluates the detection results through the mean Average Precision (mAP); this indicator takes into account the recall rate and detection precision of all categories and is used as the only indicator in multiple object detection competitions. Detection precision (Precision) represents the probability that the actually positive samples among all the samples predicted as positive, and the recall rate (Recall) represents the probability that the actually positive samples are predicted as positive. Their respective mathematical expressions are: ; ; Among them, TP represents the number of samples that are actually 1, predicted as 1, and predicted correctly; FP represents the number of samples that are actually 0, predicted as 1, and predicted incorrectly; FN represents the number of samples that are actually 1, predicted as 0, and predicted incorrectly. The two indicators of detection precision and recall rate evaluate the detection performance from complementary perspectives.
[0027] As shown in Figure 8 (a), in the binary classification task, the PR curve (Precision-Recall Curve) takes the recall rate as the horizontal axis and the precision as the vertical axis. By plotting the relationship between the precision and the recall rate, the area enclosed by the PR curve and the x and y axes is the average precision (AP), which shows the performance of the model at different decision thresholds. This decision threshold is usually the Intersection over Union (IoU). As shown in Figure 8(b) of the attached drawings, IoU represents the intersection over union of the detection box and the ground truth box, where A represents the ground truth box of the target, B represents the predicted box, and C represents the overlapping part of the ground truth box and the predicted box. It is expressed by the formula: ; When facing a multi-object classification task, calculating the average AP of all objects can obtain the mAP. mAP50 and mAP50~95 respectively represent the mAP when IoU takes 50%, and the average value of mAP when IoU takes 50%, 55%, 60%,......, 95% respectively.
[0028] Quantitatively evaluate the network scale and performance using the number of network parameters, floating-point operations (GFLOPs), mAP50, and mAP50-95, and compare with the current relatively advanced YOLOv11 model. The results are shown in Table 1: The detection model established in the present invention has a similar scale to YOLOv11n and belongs to a lightweight model; in terms of the mAP50 index, it is improved by 10.18%, 7.82%, 7.55%, 6.34%, and 5.56% compared with YOLOv11n, YOLOv11s, YOLOv11m, YOLOv11L, and YOLOv11x respectively; in terms of the mAP50-mAP95 index, it is improved by 16.33%, 8.54%, 8.67%, 5.91%, and 6.45% compared with YOLOv11n, YOLOv11s, YOLOv11m, YOLOv11L, and YOLOv11x respectively, and the effect is significant.
[0029] Table 1
Claims
1. An underwater target detection method that integrates lightweight multi-scale and soft communication attention, characterized in that, It includes the following steps: S1. Establish a neural network model Build a basic network model including an integrated convolutional layer module, a C2f layer module, and a target detection head module; Based on the basic network model, embed a multi-scale feature extraction module and a channel soft attention module to construct an improved neural network model. The multi-scale feature extraction module provides a rich receptive field for the improved neural network model, and the channel soft attention module is used to reduce the difficulty of the improved neural network model focusing on the important channels of the feature map, making it easy to identify each pixel point on the feature map; S2. Dataset preparation Adopt the SCTD sub-bottom side-scan sonar dataset and divide it into a training set and a test set; S3. Train the improved neural network model Use the training set to train the improved neural network model to obtain an underwater target detection model; S4. Test the underwater target detection model Use the test set to test the underwater target detection model and evaluate the performance of the model from two aspects: visual and objective evaluation metrics.
2. The underwater target detection method integrating lightweight multi-scale and soft communication attention according to claim 1, wherein The SCTD sub-bottom side-scan sonar dataset includes three types of target data: ships, airplanes, and dummies. Convert the format of the SCTD sub-bottom side-scan sonar dataset from xml format tags to txt format.
3. A method for underwater target detection that integrates lightweight multi-scale and soft communication attention, characterized in that Improve the channel soft attention module using the Softmax function to form a spatial fusion soft channel attention module (SSCA). The spatial fusion soft channel attention module uses the Softmax function to assign a weight to each pixel point on each feature map. The sum of the weights gives the pooling result, and the pooled value contains the information of all pixel points of the original feature map of this channel. At the same time, it adds a non-linear component to the model; The pooling result passes through a multi-layer perceptron and then passes through SiLU activation to obtain the final weight of each channel. The final weight is multiplied by the original feature map to obtain a new feature map. The mathematical expression of the Softmax function is: ; Among them, and are the pixel values of a certain point on a single channel, and N is the number of pixel points of the single-channel feature image.
4. An underwater target detection method integrating lightweight multi-scale and soft communication attention according to claim 3, characterized in that, The multi-scale feature extraction module and the channel soft attention module of the improved neural network model are fused into a multi-scale extraction and soft channel attention layer module; The multi-scale extraction and soft channel attention layer module combines with a bottleneck structure to construct a multi-scale extraction and soft channel attention layer module based on the bottleneck structure.
5. The underwater target detection method integrating lightweight multi-scale and soft communication attention according to claim 4, characterized in that, The conventional integrated convolutional layer module, C2f layer module, and target detection head module adopt the structures in the YOLO target detection model. The conventional integrated convolutional layer module integrates two-dimensional convolution, batch normalization, and the SiLU activation function; the C2f layer module uses a convolutional layer to adjust the channels of the feature map, divides the channels of the adjusted feature map into two parts, one part is directly output, and the other part is output after passing through multiple bottleneck structures, and then the feature map is concatenated in the channel dimension through a concatenation layer, and finally the number of channels is restored to the number of channels before the Split layer; the target detection head module has two branches, each branch consists of several consecutive convolutional layers, and the output after the feature map is processed by the target detection head is used to calculate the loss function.
6. The underwater target detection method integrating lightweight multi-scale and soft communication attention according to claim 5, wherein, The calculation of the loss function includes calculating the classification loss, the complete intersection over union loss, and the distribution focal loss.
7. The underwater target detection method integrating lightweight multi-scale and soft communication attention as claimed in claim 4, wherein, The multi-scale extraction and soft channel attention layer module contains three downsampling convolutional branches. The convolutional kernel sizes (k) of the downsampling convolutional branches are 3, 5, and 7 respectively, the stride (s) is 2 for all, and the padding (p) sizes are 1, 2, and 3 respectively. The downsampling convolutional branches use convolutional kernels of different scales to extract features from the feature map. The three downsampling convolutional branches have different receptive fields, which are 9, 25, and 49 respectively. After subsequent splicing and fusion, it can take into account both the local details and global information of the image.
8. The underwater target detection method integrating lightweight multi-scale and soft communication attention according to claim 4, wherein The multi-scale extraction and soft channel attention layer module based on the bottleneck structure first uses ConvD to reduce the dimension of the channels of the input feature map, then performs multi-scale extraction, and then uses ConvU to increase the dimension of the channels of the feature map, which is used to limit the network scale when the number of channels of the feature map is too large.
9. The underwater target detection method integrating lightweight multi-scale and soft communication attention according to claim 8, characterized in that The multi-scale extraction and soft channel attention layer module based on the bottleneck structure reserves the c-down and c-up parameters to facilitate the control of the reduction and increase multiples of the channels. The number of channels of the feature map before and after being processed by ConvD are respectively: output channels = input channels × c-down, where 0 < c-down ≤ 1; the number of channels of the feature map before and after being processed by ConvU are respectively: output channels = input channels × c-up, where c-up ≥ 1; Among them, c-down is the channel reduction multiple, and c-up is the channel increase multiple.
Citation Information
Patent Citations
Seabed organism target detection method based on improved Faster R-CNN
CN112257810A
Seabed target detection method based on BES-YOLOv8 model
CN118736187A
Image-text retrieval method based on multi-scale feature set extraction and alignment
CN119646484A