Underwater target image detection method and device
By designing detection models of multi-feature extraction module and feature fusion module, the problems of low detection accuracy and positioning error of small underwater targets are solved, and more efficient underwater target detection is achieved.
Patent Information
- Application Number
- CN202510509883.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing underwater target detection technology has low detection accuracy when dealing with small underwater targets, and the traditional YOLO v5 model has positioning errors when positioning small targets.
By designing an underwater target image detection method and device, a detection model of a multi-feature extraction module, a sampling module, a feature fusion module and a detection module are adopted. The feature fusion module optimizes the feature map transmission path, uses different fusion ratios to fuse feature maps of different scales, and dynamically adjusts the contribution of features of different scales.
It significantly improves the accuracy and stability of underwater target detection, especially in small target detection, reducing positioning errors.
Smart Images

Figure CN120032118A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of target detection, and in particular, to an underwater target image detection method and device. Background Art
[0002] Due to the complexity of the underwater imaging environment, such as light attenuation, turbulence interference, and noise effects, traditional optical image detection methods are difficult to apply, and sonar imaging is usually required for target recognition. Although sonar technology can break through the limitations of underwater optical imaging, sonar images often have problems such as low resolution, poor contrast, and small target sizes, which pose great challenges to the automatic detection of underwater targets.
[0003] Currently, underwater target detection technology mainly relies on object detection algorithms based on computer vision. Traditional methods usually use methods based on template matching, edge detection, and artificial feature extraction (such as SIFT, HOG, etc.) for target recognition, but these methods have poor adaptability to complex backgrounds and small target detection, and the generalization ability is limited. With the development of deep learning technology, object detection methods based on deep neural networks (such as Faster R-CNN, YOLO, SSD, etc.) have become the mainstream. Among them, the YOLO (You Only Look Once) series of models are widely used in real-time object detection tasks due to their efficient end-to-end detection capabilities.
[0004] However, models such as YOLO v5 still face challenges in underwater small target detection tasks. Since the pixel proportion of underwater small targets in sonar images is relatively small, it is difficult for the model to effectively extract features, resulting in low detection accuracy. In addition, the traditional YOLO v5 model uses Intersection over Union (IoU) as the target localization metric, and this method is not accurate enough when dealing with small targets, which is prone to localization errors. Therefore, there is an urgent need for a method to effectively improve the detection accuracy of underwater target images. Summary of the Invention
[0005] In view of this, this application provides an underwater target image detection method and device to effectively improve the detection accuracy of underwater target images.
[0006] Specifically, this application is implemented through the following technical solutions: The first aspect of this application provides an underwater target image detection method, and the method includes: Obtain an underwater target image of a target water area to obtain an underwater target image dataset; Train a detection model based on the underwater target image dataset; The detection model includes multiple feature extraction modules, sampling modules, multiple feature fusion modules, and multiple detection modules with the same structure; each of the feature extraction modules is used to extract feature maps of different scales from the underwater target image; each of the feature fusion modules fuses the received feature maps of different scales based on different fusion ratios; The underwater target image to be detected is input into the trained detection model to obtain the underwater target detection result.
[0007] A second aspect of the present application provides an underwater target image detection device, the device comprising an acquisition module, a training module and a detection module; Wherein, the acquisition module is used to acquire underwater target images of the target water area to obtain an underwater target image data set; The training module is used to train the detection model based on the underwater target image dataset; The detection model includes multiple feature extraction modules, sampling modules, multiple feature fusion modules, and multiple detection modules with the same structure; each of the feature extraction modules is used to extract feature maps of different scales from the underwater target image; each of the feature fusion modules fuses the received feature maps of different scales based on different fusion ratios; The detection module is used to input the underwater target image to be detected into the trained detection model to obtain the underwater target detection result.
[0008] The underwater target image detection method and device provided by the present application, the feature fusion module optimizes the path of feature map transmission, and adopts different fusion ratios to fuse feature maps of different scales, so that the detection model performs better in multi-scale target detection tasks, especially for small target detection in underwater environments. It has significant advantages. First, traditional feature fusion methods often use fixed weights for multi-scale feature fusion, which cannot fully adapt to the importance differences of features of different scales. The present application introduces different fusion ratios in the feature fusion module, so that the detection model can dynamically adjust the contribution of features of different scales, thereby retaining the key information of the target more accurately. For example, high-resolution feature maps contain richer detail information, while low-resolution feature maps have stronger semantic information. Different detection tasks may have different information requirements for different scales. The use of variable fusion ratios can better balance the impact of information of different scales and improve detection accuracy. Secondly, the feature fusion module of the present application constructs a progressive information transmission path between the feature extraction module and the detection module. Feature maps flow between different fusion modules, ensuring that shallow high-resolution information can be gradually transferred to deep low-resolution features, while deep semantic information can also be supplemented to the shallow layer. Feature interaction is enhanced through cross-scale connection paths, further improving the detection capability of the detection model, especially for small target detection. In addition, different feature fusion modules use different fusion ratios, so that the detection model can adopt the optimal fusion strategy for features at different levels. For example, the first feature fusion module may retain more shallow features to enhance detail information, while subsequent feature fusion modules may be more inclined to combine deep semantic information to improve the robustness of target classification. This hierarchical fusion method avoids information redundancy or loss of key information, enables the detection model to use feature information more efficiently, and improves the accuracy and stability of target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 A flowchart of an underwater target image detection method provided in Example 1 of the present application; Figure 2 A schematic diagram of the structure of the detection model shown in this application; Figure 3 A schematic diagram of a first underwater target image to be detected provided in the present application; Figure 4 A schematic diagram of a detection process of a first underwater target image to be detected provided in the present application; Figure 5 A schematic diagram of underwater target detection results of a first underwater target image provided by the present application; Figure 6 A schematic diagram of a second underwater target image to be detected provided by the present application; Figure 7A schematic diagram of a detection process of a second underwater target image to be detected provided in the present application; Figure 8 A schematic diagram of an underwater target detection result of a second underwater target image provided by the present application; Fig. 9 A schematic diagram of a third underwater target image to be detected provided by the present application; Fig.10 A schematic diagram of a detection process of a third underwater target image to be detected provided in the present application; Fig.11 A schematic diagram of underwater target detection results of a third underwater target image provided by the present application; Fig.12 A schematic diagram of a fourth underwater target image to be detected provided by the present application; Fig.13 A schematic diagram of a detection process of a fourth underwater target image to be detected provided in the present application; Fig.14 A schematic diagram of underwater target detection results of a fourth underwater target image provided by the present application; Fig.15 A schematic diagram of the structure of the underwater target image detection device provided in Example 2 of the present application. DETAILED DESCRIPTION
[0010] Here, exemplary embodiments are described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application.
[0011] The terms used in this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The singular forms of "a", "said" and "the" used in this application are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in this article refers to and includes any or all possible combinations of one or more associated listed items.
[0012] It should be understood that although the terms first, second, third, etc. may be used in the present application to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0013] Specific embodiments are given below to introduce the technical solution of the present application in detail.
[0014] Figure 1 This is a flow chart of the underwater target image detection method provided in Example 1 of the present application. Figure 1 , the method provided in this embodiment may include: S101. Acquire underwater target images of a target water area to obtain an underwater target image dataset.
[0015] Specifically, the target water area is set according to actual needs, and is not limited in this embodiment.
[0016] It should be noted that since in underwater environments, small target detection objects have low resolution and small pixel proportion in underwater target images, the underwater targets detected in this application mostly refer to small targets, such as spheres, human models, tires, square cages, iron barrels, etc.
[0017] In a specific implementation, in one possible implementation, a multi-beam sonar, a side-scan sonar or an optical camera is used to scan the target waters to obtain raw image data containing underwater targets, thereby obtaining an underwater target image dataset. In another possible implementation, the underwater target image dataset of the target waters is obtained by consulting a relevant website or resource database and using the interface provided by the website or resource database.
[0018] S102: Training a detection model based on the underwater target image dataset.
[0019] Among them, the detection model includes multiple feature extraction modules, sampling modules, multiple feature fusion modules, and multiple detection modules with the same structure; each of the feature extraction modules is used to extract feature maps of different scales from the underwater target image; each of the feature fusion modules fuses the received feature maps of different scales based on different fusion ratios.
[0020] The feature map output by the second feature extraction module and the feature map output by the third feature extraction module after being sampled and processed by the second sampling module are transmitted together to the first feature fusion module; the feature map output by the first feature extraction module and the feature map output by the first feature fusion module after being processed by the fourth feature extraction module and the first sampling module are transmitted together to the second feature fusion module; the feature map output by the second feature extraction module, the feature map output by the first feature fusion module after being processed by the fourth feature extraction module, and the feature map output by the second feature fusion module are transmitted together to the third feature fusion module; the feature map output by the third feature extraction module and the feature map output by the third feature fusion module are transmitted together to the fourth feature fusion module; the feature maps output by the second feature fusion module, the third feature fusion module, and the fourth feature fusion module are respectively transmitted to the corresponding detection modules; the feature map output by the first feature fusion module is the input of the second feature fusion module, the feature map output by the second feature fusion module is the input of the third feature fusion module, and the feature map output by the third feature fusion module is the input of the fourth feature fusion module; Specifically, the detection model is used to detect underwater target images and determine the type, position and confidence of the underwater target in the underwater target image. Figure 2 This is a schematic diagram of the structure of the detection model shown in this application. Figure 2 ,The detection model includes an input module, a splicing module, multiple feature extraction modules, a sampling module, multiple feature fusion modules, and multiple detection modules with the same structure.
[0021] Furthermore, the input module is the entrance to the entire detection model. The input module is used to receive underwater target images and pre-process the underwater target images, including data enhancement methods such as random cropping, flipping, and color gamut transformation. In addition, the input module will also adaptively scale underwater target images of different sizes, unify them to the same appropriate size, and send the processed image data to the stitching module.
[0022] The stitching module is connected to the input module. The stitching module is used to receive the image data preprocessed by the input module, slice the processed image data according to a specific step size, and connect four adjacent slices in the channel dimension to obtain a stitched image.
[0023] The first feature extraction module includes the CBL (Convolution Batch Normalization LeakyRelu) module and the CSP (Cross Stage Partial Network) module. The CBL module consists of a convolution layer, a batch normalization layer, and an activation function. The convolution layer is used to extract the features of the input image. The batch normalization layer is used to accelerate the convergence of the detection model and reduce the internal covariate shift. The activation function is used to introduce nonlinearity to the detection model so that the detection model can learn more complex features and finally output the feature map after convolution and activation processing. The CSP module consists of a convolution layer and a cross-stage local network. The convolution layer is used to perform preliminary feature extraction on the input feature map. The cross-stage local network is used to divide the extracted feature map into two parts, one part is directly transmitted, and the other part is convolved and then spliced with the directly transmitted part to obtain a fused feature map.
[0024] Please refer to Figure 2 The first feature extraction module, the second feature extraction module, and the fourth feature extraction module are similar in structure and all include convolutional layers and cross-stage local networks. Although the first feature extraction module and the fourth feature extraction module differ in the number of internal convolutional layers included in the cross-stage local networks, their working principles are similar. Therefore, in this embodiment, the second feature extraction module and the fourth feature extraction module are no longer introduced.
[0025] The third feature extraction module includes a CBL module, an SPP module and a CSP module. The introduction of the CBL module and the CSP module can refer to the above embodiment. The SPP module is a spatial pyramid pooling module. The SPP module is used to apply the maximum pooling operation of different scales to the input feature map, and fuse the features from different receptive fields, so that the detection model can obtain multi-scale context information and enhance the adaptability to targets of different sizes.
[0026] The first sampling module is an upsampling module, which is used to upsample the feature map output by the first feature fusion module so that its size matches the size of the feature map output by the first feature extraction module. The size of the feature map output by the first feature fusion module is lower than the size of the feature map output by the first feature extraction module, and the two feature maps are spliced in the channel dimension.
[0027] The second sampling module is an upsampling module, which is used to upsample the feature map output by the third feature extraction module so that its size matches the size of the feature map output by the second feature extraction module. The size of the feature map output by the third feature extraction module is lower than the size of the feature map output by the second feature extraction module, and the two feature maps are spliced in the channel dimension.
[0028] Multiple feature fusion modules are used to fuse feature maps of multiple scales to obtain a fused feature map. The multiple feature fusion modules also include a convolution layer and an activation function. After obtaining the fused feature map, the method further includes: performing a convolution operation on the fused feature map based on the convolution layer, extracting the target features in the feature map using a convolution kernel, and adjusting the number of channels of the fused feature map; applying an activation function to the convolved feature map, performing a nonlinear transformation on each convolution result, and obtaining an optimized feature map.
[0029] As an optional embodiment, before outputting the feature map, each feature fusion module also includes: calculating the sub-block semantics of each sub-block in the feature map; calculating the semantic similarity distance between the semantics of each sub-block and the underwater target to be detected, and calculating the semantic weight based on the ratio of the semantic similarity distance to the semantic similarity distance between the entire feature map and the underwater target to be detected; identifying each sub-block in the feature map with the semantic weight as a mark, and sending the feature map with the semantic weight mark to the next feature fusion module; the next feature fusion module determines the feature extraction depth and granularity of each sub-block according to the semantic weight of each sub-block in the received feature map, wherein the feature extraction depth and granularity are proportional to the semantic weight, and the larger the semantic weight, the deeper the scale of feature extraction and the finer the granularity of feature extraction in the sub-block.
[0030] For example, the following is an example of the feature map output by the first feature fusion module to the second feature fusion module. The first feature fusion module divides the feature map extracted by the second feature extraction module and the third feature extraction module into multiple non-overlapping sub-blocks. For each sub-block, the sub-block semantics of the sub-block is calculated by analyzing the feature vector in the sub-block, for example, using a predefined feature statistical method (such as mean, variance, etc.) or a more complex semantic analysis model (such as a semantic encoder based on deep learning), and the sub-block semantics represents the abstract representation of the local feature information represented by the sub-block. Further, based on the predefined semantic vectors such as the category and features of the target, the semantic representation of the underwater target to be detected is determined. The semantic similarity distance between each sub-block semantics and the underwater target to be detected is calculated using measurement methods such as cosine similarity and Euclidean distance. At the same time, the semantic similarity distance between the entire feature map and the underwater target to be detected is calculated. Based on the ratio of the semantic similarity distance of each sub-block to the semantic similarity distance of the entire feature map, the semantic weight of each sub-block is calculated. Using the calculated semantic weight as a mark, each sub-block in the feature map is identified, and the feature map with the semantic weight mark is sent to the second feature fusion module. After receiving the feature map with the semantic weight mark, the second feature fusion module determines the feature extraction depth and granularity of each sub-block according to the semantic weight of each sub-block. Since the feature extraction depth and granularity are proportional to the semantic weight, for sub-blocks with larger semantic weights, the second feature extraction module will perform deeper feature extraction, such as increasing the number of convolutional layers or using larger convolution kernels, to obtain richer and more detailed feature information; at the same time, the granularity of feature extraction within the sub-block will be finer, and more subtle feature changes can be captured. For sub-blocks with smaller semantic weights, relatively shallow feature extraction is performed to reduce the amount of calculation.
[0031] Furthermore, the multiple feature fusion modules include a learnable weight module, a weighted sum module, a convolution module, and an activation function module connected in sequence. Among them, the learnable weight module is used to assign learnable weights to each input feature map of different scales. In the underwater target image detection scenario, feature maps of different scales contain information at different levels. For example, a large-scale feature map may contain more detailed information, which is conducive to detecting small targets, while a small-scale feature map has a larger receptive field and can provide global information. The learnable weight module automatically adjusts the weight of each feature map by learning the feature patterns in the data to determine its importance in the fusion process. For example, for some underwater targets, the weight of the large-scale feature map may be higher because the detailed features of the target are more critical to its detection. The learnable weight module receives feature maps from different feature extraction modules or processed by the sampling module as input. For example, the first feature fusion module receives the feature map output by the second feature extraction module and the feature map output by the third feature extraction module and sampled and processed by the second sampling module, and the learnable weight module assigns weights to the two feature maps respectively.
[0032] The weighted summation module is used to perform a weighted summation operation on the input feature maps of different scales according to the weights assigned by the learnable weight module. In this way, feature maps of different importance are fused to obtain a comprehensive feature map. The weighted summation operation can effectively integrate the information of feature maps of different scales, so that the fused feature map contains both rich detail information and global semantic information. The input of the weighted summation module comes from the weighted feature map output by the learnable weight module. After receiving these feature maps, the summation calculation is performed according to the weights, and the fused feature map is output.
[0033] The convolution module is used to perform convolution operations on the weighted summed feature map, and the convolution kernel of the convolution module is a 1×1 convolution kernel. The main function of 1×1 convolution is to adjust the number of channels of the feature map and integrate cross-channel information. Without changing the spatial size of the feature map, the features are linearly transformed to enhance the expressiveness of the features. In underwater target detection, 1×1 convolution can be used to combine information from different channels to extract more representative features. The input of the convolution module is the fused feature map output by the weighted summation module. After 1×1 convolution processing, the feature map with adjusted channel number and integrated information is output.
[0034] The activation function module is used to apply the activation function (such as SiLU activation function) to the feature map after convolution, and perform nonlinear transformation on each convolution result. The activation function can introduce nonlinear factors, so that the detection model can learn more complex feature patterns. The SiLU activation function exhibits different response characteristics under different input values. For larger input values, it is close to linear, retaining the original information of the feature; for smaller input values, it is suppressed to reduce the impact of noise. The input of the activation function module is the feature map output by the convolution module. After being processed by the activation function, the optimized feature map is output, which will be passed to subsequent modules (such as other feature fusion modules or detection modules) as the final output of the feature fusion module.
[0035] The first feature fusion module is used to receive the feature map output by the second feature extraction module and the feature map output by the third feature extraction module after being processed by the CBL module and the second sampling module, and fuse the two feature maps of different scales.
[0036] The second feature fusion module is used to receive the feature map output by the first feature extraction module and the feature map output by the first feature fusion module after being processed by the fourth feature extraction module and the first sampling module, and fuse the two feature maps of different scales.
[0037] The third feature fusion module is used to receive the feature map output by the second feature extraction module, the feature map output by the second feature fusion module after being processed by the CBL module, and the feature map output by the first feature fusion module after being processed by the fourth feature extraction module, and fuse the three feature maps of different scales.
[0038] The fourth feature fusion module is used to receive the feature map output by the third feature extraction module and processed by the CBL module, and the feature map output by the third feature fusion module and processed by the CBL module, and fuse the two feature maps of different scales.
[0039] It should be noted that, referring to the previous description, the feature map output by the first feature fusion module is the input of the second feature fusion module and the third feature fusion module, the feature map output by the second feature fusion module is the output of the third feature fusion module, and the feature map output by the third feature fusion module is the input of the fourth feature fusion module.
[0040] For further information, please refer to Figure 2 , multiple detection modules are connected to the second feature fusion module, the third feature fusion module, and the fourth feature fusion module through the CBL module. The detection module is used to receive the fused feature map output by the corresponding feature fusion module, and predict the category, position, and confidence of the underwater target in the underwater target image in combination with the preset anchor frame. It should be noted that the structure of each detection module is the same and the working principle is also the same.
[0041] Optionally, the feature fusion process of the multiple feature fusion modules includes: establishing a matching relationship between a feature fusion module and a feature extraction module; for each feature fusion module, determining a target feature extraction module that matches the feature fusion module based on the matching relationship, and extracting a feature map of the underwater target image based on the target feature extraction module; there are multiple feature maps, each with a different scale; based on the scale and semantic information of the feature maps, calculating the contribution of each feature map to target detection, determining the weights of different feature maps, and determining a fusion ratio of different feature maps based on the weights; different feature maps have different weights under different feature fusion modules; the feature fusion module performs weighted connection and fusion on the feature maps based on the fusion ratio to obtain a fused feature map.
[0042] In specific implementation, according to the structural design of the detection model, the connection mode of the feature fusion module and the feature extraction module is determined, and a matching relationship is established (for example, the first feature fusion module has a matching relationship with the second feature extraction module and the third feature extraction module). For each feature fusion module, the corresponding target feature extraction module is searched according to the determined matching relationship, and the feature map of the corresponding scale is extracted using the target feature extraction module. The contribution of each feature map to target detection is calculated using a statistical analysis method or a learning mechanism based on training data, and the gradient size of each feature map in back propagation is calculated. A larger gradient means that the feature map has a greater impact on target detection and a greater contribution. Combined with the scale and semantic information of the feature map, the weights of different feature maps are determined based on the fact that the scale of the feature map is proportional to the weight size, and the depth of the semantic information is proportional to the weight size. Different weights are assigned to different feature maps in different feature fusion modules. According to the determined feature map weights, the fusion ratio of different feature maps is calculated. Each feature fusion module adopts a learnable weighted connection method to fuse the corresponding feature maps. Through convolution operations, weighted summation or splicing, it fuses feature maps of multiple scales to obtain the fused feature map for use by subsequent detection modules.
[0043] The method provided in this embodiment, first, by establishing a matching relationship between the feature fusion module and the feature extraction module, can ensure that different feature fusion modules receive information from the most relevant feature extraction module, rather than performing indiscriminate information fusion. This precise matching can reduce the interference of irrelevant information, improve the effectiveness of feature extraction, and enable the network to focus more on the key features of underwater targets. Secondly, in traditional feature fusion methods, feature maps of different scales are often fused at a fixed ratio, which may cause some features to be overemphasized or ignored. The present application calculates the contribution of each feature map to target detection and dynamically adjusts the weights of different feature maps so that the fusion ratio can adapt to the needs of targets of different scales. For example, when detecting small targets, the contribution of high-resolution features may be greater, while for large targets, low-resolution features may be more important. This strategy of adaptive weight adjustment enables the detection model to more reasonably allocate feature weights, thereby improving detection performance. In addition, since the weights of different feature maps are different under different feature fusion modules, this means that the model can adopt different fusion strategies at different levels. For example, shallow features may pay more attention to the edge and texture information of the target, while deep features contain more semantic information. Through weighted connection and fusion, features of different scales can interact effectively, so that high-level features can supplement the detailed information of low-level features, and low-level features can also improve semantic understanding capabilities with the help of high-level features. Especially in complex underwater environments, targets often have scale changes, low contrast and noise interference. This optimized feature fusion method can better adapt to these challenges and improve the robustness of detection. In addition, traditional feature fusion methods may indiscriminately splice or weighted average all scale features, which may cause information redundancy or even introduce invalid information. This application calculates the contribution and reasonably allocates the fusion ratio to ensure that only the most valuable features are effectively utilized, while unimportant features are weakened or ignored, thereby improving the computational efficiency and detection accuracy of the model. In addition, this method can also reduce the risk of overfitting caused by feature redundancy and improve the generalization ability of the detection model in different underwater scenes.
[0044] Optionally, the training of the detection model based on the underwater target image dataset includes: (1) For the prediction box and target box output by each detection module, the NWD metric and IoU metric corresponding to each detection module are calculated; the NWD metric establishes a Gaussian distribution to calculate the similarity between the prediction box and the target box, and the IoU metric represents the degree of overlap between the prediction box and the target box; Specifically, the prediction box refers to the prediction result of the underwater target in the underwater target image by the detection module based on the input feature map, and the target box refers to the actual result of the underwater target contained in the underwater target image. The NWD metric measures the similarity between the prediction box and the target box by establishing a Gaussian distribution model for the prediction box and the target box, calculating the distance between the two. The IoU metric characterizes the degree of overlap between the prediction box and the target box.
[0045] In the specific implementation, for the prediction box and target box output by each detection module, the NWD metric corresponding to each detection module is calculated, including: (1) The target box and the prediction box are modeled as two-dimensional Gaussian distributions respectively.
[0046] In specific implementation, the prediction box is represented as ,in, , is the coordinate of the center point of the prediction box, , is the width and height of the prediction box. For the pixel distribution characteristics, it can be expressed by the ellipse equation: ; in, , is the center point of the ellipse, , ; is the x-axis radius, , is the y-axis radius, ; is the horizontal axis, Is the vertical axis.
[0047] Furthermore, the probability density function of the two-dimensional Gaussian distribution is: ; in, For coordinates ; is the mean, , , is the coordinate of the center point of the prediction box; is the variance, , , is the width and height of the prediction box.
[0048] When the ellipse in the ellipse equation is the probability density function of a two-dimensional Gaussian distribution, the prediction box It can be effectively modeled as a two-dimensional Gaussian distribution.
[0049] (2) Calculate the second-order NWD distance between the two-dimensional Gaussian distribution of the target box and the predicted box.
[0050] In specific implementation, the Wasserstein distance is used to calculate the distance between two two-dimensional Gaussian distributions. The second-order NWD distance between the two-dimensional Gaussian distributions of the target box and the predicted box can be expressed as: ; in, , is the mean of the target box and the predicted box; , is the variance of the target box and the predicted box; , are the mean vectors of two 2D Gaussian distributions.
[0051] Since in this embodiment, the target box and the predicted box are bounding boxes, the second-order NWD distance can be expressed as: ; in, is the second-order NWD distance; , For the target box and the predicted box; is the center point coordinate of the target frame; is the center point coordinate of the prediction box; is the width of the target box; is the height of the target box; is the width of the prediction box; is the height of the prediction box.
[0052] (3) Performing exponential normalization processing on the second-order NWD distance to obtain a normalized NWD metric.
[0053] Specifically, the second-order NWD distance is represented by a normalized exponential, and the normalized NWD metric can be expressed as: ; in, , For the target box and the predicted box; is the second-order NWD distance; is a constant.
[0054] The method provided in this embodiment uses the NWD metric to calculate the loss function. Compared with the traditional IoU metric, the NWD metric is based on the Wasserstein distance principle and calculates the similarity by constructing the Gaussian distribution of the prediction box and the target box, so that the detection model can more finely measure the degree of match between the two, which is particularly suitable for small target detection tasks. Since the IoU metric is too sensitive to slight offsets when the target box overlaps less, it may cause unstable gradient updates, while the NWD metric can provide smoother gradient optimization, which helps the detection model to converge more stably. In addition, the NWD metric takes into account the spatial relationship between the target boxes. Even if the prediction box and the target box do not overlap directly, the similarity between them can still be quantified, thereby improving the positioning accuracy in small target detection. Combined with the IoU metric, NWD makes up for the shortcomings of IoU that is not robust enough for small targets, making the loss function more comprehensive and the optimization process more stable, which helps to improve the detection performance of the detection model for targets in complex underwater environments.
[0055] In specific implementation, for the prediction box and the target box output by each detection module, the IoU metric corresponding to each detection module is calculated, including: calculating the intersection between the target box and the prediction box; calculating the union between the target box and the prediction box; and determining the quotient of the intersection and the union as the IoU metric.
[0056] Specifically, the IoU metric can be expressed as: ; Among them, the is the target frame; is the prediction box; is the intersection between the target box and the predicted box; is the union between the target box and the predicted box.
[0057] (2) Determine a weight factor based on the degree of matching between the predicted box and the target box.
[0058] Specifically, the weight factor refers to the weight of the prediction box, that is, the relative proportion of the prediction box and the target box in the loss function.
[0059] In a specific implementation, determining the weight factor based on the degree of match between the prediction box and the target box includes: determining a first value of the weight factor based on the NWD metric and the IoU metric; calculating the degree of match between the prediction box and the target box, and adjusting the first value of the weight factor based on the degree of match; calculating the loss function under the first value, and adjusting the value of the weight factor based on the size of the loss function.
[0060] Specifically, based on the calculated NWD metric and IoU metric, the ratio of the two is calculated, and the first value of the weight factor is determined according to the ratio. Furthermore, the scale ratio between the prediction box and the target box is calculated to characterize the degree of matching between the two, and the value of the weight factor is adjusted according to the size of the scale ratio. If the scale is relatively large, that is, the matching degree is high, the weight of the IoU metric is increased and the value of the weight factor is reduced. If the scale is relatively small, that is, the matching degree is low, the weight of the NWD metric is increased and the value of the weight factor is increased. After determining the first value of the weight factor, the loss function under the current value is calculated, and the gradient of the loss function is calculated using a gradient descent or adaptive optimization algorithm (such as Adam) to update the weight factor. In multiple training rounds, the weight factor is continuously adjusted to determine the optimal balance point between the NWD metric and the IoU metric. After the training is completed, the optimal weight factor is determined.
[0061] The method provided in this embodiment determines the weight factor based on the matching degree between the prediction box and the target box, which can realize adaptive detection optimization of targets of different scales and improve detection accuracy and robustness. First, the initial weight factor is calculated using the NWD metric and the IoU metric, so that it can comprehensively consider the spatial distribution similarity and regional overlap between the target box and the prediction box to ensure that the initial weight is reasonable. Secondly, the weight factor is dynamically adjusted by calculating the matching degree (such as center point offset, scale ratio, etc.), so that when the target scale is small, it is more dependent on the distribution information of the NWD metric to improve the detection ability of small targets, and when the target scale is large, it is more dependent on the geometric overlap information of the IoU metric to improve the positioning accuracy of large targets. In addition, the weight factor is further optimized by calculating the size of the loss function, so that it gradually converges to the optimal value during the training process, ensuring that the loss is minimized and improving the learning ability of the model. This method avoids the problem of insufficient generalization that may be caused by fixed weight factors, so that the detection model can be adaptively adjusted according to different target features, and enhance the detection ability of multi-scale targets in complex underwater environments.
[0062] (3) Based on the NWD metric, the IoU metric, and the weight factor, calculate the loss function corresponding to each detection module, and adjust the detection model based on the loss function.
[0063] In the specific implementation, the loss function is calculated based on the weight factor, NWD metric, and IoU metric through a linear weighted method: ; Among them, the is the loss function; is the weight factor; is the NWD metric; is the IoU metric.
[0064] Furthermore, after calculating the loss function of each detection module, the parameters of the detection model are adjusted through back propagation, and the network weights are updated based on the loss function, so that the optimization direction of the detection model converges towards the minimum loss. The training steps are repeated until the detection model converges to obtain a trained detection model.
[0065] The method provided in this embodiment combines the NWD metric and the IoU metric to determine the loss function, and dynamically adjusts it through the weight factor, which helps to improve the optimization effect of the underwater target detection model. The IoU metric mainly measures the overlap between the prediction box and the target box, which is suitable for situations where the target size is large and the overlapping part is obvious. However, for small target detection, the IoU gradient vanishing problem is more serious, especially when the prediction box and the target box hardly overlap, IoU cannot provide an effective optimization direction. The NWD metric is based on the similarity between the Gaussian distribution calculation boxes. Even if the prediction box and the target box are offset greatly, it can still provide a continuous and smooth gradient, so that the model can more stably optimize the positioning accuracy of small targets. Therefore, combining the NWD metric and the IoU metric enables the loss function to obtain more effective feedback in different detection scenarios, making up for the shortcomings of a single metric method. In addition, the process of determining the weight factor further enhances the adaptive ability of the loss function. The calculation of the weight factor is based on the matching degree between the prediction box and the target box, so that the loss function can dynamically adjust the contribution ratio of NWD and IoU for different target types. For example, in small object detection, the NWD metric can occupy a higher weight to enhance the fine adjustment of the target position; while for large object detection, the contribution of the IoU metric can be increased to ensure the accuracy of the shape and overlapping areas. In addition, the optimization of the weight factor can also avoid a certain metric dominating the training process, causing the loss function to be biased towards a specific optimization direction, thereby improving the generalization ability of the model and enabling it to adapt to object detection of different scales.
[0066] S103: Input the underwater target image to be detected into the trained detection model to obtain the underwater target detection result.
[0067] Specifically, the underwater target detection result includes the type, position and confidence of the underwater target in the underwater target image.
[0068] In specific implementation, the underwater target image to be detected is input into the trained detection model, and the detection model outputs the underwater target detection result, including the type, location and confidence of the underwater target. Figure 3 A schematic diagram of a first underwater target image to be detected provided in this application, Figure 4 A schematic diagram of the detection process of the first underwater target image to be detected provided in this application, Figure 5 A schematic diagram of underwater target detection results of the first underwater target image provided in this application. Figure 6A schematic diagram of a second underwater target image to be detected provided in this application, Figure 7 A schematic diagram of the detection process of the second underwater target image to be detected provided in this application, Figure 8 A schematic diagram of underwater target detection results of the second underwater target image provided in this application. Fig. 9 A schematic diagram of a third underwater target image to be detected provided in this application, Fig.10 A schematic diagram of the detection process of the third underwater target image to be detected provided in this application, Fig.11 A schematic diagram of underwater target detection results of the third underwater target image provided in this application. Fig.12 A schematic diagram of a fourth underwater target image to be detected provided in this application, Fig.13 A schematic diagram of a detection process of a fourth underwater target image to be detected provided in this application, Fig.14 A schematic diagram of underwater target detection results of the fourth underwater target image provided in the present application.
[0069] The underwater target image detection method provided in this embodiment, on the first hand, the feature fusion module optimizes the path of feature map transmission, and fuses feature maps of different scales with different fusion ratios, making the detection model perform better in multi-scale target detection tasks, especially having significant advantages in small target detection in underwater environments. First of all, traditional feature fusion methods often use fixed weights for multi-scale feature fusion and cannot fully adapt to the importance differences of different scale features. By introducing different fusion ratios in the feature fusion module in this application, the detection model can dynamically adjust the contribution degrees of different scale features, thereby more accurately retaining the key information of the target. For example, high-resolution feature maps contain richer detailed information, while low-resolution feature maps have stronger semantic information. Different detection tasks may have different information requirements for different scales. Using variable fusion ratios can better balance the influence of different scale information and improve the detection accuracy. Secondly, the feature fusion module of this application constructs a progressive information transmission path between the feature extraction module and the detection module. Feature maps flow between different fusion modules, ensuring that the shallow high-resolution information can gradually be transmitted to the deep low-resolution features, and at the same time the deep semantic information can also be supplemented to the shallow layer. The feature interaction is enhanced through the cross-scale connection path, further improving the detection ability of the detection model, especially being more friendly to small target detection. In addition, different feature fusion modules use different fusion ratios, enabling the detection model to adopt the optimal fusion strategy for features at different levels. For example, the first feature fusion module may retain more shallow features to enhance the detailed information, while subsequent feature fusion modules may be more inclined to combine deep semantic information to improve the robustness of target classification. Such a hierarchical fusion method avoids information redundancy or loss of key information, enables the detection model to more efficiently utilize feature information, and improves the accuracy and stability of target detection. On the second hand, combining the NWD metric and the IoU metric to determine the loss function and dynamically adjusting it through a weight factor helps to improve the optimization effect of the underwater target detection model. The IoU metric mainly measures the overlap degree between the predicted box and the target box and is applicable to cases where the target size is large and the overlapping part is obvious. However, for small target detection, the problem of IoU gradient disappearance is relatively serious. Especially when the predicted box and the target box hardly overlap, IoU cannot provide an effective optimization direction. The NWD metric calculates the similarity between boxes based on the Gaussian distribution. Even when the predicted box and the target box are greatly offset, it can still provide a continuous and smooth gradient, enabling the model to more stably optimize the localization accuracy of small targets. Therefore, combining the NWD metric and the IoU metric enables the loss function to obtain more effective feedback in different detection scenarios and makes up for the deficiencies of single metric methods. In addition, the process of determining the weight factor further improves the adaptive ability of the loss function. The calculation of the weight factor is based on the matching degree between the predicted box and the target box, enabling the loss function to dynamically adjust the contribution ratios of NWD and IoU for different target types.For example, in small object detection, the NWD metric can occupy a higher weight to enhance the fine adjustment of the target position; while for large object detection, the contribution of the IoU metric can be increased to ensure the accuracy of the shape and overlapping areas. In addition, the optimization of the weight factor can also avoid a certain metric dominating the training process, causing the loss function to be biased towards a specific optimization direction, thereby improving the generalization ability of the model and enabling it to adapt to object detection of different scales.
[0070] Corresponding to the aforementioned embodiment of an underwater target image detection method, the present application also provides an embodiment of an underwater target image detection device.
[0071] Fig.15 This is a schematic diagram of the structure of the underwater target image detection device provided in Example 2 of this application. Fig.15 , the device provided in this embodiment includes an acquisition module 310, a training module 320 and a detection module 330; Wherein, the acquisition module 310 is used to acquire underwater target images of the target water area to obtain an underwater target image dataset; The training module 320 is used to train a detection model based on the underwater target image dataset; The detection model includes multiple feature extraction modules, sampling modules, multiple feature fusion modules, and multiple detection modules with the same structure; each of the feature extraction modules is used to extract feature maps of different scales from the underwater target image; each of the feature fusion modules fuses the received feature maps of different scales based on different fusion ratios; The detection module 330 is used to input the underwater target image to be detected into the trained detection model to obtain the underwater target detection result.
[0072] The device of this embodiment can be used to perform Figure 1 The steps, specific implementation principles and implementation processes of the method embodiment shown are similar and will not be repeated here.
[0073] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.
[0074] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiment described above is only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present application scheme. A person of ordinary skill in the art can understand and implement it without paying any creative work.
[0075] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A method for underwater target image detection, characterized in that: The method comprises: Acquire underwater target images of target waters to obtain an underwater target image data set; Training a detection model based on the underwater target image dataset; The detection model includes multiple feature extraction modules, sampling modules, multiple feature fusion modules, and multiple detection modules with the same structure; each of the feature extraction modules is used to extract feature maps of different scales from the underwater target image; each of the feature fusion modules fuses the received feature maps of different scales based on different fusion ratios; The underwater target image to be detected is input into the trained detection model to obtain the underwater target detection result.
2. The method according to claim 1, characterized in that The feature fusion process of the multiple feature fusion modules includes: Establishing a matching relationship between the feature fusion module and the feature extraction module; For each feature fusion module, a target feature extraction module matching the feature fusion module is determined based on the matching relationship, and a feature map of the underwater target image is extracted based on the target feature extraction module; there are multiple feature maps, each with a different scale; Based on the scale and semantic information of the feature map, the contribution of each feature map to the target detection is calculated, the weights of different feature maps are determined, and the fusion ratio of different feature maps is determined based on the weights; the weights of different feature maps are different under different feature fusion modules; The feature fusion module performs weighted connection and fusion on the feature maps based on the fusion ratio to obtain a fused feature map.
3. The method according to claim 1, characterized in that The method of training a detection model based on the underwater target image data set comprises: For the prediction box and target box output by each detection module, the NWD metric and IoU metric corresponding to each detection module are calculated; the NWD metric establishes a Gaussian distribution to calculate the similarity between the prediction box and the target box, and the IoU metric represents the degree of overlap between the prediction box and the target box; Determining a weight factor based on a degree of matching between the prediction frame and the target frame; Based on the NWD metric, the IoU metric, and the weight factor, a loss function corresponding to each detection module is calculated, and the detection model is adjusted based on the loss function.
4. The method according to claim 3, characterized in that The determining of a weight factor based on a matching degree between the prediction frame and the target frame includes: Determine a first value of the weight factor based on the NWD metric and the IoU metric; Calculating a matching degree between the predicted frame and the target frame, and adjusting a first value of the weight factor based on the matching degree; Calculate the loss function under the first value, and adjust the value of the weight factor based on the size of the loss function.
5. The method according to claim 2, characterized in that: The plurality of feature fusion modules further include a convolution layer and an activation function. After obtaining the fused feature map, the method further includes: Performing a convolution operation on the fused feature map based on a convolution layer, extracting target features in the feature map using a convolution kernel, and adjusting the number of channels of the fused feature map; Apply the activation function to the convolved feature map, perform nonlinear transformation on each convolution result, and obtain the optimized feature map.
6. The method according to claim 3, characterized in that For the prediction box and target box output by each detection module, the NWD metric corresponding to each detection module is calculated, including: Model the target box and prediction box as two-dimensional Gaussian distributions respectively; Calculate the second-order NWD distance between the two-dimensional Gaussian distribution of the target box and the predicted box; The second-order NWD distance is subjected to exponential normalization processing to obtain a normalized NWD metric.
7. The method according to claim 3, characterized in that For the prediction box and target box output by each detection module, the IoU metric corresponding to each detection module is calculated, including: Calculate the intersection between the target box and the predicted box; Calculate the union between the target box and the predicted box; The quotient of the intersection and the union is determined as the IoU metric.
8. The method according to claim 1, characterized in that The first feature extraction module includes a CBL module and a CSP module. The CBL module is composed of a convolution layer, a batch normalization layer, and an activation function. The convolution layer is used to extract the features of the input image. The batch normalization layer is used to accelerate the convergence of the detection model and reduce the internal covariate shift. The activation function is used to introduce nonlinearity into the detection model so that the detection model can learn multiple features.
9. The method according to claim 1, characterized in that: The feature map output by the second feature extraction module and the feature map output by the third feature extraction module after being sampled and processed by the second sampling module are transmitted together to the first feature fusion module; The feature map output by the first feature extraction module and the feature map output by the first feature fusion module and processed by the fourth feature extraction module and the first sampling module are transmitted to the second feature fusion module; The feature map output by the second feature extraction module, the feature map output by the first feature fusion module and processed by the fourth feature extraction module, and the feature map output by the second feature fusion module are transmitted together to the third feature fusion module; The feature map output by the third feature extraction module and the feature map output by the third feature fusion module are transmitted together to the fourth feature fusion module; the feature maps output by the second feature fusion module, the third feature fusion module, and the fourth feature fusion module are transmitted to the corresponding detection modules respectively; the feature map output by the first feature fusion module is the input of the second feature fusion module, the feature map output by the second feature fusion module is the input of the third feature fusion module, and the feature map output by the third feature fusion module is the input of the fourth feature fusion module.
10. An underwater target image detection device, characterized in that: The device comprises an acquisition module, a training module and a detection module; Wherein, the acquisition module is used to acquire underwater target images of the target water area to obtain an underwater target image data set; The training module is used to train the detection model based on the underwater target image dataset; The detection model includes multiple feature extraction modules, sampling modules, multiple feature fusion modules, and multiple detection modules with the same structure; each of the feature extraction modules is used to extract feature maps of different scales from the underwater target image; each of the feature fusion modules fuses the received feature maps of different scales based on different fusion ratios; The detection module is used to input the underwater target image to be detected into the trained detection model to obtain the underwater target detection result.
Citation Information
Patent Citations
Target labeling method and system based on multi-feature loss function fusion
CN116681921A
Ceramic tile flaw detection method based on deep learning
CN117173117A
Ship target detection method based on improved YOLOv5 deep learning network
CN117523179A
Target detection method and device, electronic equipment and readable storage medium
CN117746359A
On-site auditing equipment intelligent identification method based on improved YOLOv5s
CN118351291A
Cited By
Marine organism intelligent detection system based on multi-scale convolution fusion and YOLOv8
CN120953780A