Image enhancement and sample equalization method for underwater sonar small target detection and medium
Through adaptive bilateral filtering and multi-scale Retinex image enhancement, combined with a single-stage anchor-free object detection network and an adaptive sample allocation strategy, the robustness and accuracy problems of small-scale object detection in underwater sonar images are solved, and efficient underwater target recognition is achieved.
Patent Information
- Application Number
- CN202311493355.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-09
- Publication Date
- 2025-07-08
AI Technical Summary
In the complex and changeable underwater environment, sonar image object detection has problems such as poor robustness, high computational complexity, low detection accuracy and unbalanced samples. Especially in small object detection, the problem of inaccurate recognition is prone to occur.
Adaptive bilateral filtering and multi-scale Retinex image enhancement methods are used to remove speckle noise, and a single-stage anchor-free object detection network is built. Combined with multi-scale feature fusion and dynamic central sample allocation strategy of adaptive target scales, and optimized model training is used using the category self-equilibrium focus loss function.
It improves the accuracy and speed of underwater sonar small object detection, enhances the robustness and generalization ability of the model, improves the proportion of positive and negative samples, and improves the detection accuracy and classification effect.
Smart Images

Figure CN120279348A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of sonar image processing, and relates to an image enhancement and sample balancing method and medium for underwater sonar small target detection. Background Art
[0002] In recent years, with the development of the times and the proposal of the marine power strategy, the detection task of underwater targets has become one of the very active research fields, and its applications are very extensive, including fields such as fish school positioning and submarine detection. However, due to the complex and changeable underwater environment and the difficulties in obtaining and processing underwater signals, the data information collected underwater is often not accurate enough. Therefore, how to accurately perform target detection under the influence of complex and changeable underwater environment and noise is a research problem.
[0003] Underwater images include underwater optical images and sonar images. Underwater optical images have the characteristics of high resolution and color pixels, and can easily observe the actual seabed conditions, but are limited by water quality and shooting distance. Due to the propagation characteristics of sound waves themselves, sonar images are more widely used in underwater detection.
[0004] Traditional sonar image target detection methods mainly rely on manually designed features. Palomeras and Myers et al. have successively proposed sonar image target detection algorithms based on template matching (TM), and used manually designed template features to locate and classify target objects respectively. Since traditional sonar image target detection algorithms need to be manually extracted, they have defects such as poor robustness and high computational complexity. In recent years, researchers have introduced deep learning-based target detection methods into the target detection task of sonar images.
[0005] Deep learning detection algorithms are mainly divided into two-stage and one-stage algorithms. The two-stage algorithm is based on the Region Proposal Network (RPN) and performs detection in two steps. After foreground-background classification, it performs fine-tuning prediction. Its detection accuracy is relatively high, but it has the disadvantages of complex algorithm structure and long inference time. Single-stage object detectors do not need to extract the target area, directly generate prediction boxes on the feature map for regression detection, and have better real-time performance. However, these algorithms use preset anchor boxes (anchor-based), and need to design and adjust the size, aspect ratio and quantity of the preset anchor boxes according to prior knowledge. In multi-scale target detection, the number of anchor boxes generated is large, and the system needs to calculate a large number of hyperparameters, occupying the running memory. FCOS follows the structure of the RetinaNet algorithm, abandons the modeling method of preset anchor boxes, and proposes a modeling method of predicting anchor boxes with the center point, that is, anchor-free. The reduction of hyperparameters such as the scale and aspect ratio of the anchor box makes the inference speed of the FCOS model faster.
[0006] Due to the special underwater environment and the special imaging principle of sonar, the obtained images are affected by speckle noise compared with ordinary visible light images. Moreover, since the number and area ratio of underwater targets in sonar images are small, the detection network tends to learn negative sample information, resulting in insufficient learning of positive sample information and inaccurate target recognition during prediction. Therefore, it is necessary to design a multi-scale target detection algorithm suitable for underwater sonar small target detection tasks. Summary of the Invention
[0007] In view of this, the present invention provides an image enhancement and sample balancing method for underwater sonar small target detection, including:
[0008] Step 1, augment the sample set of sonar images through geometric transformation;
[0009] Step 2, adaptively adjust the spatial standard deviation and grayscale standard deviation of the augmented sample set, and then enhance the sonar images after adjustment to remove the speckle noise therein;
[0010] Step 3, construct a single-stage anchor-free target detection network model to directly perform regression prediction on the feature map to improve the detection speed;
[0011] Step 4, in the target detection network model, the feature fusion module adopts multi-scale feature fusion to identify targets of different scales, identify small targets in the low-level feature map, identify large targets in the high-level feature map, and use accurate low-level positioning signals to enhance the entire feature hierarchy;
[0012] Step 5, for the dataset of sonar images, adopt a dynamic center sample allocation strategy with an adaptive target scale to improve the positive and negative sample allocation ratio;
[0013] Step 6, in the target detection network model, the detection module predicts the underwater target category for each layer of feature map output by the feature fusion module and locates it;
[0014] Step 7, improve the classification loss function in the target detection network model, adopt a class self-balanced focal loss function, calculate the total loss, and update the model parameters through backpropagation.
[0015] Particularly, the specific content of Step 2 includes:
[0016] First, the adaptive adjustment of the spatial standard deviation and the gray standard deviation is achieved by adjusting the photometric similarity characteristics. Then, the sample information in the adjusted sample set is applied to improve the bilateral filter, and the original image is decomposed into images of multiple scales. The center-surround function of Retinex is used to divide the image into a base layer and a detail layer according to the principle of human visual perception, and the base layer and the detail layer are enhanced respectively to eliminate the speckle noise in the samples.
[0017] Specifically, in step 2, the adaptive adjustment of the spatial standard deviation and the gray standard deviation by adjusting the photometric similarity characteristics is achieved as follows: the sample information is realized through the following first equation:
[0018]
[0019] where I i,j is the intensity value of the sample within the image reference window, and a is the depth kernel function for sample pruning; μ ω and σ ω are the standard deviation and the average value of all samples in the image reference window, and the specific definitions are as follows
[0020]
[0021]
[0022] where N is the size of the local window. If μ ω and σ ω satisfy the first equation, then all samples in the local reference window will be retained and processed by bilateral filtering; otherwise, the samples will be deleted. After pruning the samples, a combined filtering weight is generated.
[0023] Specifically, step 3 includes the following steps:
[0024] Construct a single-stage anchor-free object detection network model, where the object detection algorithm framework adopts a single-stage anchor-free and includes three parts: a feature extraction module, a feature fusion module, and a detection module. The feature extraction module is the backbone network, which is used to extract object features, and three-level scale {C3, C4, C5} feature maps are extracted from it, and feature maps with scales of 1 / 8, 1 / 16, and 1 / 32 of the original image scale are output respectively. The feature fusion module fuses multi-scale object features to enhance the semantic information of low-level features. The detection module contains three branches: classification, regression, and localization quality, classifies and identifies underwater objects, and identifies the specific positions of the objects in the sonar image in the form of bounding boxes.
[0025] Specifically, step 4 includes the following:
[0026] In the feature fusion module, each level of feature map is upsampled by a factor of 2 using nearest neighbor interpolation through a 1×1 fully connected layer with horizontal connection. Then, the high-level feature map is added element-wise to the feature map of the immediately lower level, and feature maps with scales of 1 / 64 and 1 / 128 are obtained through downsampling, resulting in five-level scale feature maps {P3, P4, P5, P6, P7}, corresponding to feature maps of sizes 1 / 8, 1 / 16, 1 / 32, 1 / 64, and 1 / 128 of the output scale respectively, and then multi-scale cross-feature fusion is performed. Since different input features have different influences on the output at different scales, learnable weights are introduced to learn the influence of different input features.
[0027] Specifically, step 5 includes the following:
[0028] Adjust the setting range of positive sample points according to the target aspect ratio, adaptively set the sampling area, increase the positive and negative sample ratio, and obtain more feature information. Adaptive setting of the sampling area according to the target aspect ratio includes:
[0029] Assume that the center coordinates of the target ground truth box are (x, y), the width and height are w and h respectively, and the aspect ratio is r = h / w. When r > 1, a square area is set as the central expansion area S with the target center as the midpoint according to the aspect ratio, and its side length w S is
[0030]
[0031] where μ is a smoothing factor to prevent the expansion area from moving away from the center point and can be adjusted according to the actual data situation.
[0032] Specifically, in step 6, the detection module makes predictions on each layer of the feature map output by the feature fusion module, including three branches: the bounding box regression branch, the class prediction branch, and the localization quality prediction branch, so as to realize predicting the underwater target category and localizing it. The localization quality prediction branch uses centerness to suppress low-quality prediction boxes far from the target center and improve the target localization accuracy. The definition of centerness is
[0033]
[0034] where l*, r*, t*, and b* correspond to the distances from the predicted center point to the left, right, top, and bottom sides of the predicted box respectively, and the square root is used to slow down the attenuation of centerness, with the numerical range from 0 to 1. The loss function of the regression branch uses the intersection over union loss function (IoU Loss).
[0035] Specifically, step 7 specifically includes: Among them, the class self-balanced focal loss function is adopted, and the weighting factor and modulation factor are dynamically adjusted according to different classes; in model training, the class self-balanced focal loss function uses the cumulative gradient ratio of positive samples and negative samples to judge the balance degree of positive and negative samples of each class, so as to adjust the factors, balance the loss contributions of different classes, give larger weights to difficult classes, and reduce smaller weights for easy classes;
[0036] Weighted fusion is performed on the loss values of each scale to calculate the overall loss, and the model parameters are updated by backpropagation; if the number of training times reaches the number of iterations, the training ends, otherwise the training continues; the generated object detection model can be directly used for the underwater sonar small target detection task and can be directly deployed on computer equipment.
[0037] Specifically, the class self-balanced focal loss function is defined as:
[0038]
[0039] where p t represents the sample t and its predicted probability, C is the total number of target classes; α t and γ c are the weighting factor and modulation factor respectively, γ c and α c are self-balanced coefficients, and the weighting factor and modulation factor will be dynamically adjusted according to different classes.
[0040] The present invention also proposes a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the image enhancement and sample balancing method for underwater sonar small target detection is realized.
[0041] Beneficial effects:
[0042] (1) Based on adaptive bilateral filtering to improve multi-scale Retinex image enhancement, removing the influence of speckle noise on sonar images;
[0043] (2) Adaptive target scale dynamic center sample allocation strategy, improving positive and negative sample allocation, preventing deviation from the center point while increasing the sampling range, and strengthening target feature learning;
[0044] (3) The class self-balanced focal loss function can adaptively adjust the loss weight according to the sample balance situation of the detected target class.
[0045] (4) Construct a single-stage anchor-free object detection network framework, directly generate prediction boxes on the feature map, and improve the detection speed;
[0046] (5) Feature fusion for cross-scale information exchange to identify targets of different scales;
[0047] (6) Data augmentation improves the robustness and generalization of the model;
[0048] (7) The detection branch predicts the target classification and bounding box regression;
[0049] (8) The centrality branch suppresses low-quality prediction boxes and improves the localization accuracy;
[0050] (9) The improved classification loss function balances the weights of the loss functions for different category samples. Description of the Drawings
[0051] Figure 1 is a flowchart for implementing an image enhancement and sample balancing method for underwater sonar small target detection;
[0052] Figure 2 is a framework diagram of a multi-scale underwater sonar small target detection model;
[0053] Figure 3 is a multi-scale cross-feature fusion structure diagram. Detailed Implementation Manner
[0054] The following takes the drawings and examples as an illustration to describe the present invention in detail.
[0055] As Figure 1 shown, the specific implementation steps of the image enhancement and sample balancing method for underwater sonar small target detection are as follows:
[0056] (1) Augment sonar images to improve the robustness and generalization of the model.
[0057] The larger the image data, the more information is learned, and the higher the robustness and generalization of the model will be. To enrich the data, the present invention generates some new images through image transformation to augment the data set. According to different augmentation principles, it can be mainly divided into geometric transformation, color transformation, sharpness transformation, noise injection and other transformations. Since the data set of the present invention is sonar images, sonar images are low-frequency images and lack high-frequency details compared with optical images; the gray level series of the background noise part of sonar images is less, and data augmentations such as color transformation and noise interference types will affect the target features. Therefore, the present invention mainly uses geometric transformation to process the pictures and augment the data set.
[0058] Geometric transformation enriches information such as the spatial position, perspective, and scale of image data while ensuring that the original semantic information remains unchanged. It mainly includes various operations such as mirroring, random rotation, and random cropping. Mirroring can include horizontal mirroring and vertical mirroring. Random rotation can increase the perspective and position information of the image. To ensure the uniform input size of the image, zero values are used to fill the non-real image positions after rotation, and the rotation angle range is set between 0° and 30°. Randomly cropping a part of the image can change the relative position of the target. To prevent the cropped image from containing only the background, the cropping size is set to three-quarters of the original image.
[0059] (2) Image enhancement based on the improved multi-scale Retinex method using adaptive bilateral filtering adjusts the photometric similarity weight characteristics to achieve the adaptive adjustment of the spatial standard deviation and the gray standard deviation, and removes the influence of speckle noise on the sonar image.
[0060] First, the adaptive adjustment of the spatial standard deviation and the gray standard deviation is achieved by adjusting the photometric similarity characteristics, and then the adjusted sample information is applied to improve the bilateral filter. The processing of the sample information of the fast adaptive threshold is defined as follows:
[0061]
[0062] where, I i,j is the intensity value of the sample within the image reference window, and a is the depth of sample pruning. μ ω and σ ω are the standard deviation and the average value of all samples in the image reference window, and the specific definitions are as follows
[0063]
[0064]
[0065] where, N is the size of the local window. If μ ω and σ ω satisfy Equation (1), then all samples in the local reference window will be retained and processed through bilateral filtering; otherwise, the samples will be deleted. After pruning the samples, a combined filtering weight is generated. The depth kernel function a is constructed as follows:
[0066]
[0067] where, σ h is the standard deviation value of the entire sonar image, and σ ω is the standard deviation of the local reference window. The smaller the intensity weight value θ of the truncated sample, the stronger the ability to truncate the sample information. If the filtered pixel is a pixel at the edge of the image, σ ω is greater than σ hIf the value of a is larger, the number of automatically truncated sample information is smaller, and the edge information and texture of the image are well preserved. Otherwise, more pixels deviating from the image average value will be eliminated, and the number of automatically truncated samples will be large, which can be very smooth speckle noise to a great extent.
[0068] For the improved bilateral filtering, the original image is decomposed into images of multiple scales, which are used as the center-surround function of the Retinex method, and the speckle noise, especially the strong speckle noise, can be largely eliminated.
[0069] The image enhancement method based on the Retinex theory divides the image into a basic layer and a detail layer according to the principle of human visual perception, and performs enhancement processing on the basic layer and the detail layer respectively. Specifically,
[0070] I = L × R (5)
[0071] where, I is the original image; L is the basic layer of the image; R is the detail layer of the image. For easy calculation, it can be transformed into
[0072] lnI = lnL + lnR (6)
[0073] where, the basic layer L is the filtering result of the center-surround function on the image I.
[0074] (3) Construct a single-stage anchor-free object detection network framework, which does not need to extract the target area, directly generates prediction boxes on the feature map for regression detection, and has better real-time performance, such as Figure 2 as shown.
[0075] The object detection algorithm framework adopts a single-stage anchor-free structure and consists of three parts: feature extraction, feature fusion, and detection module. The feature extraction module is the backbone network, which is used to extract target features, and extracts three-level scale {C3, C4, C5} feature maps from it, and outputs feature maps with scales of 1 / 8, 1 / 16, and 1 / 32 of the original image scale respectively; the feature fusion module fuses multi-scale target features to enhance the semantic information of low-level features; the detection module contains three branches: classification, regression, and localization quality, classifies and recognizes underwater targets, and identifies the specific positions of the targets in the sonar image in the form of bounding boxes.
[0076] (4) Feature fusion performs cross-scale cross-fusion, recognizes small targets in the low-level feature map, recognizes large targets in the high-level feature map, and uses accurate low-level localization signals to enhance the entire feature hierarchy, thereby shortening the information path between the low-level and top-level features.
[0077] Each level of the feature layer is upsampled by a factor of 2 using the nearest neighbor interpolation method through a horizontally connected 1×1 fully connected layer. The scale of the high-level feature map is then enlarged by a factor of 2, and the high-level feature map is added to the feature map of the next lower layer element by element. Feature maps with scales of 1 / 64 and 1 / 128 are obtained through downsampling, resulting in five levels of scale feature maps {P3, P4, P5, P6, P7}, corresponding to feature maps with sizes of 1 / 8, 1 / 16, 1 / 32, 1 / 64, and 1 / 128 of the output scale, respectively. Then, multi-scale cross-feature fusion is performed. Taking the P6-level scale feature as an example, the intermediate node of the sixth layer is calculated using the feature of the next layer and the input feature of the sixth layer. Then, the fused feature of the sixth layer is calculated using the input of the sixth layer, the output of the intermediate node, and the output of the previous layer, that is
[0078]
[0079]
[0080] Among them, P6 in and P7 in are the input layer features corresponding to P6 and P7 levels in Figure 2 .3 respectively; P6 td is the intermediate layer feature corresponding to P6 level in Figure 3 ; P5 out and P6 out correspond to the output layer features of P5 and P6 in Figure 2 .3 respectively; Resize is upsampling or downsampling to adjust the feature scales of different layers to be consistent; Conv represents convolution; w i and w’ i are learnable weights, where i takes values from 1 to 3, corresponding to the weights of different hierarchical features; ε is a value approaching 0 to prevent the denominator from being 0. The influence of different input features on the output at different scales is different. Introducing learnable weights can learn the influence of different input features.
[0081] (5) Improve the positive and negative sample assignment, and propose a dynamic center sample assignment strategy with an adaptive target scale to prevent deviation from the center point while increasing the sampling range.
[0082] The number of positive samples is much smaller than the number of negative samples. This will cause the network to be biased towards learning negative sample information, resulting in insufficient learning of positive sample information and inaccurate target recognition during prediction.
[0083] The present invention proposes a dynamic center sample assignment strategy, which adjusts the setting range of positive sample points according to the aspect ratio of the target to increase the positive-negative sample ratio and obtain more feature information. Before the improvement, the sampling range needs to be within the ground truth box. For the tall and thin shape characteristics of composite insulators, it is necessary to increase the positive-negative sample ratio while ensuring the quality of positive samples. Therefore, the present invention adaptively sets the sampling area according to the aspect ratio of the target.
[0084] Assume that the center coordinates of the target ground truth box are (x, y), the width and height are w and h respectively, and the aspect ratio is r = h / w. When r > 1, a square area is set as the central expansion area S with the target center as the midpoint according to the aspect ratio, and its side length is w S is
[0085]
[0086] where μ is a smoothing factor to prevent the expansion area from moving away from the center point and can be adjusted according to the actual data situation.
[0087] (6) Detection branch, predicting the underwater target category and localizing it.
[0088] The detection module makes dense predictions on the features of each layer of the feature fusion module, including the bounding box regression branch, the class prediction branch, and the localization quality prediction. The bounding box regression branch obtains the output of the target prediction box. The localization quality prediction branch uses Centerness to suppress low-quality prediction boxes far from the target center and improve the target localization accuracy. The definition of Centerness is
[0089]
[0090] where l*, r*, t*, and b* respectively correspond to the distances from the predicted center point to the left, right, upper, and lower sides of the predicted box. Taking the square root is used to slow down the attenuation of Centerness, and the value range is from 0 to 1. The Centerness branch uses the binary cross-entropy loss function, and the loss function of the regression branch uses the IoU Loss.
[0091] (7) Improve the classification loss function and propose the class self-balanced focal loss function.
[0092] To solve the problem of class-imbalanced samples, the classification loss function is often used to re-weight the loss value to enhance the recognition accuracy of difficult samples. The Focal loss function can reduce the weight of simple negative samples and make the network focus on the few-sample classes. Therefore, the Focal loss is generally used as the classification loss function. However, the Focal loss is not suitable enough for the composite insulator detection scenario.
[0093] In view of the defects of the two hyperparameters, the weighting factor and the modulation factor, of the Focal loss classification loss function in the composite insulator detection task, the present invention makes improvements. Therefore, the class self-balanced focal loss function (AFL) is proposed, and its definition is
[0094]
[0095] where p tDenote the sample \(t\) and its predicted probability, where \(C\) is the total number of target categories; \(\alpha\) t and \(\gamma\) c are the weighting factor and modulation factor respectively, which are the same as the Focal loss; \(\gamma\) c and \(\alpha\) c are self-balancing coefficients, which will dynamically adjust the weighting factor and modulation factor according to different categories. During model training, the cumulative gradient ratio of positive and negative samples better reflects the imbalance degree of positive and negative samples than the ratio of the number of positive and negative samples. Therefore, the category self-balancing focal loss function uses the cumulative gradient ratio of positive and negative samples to judge the balance degree of positive and negative samples of each category, so as to adjust the factors, \(\gamma\) c and \(\alpha\) c The expressions are respectively
[0096]
[0097]
[0098] where respectively represent the cumulative gradients of positive and negative samples of the \(c\)-th category. The more balanced the training of simple category samples is, the smaller the value of its \(\gamma\) c will be, and the proportion of the loss value will decrease; while the training of difficult category samples is more unbalanced, the value of its \(\gamma\) c is larger, and the proportion of its loss value will increase. \(\alpha\) c and \(\gamma\) c are consistent, which is used to balance the loss contributions of different categories, giving larger weights to difficult categories and reducing smaller weights for simple categories.
[0099] Weight and fuse the loss values of each scale, calculate the overall loss, and update the model parameters by backpropagation. If the number of training times reaches the number of iterations, the training ends; otherwise, continue training to generate a target detection model, which can be directly used for the underwater sonar small target detection task and directly deployed on computer devices.
[0100] The present invention also proposes a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the image enhancement and sample balancing method for underwater sonar small target detection as described above is implemented.
[0101] In summary, the above is only a preferred embodiment of the present invention, and is not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0102] For those skilled in the art, it is obvious that the embodiments of the present invention are not limited to the details of the above-mentioned exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the embodiments of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the embodiments of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the embodiments of the present invention. Any reference signs in the claims should not be construed as limiting the claims involved. In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units, modules or devices described in the system, apparatus or terminal claims can also be implemented by the same unit, module or device through software or hardware. The words such as "first" and "second" are used to denote names and do not represent any particular order.
[0103] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of the present invention and not to limit them. Although the technical solutions of the embodiments of the present invention have been described in detail with reference to the above preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the embodiments of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An image enhancement and sample equalization method for underwater sonar small target detection, characterized in that, Including: Step 1: Augment the sample set of sonar images through geometric transformation; Step 2: Perform adaptive adjustment of the spatial standard deviation and grayscale standard deviation on the augmented sample set, and then enhance the sonar images after adjustment to remove the speckle noise therein; Step 3: Construct a single-stage anchor-free object detection network model, and directly perform regression prediction on the feature map to improve the detection speed; Step 4: In the object detection network model, the feature fusion module adopts multi-scale feature fusion to identify objects of different scales, identify small objects in the low-level feature map, identify large objects in the high-level feature map, and use accurate low-level positioning signals to enhance the entire feature hierarchy; Step 5: For the dataset of sonar images, adopt a dynamic center sample allocation strategy with an adaptive object scale to improve the positive and negative sample allocation ratio; Step 6: In the object detection network model, the detection module predicts the underwater object category for each layer of the feature map output by the feature fusion module and locates it; Step 7: Improve the classification loss function in the object detection network model, adopt a class self-balanced focal loss function, calculate the overall loss, and update the model parameters through backpropagation.
2. The image enhancement and sample equalization method for underwater sonar small target detection according to claim 1, characterized in that, Specifically included in Step 2: First, achieve the adaptive adjustment of the spatial standard deviation and grayscale standard deviation by adjusting the photometric similarity characteristics, and then apply the sample information in the adjusted sample set to improve the bilateral filter to decompose the original image into images of multiple scales; use the center surround function of the Retinex method to divide the image into a basic layer and a detail layer according to the principle of human visual perception, and perform enhancement processing on the basic layer and the detail layer respectively to eliminate the speckle noise in the samples.
3. The image enhancement and sample equalization method for underwater sonar small target detection according to claim 2, characterized in that, The adaptive adjustment of the spatial standard deviation and grayscale standard deviation of the sample information by adjusting the photometric similarity characteristics in Step 2 is achieved through the following first equation: |I i,j -μ ω |≤a·σ ω , Among them, I i,j is the intensity value of the sample within the image reference window, and a is the depth kernel function for sample pruning; μ ω and σ ω are the standard deviation and mean value of all samples in the image reference window, and the specific definitions are as follows where N is the size of the local window, if μ ω and σ ω satisfy the first equation, all samples in the local reference window will be retained and processed by bilateral filtering; otherwise, the samples will be deleted; after pruning the samples, a combined filtering weight is generated.
4. The image enhancement and sample equalization method for underwater sonar small target detection according to claim 1, characterized in that, Step 3 includes the following steps: Construct a single-stage anchor-free object detection network model, where the object detection algorithm framework adopts a single-stage anchor-free box, including a feature extraction module, a feature fusion module, and a detection module; the feature extraction module is the backbone network for extracting object features, and extracts three-level scale {C3, C4, C5} feature maps from it, and outputs feature maps of 1 / 8, 1 / 16, and 1 / 32 of the original image scale respectively; The feature fusion module fuses multi-scale object features and enhances the semantic information of the low-level features; the detection module includes three branches: classification, regression, and localization quality, classifies and identifies underwater objects, and identifies the specific position of the object in the sonar image in the form of a bounding box.
5. The image enhancement and sample equalization method for underwater sonar small target detection according to claim 1, characterized in that, Specifically included in Step 4: In the feature fusion module, feature maps at all levels are upsampled by a factor of 2 using nearest neighbor interpolation through a 1×1 fully connected layer with horizontal connections. Then, the high-level feature map is added element-wise to the feature map of the immediately lower layer, and feature maps with scales of 1 / 64 and 1 / 128 are obtained through downsampling, resulting in five levels of scale feature maps {P3, P4, P5, P6, P7}, corresponding to feature maps of sizes 1 / 8, 1 / 16, 1 / 32, 1 / 64, and 1 / 128 of the output scale, respectively, and then multi-scale cross-feature fusion is performed; since different input features have different influences on the output features at different scales, learnable weights are introduced to learn the influence of different input features.
6. The image enhancement and sample equalization method for underwater sonar small target detection according to claim 1, characterized in that, Step 5 specifically includes: Adjust the range of positive sample points according to the target aspect ratio, adaptively set the sampling area, increase the positive and negative sample ratio, and obtain more feature information; adaptively setting the sampling area according to the target aspect ratio includes: Assume that the center coordinates of the target ground truth box are (x, y), the width and height are w and h respectively, and the aspect ratio is r = h / w. When r > 1, a square region is set as the central expansion region S with the target center as the midpoint according to the aspect ratio, and its side length is w S is where μ is a smoothing factor to prevent the expansion area from moving away from the center point and can be adjusted according to the actual data situation.
7. The image enhancement and sample equalization method for underwater sonar small target detection according to claim 1, characterized in that, In step 6, the detection module makes predictions on each layer of the feature map output by the feature fusion module, including three branches: the bounding box regression branch, the class prediction branch, and the localization quality prediction branch, so as to realize predicting the underwater target class and localizing it; the localization quality prediction branch uses centerness to suppress low-quality prediction boxes far from the target center and improve the target localization accuracy. The definition of centerness is where l*, r*, t*, and b* correspond to the distances from the predicted center point to the left, right, top, and bottom sides of the predicted box respectively. Taking the square root is used to slow down the attenuation of centerness, and the numerical range is from 0 to 1; the loss function of the regression branch uses the intersection over union loss function (IoU Loss).
8. The image enhancement and sample equalization method for underwater sonar small target detection according to claim 1, characterized in that, Step 7 specifically includes: Among them, the class self-balanced focal loss function is used, and the weighting factor and modulation factor will be dynamically adjusted according to different classes; during model training, the class self-balanced focal loss function uses the cumulative gradient ratio of positive samples to negative samples to judge the balance degree of positive and negative samples of each class, so as to adjust the factors, balance the loss contributions of different classes, assign a larger weight to difficult classes, and reduce the weight of simple classes; The loss values at each scale are weighted and fused to calculate the overall loss, and the model parameters are updated by backpropagation; if the number of training times reaches the number of iterations, the training ends, otherwise the training continues; the generated target detection model can be directly used for the underwater sonar small target detection task and can be directly deployed on a computer device.
9. The image enhancement and sample equalization method for underwater sonar small target detection according to claim 1, characterized in that, The class self-balanced focal loss function is defined as: Among them, p t represents the sample t and its predicted probability, and C is the total number of target categories; α t and γ c are the weighting factor and the modulation factor respectively, and γ c and α c are the self-balancing coefficients, and the weighting factor and the modulation factor will be dynamically adjusted according to different categories.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements the image enhancement and sample balancing method for underwater sonar small target detection as described in any one of claims 1-9.