Underwater target detection method, device and equipment based on sonar image and medium
By using the YOLOV10 network model to process the sonar image sample set, the problem of low detection accuracy in the noise background of traditional methods is solved, and higher underwater target positioning accuracy and classification accuracy are achieved.
Patent Information
- Application Number
- CN202510479569.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-17
AI Technical Summary
Traditional underwater object detection methods based on sonar images are difficult to accurately describe target characteristics in the background of noise, resulting in a decrease in detection accuracy.
Using the YOLOV10 network model, the sonar image sample set is obtained and trained, image features are extracted and learned, and the target position and category probability distribution are predicted based on the sonar image to be tested.
The manual feature engineering steps are simplified and the positioning accuracy and classification accuracy of underwater targets are improved.
Smart Images

Figure CN119992307A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of underwater detection, and in particular to a method, device, equipment and medium for underwater target detection based on sonar images. Background Art
[0002] As an important means of underwater detection, sonar imaging technology plays a key role in underwater operations, marine resource exploration, underwater security and other fields. Underwater target detection based on sonar images refers to the process of using sonar technology to obtain underwater environment images and identify and locate underwater targets from these images.
[0003] Traditional target detection methods based on sonar images mainly rely on artificial feature extraction techniques, such as edge detection and shape matching, which can identify and locate underwater targets to a certain extent. However, due to the complex and changeable underwater environment, sonar images are easily affected by noise interference, water scattering, multipath effects and other factors, resulting in image quality degradation. Artificial features are difficult to accurately describe target characteristics in a noisy background, thus affecting the detection accuracy. Summary of the invention
[0004] The purpose of this application is to provide a method, device, equipment and medium for underwater target detection based on sonar images, which can improve the accuracy of target detection.
[0005] To achieve the above objectives, this application provides the following solutions: In a first aspect, the present application provides an underwater target detection method based on sonar images, comprising: Acquire a sonar image sample set, the sonar image sample set comprising sample images and sample labels attached to each sample image, each sample label comprising position information and category probability distribution, the position information refers to the position information of the underwater target in the sonar image, and the category probability distribution refers to the category probability distribution of the underwater target; Obtaining a training set from the sonar image sample set, and training an initial underwater target detection model through the training set to obtain an underwater target detection model, wherein the initial underwater target detection model is a YOLOV10 network model, and the YOLOV10 network model includes an inverted bottleneck fusion module, and the inverted bottleneck fusion module is integrated with an SE block to adaptively adjust the weight of each channel of the sample image in the feature extraction stage; Acquire a sonar image to be detected, input the sonar image to be detected into an underwater target detection model, and output the position information and the category probability distribution.
[0006] In a second aspect, the present application provides an underwater target detection device based on sonar images, comprising: A first acquisition module is used to acquire a sonar image sample set, wherein the sonar image sample set includes sample images and sample labels attached to each sample image, wherein each sample label includes position information and category probability distribution, wherein the position information refers to the position information of the underwater target in the sonar image, and the category probability distribution refers to the category probability distribution of the underwater target; A training module is used to obtain a training set from the sonar image sample set, and train an initial underwater target detection model through the training set to obtain an underwater target detection model, wherein the initial underwater target detection model is a YOLOV10 network model, and the YOLOV10 network model includes an inverted bottleneck fusion module, and the inverted bottleneck fusion module is integrated with an SE block to adaptively adjust the weight of each channel of the sample image in the feature extraction stage; The input module is used to obtain the sonar image to be detected, input the sonar image to be detected into the underwater target detection model, and output the position information and the category probability distribution.
[0007] In a third aspect, the present application provides a computer device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of any of the above-mentioned methods for underwater target detection based on sonar images.
[0008] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above-mentioned methods for underwater target detection based on sonar images.
[0009] According to the specific embodiments provided in this application, this application discloses the following technical effects: This application applies the YOLOV10 network model in the field of underwater detection. By using the YOLOV10 network model to extract and learn the image features of the sonar image sample set, and predicting the exact location of each potential target and the probability distribution of its category based on the sonar image to be tested, it simplifies the tedious manual feature engineering steps in the traditional method and improves the positioning accuracy and classification accuracy of underwater targets. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0011] Figure 1This is an application environment diagram of an underwater target detection method based on sonar images in one embodiment of the present application; Figure 2 A schematic diagram of a flow chart of an underwater target detection method based on sonar images provided in one embodiment of the present application; Figure 3 A schematic flow chart of a method for underwater target detection based on sonar images provided in yet another embodiment of the present application; Figure 4 A schematic diagram of functional modules of a method and device for underwater target detection based on sonar images provided in one embodiment of the present application; Figure 5 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0012] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0013] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0014] The underwater target detection method based on sonar images provided in the embodiments of the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, or placed on the cloud or other servers. The terminal 102 can send the sonar image sample set to be processed to the server 104. After the server 104 receives the sonar image sample set to be processed, for the sonar image sample set to be processed, the server 104 obtains the training set in the sonar image sample set, and trains the initial underwater target detection model through the training set to obtain the underwater target detection model. And obtain the sonar image to be detected, input the sonar image to be detected into the underwater target detection model, and output the position information and category probability distribution.
[0015] The server 104 may feed back the obtained position information and category probability distribution to the terminal 102. In addition, in some embodiments, the underwater target detection method based on sonar images may also be implemented by the server 104 or the terminal 102 alone. For example, the terminal 102 may directly train the initial underwater target detection model for the training set in the sonar image sample set to be processed, or the server 104 may obtain the training set in the sonar image sample set to be processed from the data storage system, and perform model training for the training set in the sonar image sample set to be processed.
[0016] The terminal 102 may be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, IoT devices, and portable wearable devices. The IoT devices may be smart speakers, smart TVs, smart air conditioners, smart vehicle-mounted devices, etc. The portable wearable devices may be smart watches, smart bracelets, head-mounted devices, etc. The server 104 may be implemented as an independent server or a server cluster consisting of multiple servers, or may be a cloud server.
[0017] In an exemplary embodiment, Figure 2 As shown, a method for underwater target detection based on sonar images is provided. The method is executed by a computer device, and can be executed by a computer device such as a terminal or a server alone, or by a terminal and a server together. In the embodiment of the present application, the method is applied to Figure 1 The server 104 in the example is used as an example to illustrate the method, which includes the following steps S201 to S208. Among them: In step S201, a sonar image sample set is obtained, where the sonar image sample set includes sample images and sample labels attached to each sample image. Each sample label includes position information and category probability distribution. The position information refers to the position information of the underwater target in the sonar image, and the category probability distribution refers to the category probability distribution of the underwater target.
[0018] The location information is the target bounding box coordinates, which are expressed in the form of center coordinates, width, and height.
[0019] Before step S201, the underwater target detection method based on sonar images further includes the following step S200: S200: Acquire an initial sonar image sample set, and preprocess the initial sonar images in the initial sonar image sample set to obtain a sonar image sample set.
[0020] In order to build a comprehensive and diverse sonar image sample set, it is first necessary to use various types of sonar equipment to collect data under various water conditions. These conditions include but are not limited to different depths, water temperatures, and changes in water quality to ensure that the sample set can cover as many diverse environmental conditions as possible. The targets of the collection cover a wide range of underwater entities, such as shipwrecks, marine life, underwater pipelines, etc., to ensure that the diversity of the samples is reflected not only in environmental factors, but also in the richness of target types.
[0021] Select sonar equipment with different frequency ranges, resolutions and detection distances according to the detection requirements to ensure that various targets from large structures (such as shipwrecks) to small biological groups (such as fish schools) can be effectively captured. Sampling is carried out in multiple representative water environments, including but not limited to freshwater lakes, saltwater bays, deep seas, etc. Each environment has unique physical properties (such as temperature, salinity, turbidity), which will affect the propagation of sound waves.
[0022] The purpose of denoising the collected sonar images is to remove interference caused by water scattering, equipment noise, etc., so as to improve the image quality. Adaptive filtering algorithms, such as adaptive Wiener filtering, can dynamically adjust the filter parameters according to the noise characteristics of the local area of the image to achieve better denoising effects. Adaptive Wiener filtering is a method of adjusting filter parameters based on the statistical characteristics of signals and noise. In the local area of the image, the filter coefficient is dynamically determined based on statistics such as the mean and variance of the area. The purpose of this method is to make the filtered image closest to the initial noise-free image in terms of mean square error.
[0023] In one embodiment, the initial sonar image sample set includes a plurality of initial sonar images, and step S200 further includes the following sub-steps S2001-S2002: S2001. De-noising all initial sonar images using an adaptive Wiener filtering algorithm.
[0024] S2002: Perform image enhancement processing on all the initial sonar images that have been denoised to obtain a sonar image sample set.
[0025] In one embodiment, step S2001 includes the following sub-steps S20011-S20015: S20011. Obtain an initial sonar image.
[0026] Specifically, the initial sonar image can be obtained through a public database, or the sonar image of a specific area can be obtained through authorization.
[0027] S20012. Calculate the local mean of each pixel in the initial sonar image.
[0028] Specifically, assuming that the initial sonar image (i.e., the observed noisy image) is express, ,in, represents the initial noise-free sonar image, In practical applications, the observed image It can be represented as the initial noise-free sonar image and noise The purpose of preprocessing is to reduce noise as much as possible. Initial sonar image The effect of noise reduction is improved to improve the image quality and restore or approach the initial noise-free sonar image. status.
[0029] For the initial sonar image, the pixels The size of the center is Local area: Each pixel in the initial sonar image The local mean of is calculated by the following formula: ; in, In pixels The average value of pixels in the local area centered at Express Round down to determine the range of the local area. If n=5, then ; If n=6, then .
[0030] Step S20012 is to estimate the background intensity of the current area, that is, the overall brightness level of the image without considering changes in details.
[0031] S20013. Calculate the local variance of each pixel in the initial sonar image.
[0032] Each pixel in the initial sonar image The local variance of is calculated by the following formula: ; in, It is used to measure the degree of dispersion of pixel values in a local area relative to the local mean. A high variance indicates that the pixel values in the area vary greatly and there are a lot of details or edge information; a low variance indicates that the pixel values in the area vary little, indicating that the area is relatively smooth.
[0033] S20014. Calculate the denoising value of each pixel according to the local mean and local variance of each pixel.
[0034] The adaptive Wiener filter can adjust the filter parameters according to the local characteristics of the image to achieve denoising. Effect, the output of the adaptive Wiener filter for: ; in, represents the image after adaptive Wiener filtering; assuming the variance of the noise is known; In pixels The average value of pixels in the local area centered; Represents each pixel The local variance of ; the adaptive Wiener filter retains the signal part (based on the comparison between the local variance and the noise variance) while reducing the influence of the part that is considered to be noise.
[0035] S20015. Repeat the step of calculating the denoising value of each pixel, traverse each pixel in the initial sonar image, and obtain the initial sonar image after denoising.
[0036] When performing adaptive Wiener filtering, the algorithm traverses every pixel in the initial sonar image. , and calculate the local mean for the pixel and local variance , so that the pixel is denoised. In order to calculate these local statistics, a small neighborhood or window (such as 5×5, 7×7, etc.) is usually defined around each pixel. This window slides across the image so that each pixel calculates statistics based on its surrounding pixels. Since each image may contain different noise patterns and content, each image in the initial sonar image sample set needs to be processed separately to achieve the best denoising effect.
[0037] When processing the initial sonar image sample set, the entire initial sonar image sample set needs to be loaded, and each image in the sample set is processed in turn.
[0038] In sonar images, interference such as water scattering and equipment noise is non-uniform, and the noise characteristics of different areas may vary greatly. Adaptive Wiener filtering can dynamically adjust the filtering parameters according to the noise characteristics of the local area of the image. Compared with fixed parameter filtering methods (such as mean filtering, median filtering, etc.), it can better preserve the edge and detail information of the image, and more effectively remove noise, thereby improving the quality of sonar images and providing a better foundation for subsequent tasks such as target detection.
[0039] In one embodiment, step S2002 includes the following sub-steps S20021-S20023: S20021. Determine the cumulative probability of each gray level appearing in the denoised initial sonar image according to the grayscale cumulative distribution function.
[0040] Specifically, the grayscale cumulative distribution function is expressed by the following formula: ; Among them, for each gray level (where k ranges from 0 to L-1, and L is the number of gray levels), calculate its cumulative probability ; Indicates grayscale Probability of occurrence; Indicates grayscale and all less than The cumulative probability of the gray level appearing in the image.
[0041] S20022. Determine the grayscale value after equalization based on the cumulative probability.
[0042] Specifically, the grayscale cumulative distribution function is linearly transformed to obtain a new grayscale value, that is, the equalized grayscale value, which is expressed by the following formula: ; in, Represents the grayscale value after equalization, Represents the number of gray levels minus one, which is used to scale the cumulative probability to the gray level range.
[0043] S20023. Apply the equalized grayscale value to each pixel of the initial sonar image that has been denoised to generate a sonar image.
[0044] For each pixel in the initial sonar image , according to its initial gray value r, find the corresponding equalized gray value s, replace the gray value of each pixel in the initial sonar image with the corresponding equalized gray value s, and generate the preprocessed image.
[0045] After the initial sonar image is processed by histogram equalization, its grayscale value distribution becomes more uniform, which means that the contrast of the image is enhanced, the details of the dark and bright areas are more clearly visible, and the overall quality of the image is improved.
[0046] In step S202, a training set in the sonar image sample set is obtained, and an initial underwater target detection model is trained by using the training set to obtain an underwater target detection model. The initial underwater target detection model is a YOLOV10 network model.
[0047] In one embodiment, the step S202 of “training the initial underwater target detection model by using the training set to obtain the underwater target detection model” includes the following sub-steps S2021-S2023: S2021, inputting the training set into a feature extraction network to obtain a first feature output, wherein the feature extraction network includes at least one convolution module, at least one bottleneck fusion module, at least one feature splicing module, a spatial pyramid pooling module, and a position squeeze attention module; Specifically, the convolution module includes a first convolution module, a second convolution module, a third convolution module, a fourth convolution module, a fifth convolution module, a sixth convolution module and a seventh convolution module, the bottleneck fusion module includes a first bottleneck fusion module, a second bottleneck fusion module, a third bottleneck fusion module, a fourth bottleneck fusion module and a fifth bottleneck fusion module, the inverted bottleneck fusion module includes a first inverted bottleneck fusion module and a second inverted bottleneck fusion module, the upsampling module includes a first upsampling module and a second upsampling module, and the feature splicing module includes a first feature splicing module, a second feature splicing module, a third feature splicing module and a fourth feature splicing module; In one embodiment, Figure 3 As shown, the step S2021 of "inputting the training set into the feature extraction network to obtain the first feature output" includes the following sub-steps S20211-S202111: S20211, input the sample image in the sample set into the first convolution module Conv1 of the feature extraction network for feature extraction, and obtain the first feature map .
[0048] The sample image in step S20211 is obtained by preprocessing the initial sonar image after denoising and image enhancement. The purpose of preprocessing is to ensure that the image format input to the feature extraction network is consistent. The preprocessing process includes the following sub-steps A1-A4: A1. Scale the sample image to be processed.
[0049] Assume that the size of the sample image to be processed is W The goal is to adjust it to a size where the largest side does not exceed S while keeping the aspect ratio unchanged. Let the scaling ratio be r, which is expressed by the following formula: ; The scaled image size is: = r·W, = r·W; A2. Normalize the scaled sample image to be processed.
[0050] Assume the original pixel value is I(x,y), where x and y represent the horizontal and vertical coordinates of the pixel in the image, respectively. Then the normalized pixel value is (x,y) can be expressed as: ; The purpose of normalization is to convert pixel values from the range [0,255] to [0,1] or another specified range.
[0051] A3. Crop the normalized sample image to be processed to obtain a sample image.
[0052] The purpose of cropping is to extract a fixed-size area from the image. Assume that the cropping size is , calculate the center point position of the image: , ; Cut out , ) as the center and the square area with side length C: = ; in, є[0,1].
[0053] The input sample image X is convolved through a series of convolutional layers to generate the first feature map 1. The calculation formula for convolution processing is: ; in, The first feature map The coordinates of 1, are the coordinates of the convolution kernel, and is the size of the convolution kernel.
[0054] The size of the convolution kernel directly affects the effect of feature extraction. Small-sized convolution kernels (3×3) have a small receptive field and can capture subtle textures and details in the image. In sonar images, when processing images with fine structures (such as the scale texture of small fish), small-sized convolution kernels can more accurately extract these subtle features. Large-sized convolution kernels (5×5 or 7×7) have a larger receptive field and can focus on the overall outline of the target. For example, when detecting targets such as large shipwrecks, large-sized convolution kernels can obtain the overall shape and general structural information of the target.
[0055] In step S20211, the sample image X in the sample set is 5×5 and input into the first convolution module Conv1 of the feature extraction network for feature extraction to obtain the first feature map , the convolution kernel is 3×3, the padding is 1, and the stride is 1. The convolution kernel will slide from left to right and from top to bottom on the sample image, covering a 3×3 area each time, because the stride is 1 and the padding is 1. In Location The value at is calculated using the following formula: ; Where b represents the bias term, which is a constant.
[0056] For each location , traverse the entire input sample image X to generate a complete first feature map .
[0057] S20212, the first feature map Input the second convolution module Conv2 of the feature extraction network to extract features and obtain the second feature map .
[0058] The convolution kernel of the second convolution module Conv2 is 3×3, the padding is 1, and the step size is 1. The specific calculation process is as shown in the previous steps and will not be repeated here.
[0059] S20213. Input the first feature map and the second feature map into the first bottleneck fusion module of the feature extraction network for feature fusion to obtain a third feature map.
[0060] Specifically, the YOLOV10 network model of the present application adds a fusion strategy to the first bottleneck fusion module c2f1, instead of implementing it through simple splicing and direct addition. In the existing bottleneck fusion module, the feature maps extracted from different layers are directly spliced together according to the channel dimension. Although this method increases the receptive field and feature expression ability of the network, it does not consider the differences in the importance of features of different scales. Directly add feature maps of different scales. Although this method can integrate multi-scale information, it may cause important information to be obscured or ignored because the importance of features of each scale has not been adjusted.
[0061] Assumptions is a high-resolution, low-semantic feature map, while It is a low-resolution, high-semantic feature map. In order to achieve multi-scale feature fusion, upsampling and downsampling methods are used to adjust the sizes of the two feature maps so that they can be aligned in the spatial dimension.
[0062] for , apply an average pooling operation with a step size of 2 to reduce its resolution, recorded as: ; for , using bilinear interpolation or upsampling techniques to increase its resolution, denoted as: ; Fusion strategies include: The adjusted feature map and Combined, that is, spliced together according to the channel dimension, expressed as [ ; ].if have channels, have channels, after splicing, channels.
[0063] Through the learnable parameter matrix Perform a linear transformation on the concatenated feature map to obtain a new feature map. The new feature map is passed through the Softmax function to obtain the attention map A. The attention map A is expressed by the following formula: ; Attention Map Each element of represents the two feature maps at the corresponding position. and The relative importance of . The input value is converted into a probability distribution through the Softmax function, so that the sum of all elements is 1, and each element is between 0 and 1. The larger the value, the higher the importance of the corresponding feature at that position.
[0064] Use the attention map A to the feature map and Weighted to obtain the final fusion result , fusion results It is expressed by the following formula: ; In the above formula, It means that each element in the attention map A is associated with Multiply the elements at corresponding positions in weighted. Similarly, Express Weighted, where 1−A is the complement of A, means Finally, the two weighted feature maps are added together to obtain the fused feature map .
[0065] Traditional simple concatenation or direct addition methods do not consider the importance differences of different feature maps. Weighted summation allows different feature maps to be assigned different weights, allowing the model to emphasize certain specific information as needed. By learning the optimal weights, the model can more effectively integrate information from different scales or depths, thereby improving the quality of the final output. Through the learned attention map, the model can focus on the most important areas or features and ignore irrelevant or noisy information, which helps improve recognition accuracy.
[0066] S20214, the third feature map Input the third convolution module Conv3 of the feature extraction network for feature extraction to obtain the fourth feature map .
[0067] The convolution kernel of the third convolution module Conv3 is 3×3, the padding is 1, and the step size is 1. The specific calculation process is as shown in the previous steps and will not be repeated here.
[0068] S20215, the third feature map And the fourth characteristic diagram Input the second bottleneck fusion module c2f2 of the feature extraction network for feature fusion to obtain the fifth feature map .
[0069] The specific calculation process is shown in the above steps and will not be repeated here.
[0070] S20216, the fifth feature map The fourth convolution module Conv4 of the feature extraction network is input to perform feature extraction to obtain the sixth feature map.
[0071] The convolution kernel of the fourth convolution module Conv4 is 5×5, the padding is 1, and the step size is 1. The specific calculation process is as shown in the previous steps and will not be repeated here.
[0072] S20217, the fifth feature map And the sixth characteristic diagram Input the third bottleneck fusion module c2f3 of the feature extraction network for feature fusion to obtain the seventh feature map .
[0073] The specific calculation process is shown in the above steps and will not be repeated here.
[0074] S20218, the seventh characteristic map Input the fifth convolution module Conv5 of the feature extraction network for feature extraction to obtain the eighth feature map .
[0075] The convolution kernel of the fifth convolution module Conv5 is 5×5, the padding is 1, and the step size is 1. The specific calculation process is as shown in the previous steps and will not be repeated here.
[0076] In YOLOv10, multiple convolution modules (Conv) are set up continuously to gradually extract and refine the features in the input data. This helps improve the model's ability to detect objects of different scales, and enhances the model's learning ability and generalization performance. Each convolution layer can capture information at different levels. Shallow convolutions usually learn low-level features such as edges and colors; while deep convolutions can recognize more complex structures, such as partial or overall shapes of objects. By stacking multiple convolution layers, the network can extract increasingly abstract and high-level feature representations from the initial image layer by layer. As the number of convolution layers increases, the receptive field of neurons gradually expands. This means that deep convolution kernels can cover larger image areas, which helps to recognize larger-sized objects or understand global context information.
[0077] S20219, the eighth characteristic map Input the first inverted bottleneck fusion module C2fCIB1 of the feature extraction network for feature optimization to obtain the ninth feature map .
[0078] The eighth feature map of the input To adjust the channel, set The size of is H×W×C. First adjust the number of channels through 1×1 convolution: ; in, Represents a 1x1 convolution operation with a stride of 1 and no padding.
[0079] Apply depthwise separable convolution and dilated convolution to The processing formula is as follows: ; in, represents depthwise separable convolution, denotes a dilated convolution with dilation rate d.
[0080] An inverted bottleneck structure is used to improve computational efficiency and nonlinear expression capabilities: ; Among them, ReLU6 is the activation function, and the output range of the activation function is between 0 and 6 to prevent the gradient from disappearing.
[0081] Add Squeeze-and-Excitation (SE) blocks to enhance important image features and suppress noise: Through the SE block Perform SE processing. The SE processing process is implemented by the following formula: ; in, Represents the sigmoid activation function, ranging from 0 to 1, ReLU activation function. Indicates that a global average pooling operation is performed on each channel, that is, the entire feature map is compressed into one value in the spatial dimension. The dimensions are H×W×C, then The dimensions are 1×1×C. , u represents a channel weight vector, by combining u with Multiply by 1 to adjust the importance of each channel.
[0082] This application adds an SE block to the inverted bottleneck fusion module to adaptively adjust the weight of each channel. The feature extraction network and feature fusion network of the YOLOV10 network model both add an inverted bottleneck fusion module. Since the SE block is integrated in the inverted bottleneck fusion module, the SE block adaptively adjusts the weight of each channel in the input feature map, that is, it acts on the feature map obtained by the previous layer processing. The SE block performs a nonlinear transformation on the global information through the ReLU6 activation function and the sigmoid activation function to generate a weight coefficient for each channel. In this way, the SE block can adaptively emphasize the feature channels that are more important to the current task, while suppressing those less important feature channels, thereby enhancing the important features of the image and suppressing noise. This mechanism enables the YOLOV10 network model to focus more on features that are substantially helpful for underwater target detection.
[0083] Furthermore, the feature extraction network and feature fusion network in this application realize the effective fusion and extraction of multi-scale features through the combination of the inverted bottleneck fusion module and the SE block. This structure not only enhances the perception ability of the YOLOV10 network model for underwater targets of different sizes, but also optimizes the expression ability of the features by adaptively adjusting the channel weights, so that the model can still maintain efficient detection performance when facing complex and changeable underwater environments. In addition, since the SE block only operates on the channel dimension of the feature map, its computational cost is relatively low, which enables the YOLOV10 network model to maintain a faster reasoning speed while maintaining high-precision detection, which is particularly important for underwater detection tasks that require real-time feedback.
[0084] The above processed feature map and Perform weighted summation to obtain the ninth feature map , expressed by the following formula: ; in, Represents a coefficient between 0 and 1.
[0085] S202110, the ninth characteristic map The spatial pyramid pooling module SPPF of the input feature extraction network performs multi-scale pooling and feature integration processing to obtain the tenth feature map .
[0086] Specifically, the SPPF module extracts feature information at different levels through multi-scale maximum pooling operations, and then integrates the feature information at different levels to better capture the spatial information of the input image. Assume that the feature map input to the SPPF module is , dimensions are H×W×C.
[0087] Use different sizes of pooling kernels to align the input feature map Perform a max pooling operation. Use three different pooling scales: k1, k2, and k3. For each pooling scale ki, apply the max pooling operation: ; Indicates the use of a pooling kernel of size ki×ki to align the input feature map Perform the maximum pooling operation.
[0088] Since the size of the feature map will become smaller after the maximum pooling, it is necessary to Upsample back to the original size H×W to match the output feature map Alignment, the upsampling process is expressed by the following formula: ; All upsampled feature maps With the initial feature map Together along the channel dimension: F = [ ; U1; U2; U3], where [;] represents the concatenation operation along the channel dimension. SPPF uses pooling kernels of different sizes to filter the input feature maps. The maximum pooling operation can capture information of different scales of the image. This multi-scale feature extraction capability enables the model to better identify objects of different sizes and shapes. In the object detection task, it helps to improve the detection accuracy of small objects, large objects, and objects of various scales.
[0089] S202111, the tenth feature map The position squeeze attention module PSA of the input feature extraction network is processed by the attention mechanism to obtain the eleventh feature map .
[0090] Specifically, for the input feature map Perform 1×1 convolution and convert to query vector Sum value vector : ; ; Calculate the spatial attention weights by the query vector and the value vector: ; The attention weight of the channel branch is obtained through 1×1 convolution, LayerNom function and Sigmoid function: ; Finally, the attention weight of the spatial branch is obtained through reshape and Sigmoid function: ; The attention weights of the channel branch and the attention weights of the spatial branch are directly multiplied by the input features, and then the results are added to get the final output (the eleventh feature map ): ; By combining the channel dimension (different feature maps) and the spatial dimension (location information within the feature map), this step enables the model to more accurately capture subtle changes and complex structures in the image, thereby improving accuracy in recognition tasks.
[0091] S2022. Input the first feature output into a feature fusion network to obtain a second feature output. The feature fusion network includes at least one convolution module, at least one upsampling module, at least one feature splicing module, at least one bottleneck fusion module and at least one inverted bottleneck fusion module.
[0092] In one embodiment, the first feature output includes the fifth feature map, the seventh feature map, and the eleventh feature map; the second feature output includes the sixteenth feature map, the nineteenth feature map, and the twenty-second feature map. In step S2022, "inputting the first feature output into the feature fusion network to obtain the second feature output" includes the following sub-steps S20221-S202211: S20221, the eleventh characteristic map The first upsampling module UnSample1 of the input feature fusion network is used to increase the size of the high-level semantic feature map to obtain the twelfth feature map .
[0093] Specifically, the bilinear interpolation method is used to input the eleventh feature map For upsampling, assume The size of is H×W×C, and it needs to be sampled to 2H×2W×C. Each pixel in , calculate its input feature map The positions and weights of the four corresponding adjacent pixels in .
[0094] Calculate the difference factor: ; ; Calculate the weights based on the difference factor: ; ; ; ; Use the above four weights to perform weighted averaging on adjacent pixels and get : ; in, Represents the input feature map Zhongyu The positions of the four corresponding adjacent pixels.
[0095] S20222, the twelfth characteristic map and the seventh characteristic graph Input the first feature concatenation module Concat1 of the feature fusion network for concatenation to obtain the thirteenth feature map .
[0096] Specifically, before performing feature concatenation, we first need to ensure that the two input feature maps are consistent in spatial dimensions (i.e., height and width). and If the sizes do not match, one of the feature maps needs to be adjusted to achieve size consistency.
[0097] Upsampling can be used Adjust to Same size: ; ; After alignment and The dimensions are H×W× and H×W× , then the spliced The dimensions are H×W× , =[ ; ], where [;] represents the concatenation operation in the channel dimension.
[0098] S20223, the thirteenth characteristic map The fourth bottleneck fusion module c2f4 of the input feature fusion network performs feature fusion to obtain the fourteenth feature map.
[0099] The specific calculation process is shown in the above steps and will not be repeated here.
[0100] S20224. Input the fourteenth feature map into the second up-sampling module UnSample2 of the feature fusion network to increase the size of the high-level semantic feature map to obtain the fifteenth feature map.
[0101] The specific calculation process is shown in the above steps and will not be repeated here.
[0102] S20225. Input the fifteenth feature map and the fifth feature map into the second feature concatenation module Concat2 of the feature fusion network for concatenation to obtain a sixteenth feature map.
[0103] The specific calculation process is shown in the above steps and will not be repeated here.
[0104] S20226, the sixteenth characteristic map Input the sixth convolution module Conv6 of the feature fusion network for feature extraction to obtain the seventeenth feature map .
[0105] The convolution kernel of the sixth convolution module Conv6 is 7×7, the padding is 1, and the step size is 1. The specific calculation process is as shown in the previous steps and will not be repeated here.
[0106] S20227, the seventeenth characteristic map and the fourteenth characteristic graph Input the third feature concatenation module Concat3 of the feature fusion network for concatenation to obtain the eighteenth feature map .
[0107] The specific calculation process is shown in the above steps and will not be repeated here.
[0108] S20228, the eighteenth characteristic map The fifth bottleneck fusion module c2f5 of the input feature fusion network is used for feature fusion to obtain the nineteenth feature map.
[0109] The specific calculation process is shown in the above steps and will not be repeated here.
[0110] S20229, the nineteenth characteristic map Input the seventh convolution module Conv7 of the feature fusion network for feature extraction to obtain the twentieth feature map .
[0111] The convolution kernel of the seventh convolution module Conv7 is 7×7, the padding is 1, and the step size is 1. The specific calculation process is as shown in the previous steps and will not be repeated here.
[0112] S202210, the twentieth feature map and the eleventh feature map Input the fourth feature concatenation module Concat4 of the feature fusion network for concatenation to obtain the twenty-first feature map .
[0113] The specific calculation process is shown in the above steps and will not be repeated here.
[0114] S202211, the twenty-first feature map Input the second inverted bottleneck fusion module C2fCIB2 for feature optimization to obtain the twenty-second feature map .
[0115] The specific calculation process is shown in the above steps and will not be repeated here.
[0116] S2023. Input the second feature output into the output network to obtain a prediction result, which includes position information and category probability distribution. Compare the difference between the sample label and the prediction result to adjust the model parameters to obtain an underwater target detection model. The output network includes multiple target detection head modules, each of which includes two target detection head sub-modules. The target detection head module is used to detect targets of different scales.
[0117] In one embodiment, the step S2023 of “outputting the second feature into the output network to obtain a prediction result” includes the following sub-steps S20231-S20233: S20231, the sixteenth characteristic map The first target head module is input to obtain two sets of different first target position information and category probability distributions, wherein the first target head module includes a first target head submodule One-to-many Head1 and a second target head submodule One-to-many Head2.
[0118] S20232, the nineteenth characteristic map The second target head module is input to obtain two different sets of second target position information and category probability distribution, wherein the second target head module includes a third target head submodule One-to-many Head3 and a fourth target head submodule One-to-many Head4.
[0119] S20233, the twenty-second feature map The third target head module is input to obtain two sets of position information and category probability distribution of different third targets, wherein the third target head module includes a fifth target head submodule One-to-manyHead5 and a sixth target head submodule One-to-many Head6.
[0120] Specifically, different target head modules are used to predict the bounding box coordinates of different targets. If the sample image obtained contains a shipwreck, a life buoy, and marine life, the bounding box coordinates and category probabilities of the three objects can be detected by three different target head modules (the first target head module, the second target head module, and the third target head module).
[0121] The center coordinates of the detected target bounding box are represented as (x, y), and the width and height of the detected target bounding box are represented as (w, h). These coordinate values are proportional to the feature map and need to be converted to the coordinate space of the original image through a mapping relationship. The model usually makes predictions on the feature map. The feature map is a simplified representation of the original image after a series of convolution, pooling and other operations. The size is much smaller than the original image. Therefore, the model directly predicts on the feature map. The target bounding box coordinates (such as the center point (x, y) and width and height (w, h)) obtained above are proportional values based on the size of the feature map.
[0122] Assume the sample image size is H×W, The size is × , then the coordinate calculation formula of the bounding box in the sample image is:
[0123]
[0124] Among them, (X, Y) represents the center coordinates of the bounding box in the sample image, represents the predicted width of the target bounding box in the sample image, that is, the actual width of the target bounding box. Represents the height of the predicted target bounding box in the sample image, that is, the actual height of the target bounding box.
[0125] For each bounding box, the output network computes an initial score for each class. , c represents the category index, Represents the model's initial estimate of whether a given bounding box belongs to category c. The initial score is converted into a probability distribution through the Softmax function. The Softmax function ensures that the sum of the predicted probabilities of all categories is equal to 1, which can be interpreted as the probability of belonging to each category. The probability that the target belongs to category c is expressed by the following formula:
[0126] in, represents the initial score of the corresponding category c in the output network, and C represents the total number of categories. If an object is predicted by two different target head submodules (such as One-to-many Head1 and One-to-many Head2), it is necessary to compare the probability distributions given by the two target head submodules and select the prediction with a higher probability as the final result. Assume that for the same bounding box, the maximum probability predicted by One-to-many Head1 is , and the maximum probability predicted by One-to-manyHead2 is , then choose the larger one as the final probability:
[0127] The category corresponding to the maximum probability is selected as the target category in the bounding box. For example, if the category corresponding to the maximum probability is a life buoy, then the target in the bounding box can be considered to be a life buoy. Similarly, One-to-many Head3 and One-to-many Head4, as well as One-to-many Head5 and One-to-many Head6 also make predictions for the same bounding box.
[0128] Generate bounding boxes from feature maps, then predict the category of each bounding box, and make the final classification decision based on the predicted probability distribution, which improves the accuracy of target detection and allows the model to dynamically adjust the position of the bounding box to further improve the positioning accuracy. In addition, by comparing the outputs of different target head submodules, the robustness and reliability of the model are enhanced.
[0129] In step S203, the sonar image to be detected is obtained, and the sonar image to be detected is input into the underwater target detection model, and the position information and category probability distribution are output.
[0130] In one embodiment, the underwater target detection method based on sonar images further includes the following steps S301-S302: S301, obtaining a verification set from a sonar image sample set, inputting the verification set into an underwater target detection model, and outputting a verification result corresponding to the sample image in the verification set.
[0131] S302. Measure the difference between the verification result and the sample label according to the cross entropy loss function to guide the underwater target detection model to update the parameters.
[0132] Specifically, the cross entropy loss value of a single sample is It is expressed by the following formula: = log( ) ; in, Indicates the true label of sample i belonging to category c (if it belongs, it is 1, otherwise 0), It represents the probability that the model predicts that sample i belongs to category c.
[0133] The cross entropy loss of the entire validation set is the average of all sample losses, expressed as follows: = ; Among them, N represents the number of samples in the validation set, and C represents the total number of categories.
[0134] Update model parameters through optimization algorithm , the update rule is expressed by the following formula: ; in, represents the model parameters, represents the learning rate, Represents the loss function with respect to the parameter gradient.
[0135] In order to speed up the convergence and avoid falling into the local optimal solution, the momentum term is selected in the parameter update process. When updating the parameters, not only the current gradient direction but also the previous update direction is considered. The update formula is as follows: ; in, represents the speed variable, Represents the momentum parameter (e.g. 0.9), and t represents the number of iterations. Repeat the process from calculating the average cross entropy loss to updating the parameters until the model's loss function reaches the predetermined convergence condition or the maximum number of iterations is reached. During each iteration, the model will gradually learn a better parameter configuration.
[0136] In the actual application scenario of deep-sea resource exploration, the underwater target detection device based on sonar images of this application is mounted on the bottom of an unmanned ship to ensure that the device is stable and can withstand the pressure and water flow impact of the deep-sea environment. Start the sonar image acquisition module, which is equipped with a high-precision sonar sensor and can continuously collect sonar images of the surrounding waters during the scheduled route. Set the acquisition frequency to 1 frame per second to ensure comprehensive coverage without missing any potential resources or biological targets. All collected image data are recorded and stored in the processor in real time for subsequent analysis and backtracking. The collected images are transmitted to the data processing unit in real time by wireless or wired means to ensure the timeliness of data flow. The processor quickly processes the image according to the preset denoising and enhancement algorithms, which may include filtering, contrast enhancement, edge detection, etc. to improve image quality. The processing time is controlled within 0.5 seconds per frame to ensure high efficiency of data processing without affecting subsequent analysis and decision-making.
[0137] The processed image is input into the underwater target detection model. The model outputs the target detection results within 0.2 seconds, including potential underwater mineral resource targets, surrounding marine biological targets, etc. The detected target information is transmitted to the result display and storage module to provide real-time feedback to the operator. The operator views the detection results in real time through the display screen, which may be displayed in the form of graphics, lists or maps for easy understanding and analysis. According to the location and category information of the target, the operator can command the unmanned ship to move closer to the key target area for more detailed detection. The storage module records all the data of this detection in full, including image data, processing results, operation records, etc., to provide a basis for subsequent data analysis. After the detection is completed, the stored data is analyzed in detail, including resource distribution, biodiversity, environmental impact, etc. Based on the analysis results, decision support is provided for resource development, environmental protection, scientific research, etc. A detailed detection report is generated, including images, data, analysis results and suggestions, for reference by relevant parties. During the entire detection process, the unmanned ship and device are monitored in real time to ensure the safe operation of the equipment and to respond to possible failures or abnormal situations in a timely manner. Monitor changes in the surrounding environment, such as water flow, temperature, pressure, etc., to ensure that the exploration activities will not have a negative impact on the marine environment. Through the above process, deep-sea resource exploration operations can be carried out efficiently and accurately, providing a scientific basis for the development and protection of marine resources.
[0138] Based on the same inventive concept, the embodiment of the present application also provides a device for realizing the underwater target detection based on sonar images involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more embodiments of the underwater target detection device based on sonar images provided below can refer to the limitations of the underwater target detection method based on sonar images above, and will not be repeated here.
[0139] In an exemplary embodiment, Figure 4 As shown, a sonar image-based underwater target detection device is provided, comprising: A first acquisition module 410 is used to acquire a sonar image sample set, wherein the sonar image sample set includes sample images and sample labels attached to each sample image, wherein each sample label includes position information and category probability distribution, wherein the position information refers to the position information of the underwater target in the sonar image, and the category probability distribution refers to the category probability distribution of the underwater target; A training module 420 is used to obtain a training set from the sonar image sample set, and train an initial underwater target detection model through the training set to obtain an underwater target detection model, wherein the initial underwater target detection model is a YOLOV10 network model, and the YOLOV10 network model includes an inverted bottleneck fusion module, and the inverted bottleneck fusion module is integrated with an SE block to adaptively adjust the weight of each channel of the sample image in the feature extraction stage; The input module 430 is used to obtain the sonar image to be detected, input the sonar image to be detected into the underwater target detection model, and output the position information and the category probability distribution.
[0140] As an optional implementation, the underwater target detection device of the sonar image further includes: A second acquisition module is used to acquire an initial sonar image; A first calculation module, used for calculating the local mean of each pixel in the initial sonar image; A second calculation module, used for calculating the local variance of each pixel in the initial sonar image; A third calculation module, used for calculating a denoising value of each pixel point according to the local mean and the local variance of each pixel point; A traversal module, used for repeatedly calculating the denoising value of each pixel, traversing each pixel in the initial sonar image, and obtaining the initial sonar image after denoising; A first determination module is used to determine the cumulative probability of each gray level appearing in the initial sonar image after denoising according to the grayscale cumulative distribution function; A second determination module, used to determine the gray value after equalization according to the cumulative probability; A generating module is used to apply the equalized grayscale value to each pixel of the initial sonar image that has been subjected to denoising to generate the sonar image.
[0141] As an optional implementation, the training module 420 is specifically configured to: Inputting the training set into a feature extraction network to obtain a first feature output, wherein the feature extraction network includes at least one convolution module, at least one bottleneck fusion module, at least one feature splicing module, a spatial pyramid pooling module, and a position squeeze attention module; Inputting the first feature output into a feature fusion network to obtain a second feature output, wherein the feature fusion network includes at least one convolution module, at least one upsampling module, at least one feature concatenation module, at least one bottleneck fusion module, and at least one inverted bottleneck fusion module; The second feature output is input into the output network to obtain a prediction result, which includes the position information and the category probability distribution. The difference between the sample label and the prediction result is compared to adjust the model parameters to obtain the underwater target detection model. The output network includes multiple target detection head modules, each of which includes two target detection head sub-modules. The target detection head module is used to detect targets of different scales.
[0142] As an optional implementation, the convolution module includes a first convolution module, a second convolution module, a third convolution module, a fourth convolution module, a fifth convolution module, a sixth convolution module and a seventh convolution module, the bottleneck fusion module includes a first bottleneck fusion module, a second bottleneck fusion module, a third bottleneck fusion module, a fourth bottleneck fusion module and a fifth bottleneck fusion module, the inverted bottleneck fusion module includes a first inverted bottleneck fusion module and a second inverted bottleneck fusion module, the upsampling module includes a first upsampling module and a second upsampling module, and the feature splicing module includes a first feature splicing module, a second feature splicing module, a third feature splicing module and a fourth feature splicing module; The first feature output includes a fifth feature map, a seventh feature map, and an eleventh feature map, wherein the fifth feature map is obtained by inputting the third feature map and the fourth feature map into the second bottleneck fusion module for feature fusion, the seventh feature map is obtained by inputting the fifth feature map and the sixth feature map into the third bottleneck fusion module for feature fusion, and the eleventh feature map is obtained by inputting the tenth feature map into the position squeezing attention module for attention mechanism processing; The second feature output includes a sixteenth feature map, a nineteenth feature map and a twenty-second feature map. The sixteenth feature map is obtained by inputting the fifteenth feature map and the fifth feature map into the second feature splicing module for splicing, the nineteenth feature map is obtained by inputting the eighteenth feature map into the fifth bottleneck fusion module for feature fusion, and the twenty-second feature map is obtained by inputting the twenty-first feature map into the second inverted bottleneck fusion module for feature optimization.
[0143] As an optional implementation, in terms of inputting the second feature output into the output network to obtain the prediction result, the training module 420 is specifically used to: Inputting the sixteenth feature map into a first target head module to obtain the position information and the category probability distribution of two different groups of first targets, wherein the first target head module includes a first target head submodule and a second target head submodule; Inputting the nineteenth feature map into a second target head module to obtain the position information and the category probability distribution of two different groups of second targets, wherein the second target head module includes a third target head submodule and a fourth target head submodule; The twenty-second feature map is input into the third target head module to obtain the position information and the category probability distribution of two groups of different third targets, wherein the third target head module includes a fifth target head sub-module and a sixth target head sub-module.
[0144] As an optional implementation, the position information is target bounding box coordinates, and the target bounding box coordinates are expressed in the form of center coordinates, width, and height.
[0145] As an optional implementation, the underwater target detection device based on sonar images further includes: A second acquisition module is used to acquire a verification set from the sonar image sample set, input the verification set into the underwater target detection model, and output a verification result corresponding to the sample image in the verification set; A guiding module is used to measure the difference between the verification result and the sample label according to a cross entropy loss function to guide the underwater target detection model to update parameters.
[0146] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 5As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store a sonar image sample set. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for underwater target detection based on sonar images is implemented.
[0147] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0148] In an exemplary embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0149] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0150] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0151] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0152] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).
[0153] The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but is not limited thereto.
[0154] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0155] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A method for underwater target detection based on sonar images, characterized in that: The underwater target detection method based on sonar images comprises: Acquire a sonar image sample set, the sonar image sample set comprising sample images and sample labels attached to each of the sample images, each of the sample labels comprising position information and category probability distribution, the position information refers to the position information of the underwater target in the sonar image, and the category probability distribution refers to the category probability distribution of the underwater target; Obtaining a training set from the sonar image sample set, and training an initial underwater target detection model through the training set to obtain an underwater target detection model, wherein the initial underwater target detection model is a YOLOV10 network model, and the YOLOV10 network model includes an inverted bottleneck fusion module, and the inverted bottleneck fusion module is integrated with an SE block to adaptively adjust the weight of each channel of the sample image in the feature extraction stage; Acquire a sonar image to be detected, input the sonar image to be detected into the underwater target detection model, and output the position information and the category probability distribution.
2. The underwater target detection method based on sonar images according to claim 1 is characterized in that: Before obtaining a sonar image sample set, the underwater target detection method of the sonar image includes: Acquire initial sonar image; Calculating the local mean of each pixel in the initial sonar image; Calculating the local variance of each pixel in the initial sonar image; Calculating a denoised value of each pixel according to the local mean and the local variance of each pixel; Repeat the step of calculating the denoising value of each pixel, traverse each pixel in the initial sonar image, and obtain the initial sonar image after denoising; Determining the cumulative probability of each gray level appearing in the denoised initial sonar image according to the grayscale cumulative distribution function; Determining a grayscale value after equalization according to the cumulative probability; The equalized grayscale value is applied to each pixel of the initial sonar image that has been subjected to denoising processing to generate the sonar image.
3. The underwater target detection method based on sonar images according to claim 1 is characterized in that: The initial underwater target detection model is trained by the training set to obtain the underwater target detection model, including: Inputting the training set into a feature extraction network to obtain a first feature output, wherein the feature extraction network includes at least one convolution module, at least one bottleneck fusion module, at least one feature splicing module, a spatial pyramid pooling module, and a position squeeze attention module; Inputting the first feature output into a feature fusion network to obtain a second feature output, wherein the feature fusion network includes at least one convolution module, at least one upsampling module, at least one feature concatenation module, at least one bottleneck fusion module, and at least one inverted bottleneck fusion module; The second feature output is input into the output network to obtain a prediction result, which includes the position information and the category probability distribution. The difference between the sample label and the prediction result is compared to adjust the model parameters to obtain the underwater target detection model. The output network includes multiple target detection head modules, each of which includes two target detection head sub-modules. The target detection head module is used to detect targets of different scales.
4. The underwater target detection method based on sonar images according to claim 3 is characterized in that: The convolution module includes a first convolution module, a second convolution module, a third convolution module, a fourth convolution module, a fifth convolution module, a sixth convolution module and a seventh convolution module, the bottleneck fusion module includes a first bottleneck fusion module, a second bottleneck fusion module, a third bottleneck fusion module, a fourth bottleneck fusion module and a fifth bottleneck fusion module, the inverted bottleneck fusion module includes a first inverted bottleneck fusion module and a second inverted bottleneck fusion module, the upsampling module includes a first upsampling module and a second upsampling module, and the feature splicing module includes a first feature splicing module, a second feature splicing module, a third feature splicing module and a fourth feature splicing module; The first feature output includes a fifth feature map, a seventh feature map, and an eleventh feature map, wherein the fifth feature map is obtained by inputting the third feature map and the fourth feature map into the second bottleneck fusion module for feature fusion, the seventh feature map is obtained by inputting the fifth feature map and the sixth feature map into the third bottleneck fusion module for feature fusion, and the eleventh feature map is obtained by inputting the tenth feature map into the position squeezing attention module for attention mechanism processing; The second feature output includes a sixteenth feature map, a nineteenth feature map and a twenty-second feature map. The sixteenth feature map is obtained by inputting the fifteenth feature map and the fifth feature map into the second feature splicing module for splicing, the nineteenth feature map is obtained by inputting the eighteenth feature map into the fifth bottleneck fusion module for feature fusion, and the twenty-second feature map is obtained by inputting the twenty-first feature map into the second inverted bottleneck fusion module for feature optimization.
5. The underwater target detection method based on sonar images according to claim 4 is characterized in that: The step of inputting the second feature output into the output network to obtain a prediction result includes: Inputting the sixteenth feature map into a first target head module to obtain the position information and the category probability distribution of two different groups of first targets, wherein the first target head module includes a first target head submodule and a second target head submodule; Inputting the nineteenth feature map into a second target head module to obtain the position information and the category probability distribution of two different groups of second targets, wherein the second target head module includes a third target head submodule and a fourth target head submodule; The twenty-second feature map is input into the third target head module to obtain the position information and the category probability distribution of two groups of different third targets, wherein the third target head module includes a fifth target head sub-module and a sixth target head sub-module.
6. The underwater target detection method based on sonar images according to claim 1, characterized in that: The position information is the target bounding box coordinates, and the target bounding box coordinates are expressed in the form of center coordinates, width, and height.
7. The underwater target detection method based on sonar images according to claim 1 is characterized in that: The underwater target detection method based on sonar images also includes: Obtaining a verification set from the sonar image sample set, inputting the verification set into the underwater target detection model, and outputting a verification result corresponding to the sample image in the verification set; The difference between the verification result and the sample label is measured according to a cross entropy loss function to guide the underwater target detection model to perform parameter update.
8. An underwater target detection device based on sonar images, characterized in that: The underwater target detection device based on sonar images comprises: A first acquisition module is used to acquire a sonar image sample set, wherein the sonar image sample set includes sample images and sample labels attached to each sample image, wherein each sample label includes position information and category probability distribution, wherein the position information refers to the position information of the underwater target in the sonar image, and the category probability distribution refers to the category probability distribution of the underwater target; A training module is used to obtain a training set from the sonar image sample set, and train an initial underwater target detection model through the training set to obtain an underwater target detection model, wherein the initial underwater target detection model is a YOLOV10 network model, and the YOLOV10 network model includes an inverted bottleneck fusion module, and the inverted bottleneck fusion module is integrated with an SE block to adaptively adjust the weight of each channel of the sample image in the feature extraction stage; The input module is used to obtain the sonar image to be detected, input the sonar image to be detected into the underwater target detection model, and output the position information and the category probability distribution.
9. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the underwater target detection method based on sonar images as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the underwater target detection method based on sonar images described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Underwater sonar image target detection method and system
CN114926728A
Underwater motion biological recognition and evaluation method based on multi-beam image sonar
CN118298289A
Target real-time detection and classification method based on sonar image, medium and system
CN118865003A
Defect detection model construction method, defect detection method, defect detection device and electronic equipment
CN118918092A
Method for detecting number of people in high-place operation hanging basket based on improved YOLOv9 model
CN118942033A
Cited By
Fish school target tracking and cruising speed calculation method based on sonar image
CN120182326A
Sonar target detection method, electronic equipment, storage medium and product
CN121305325A
Underwater bubble cluster detection method based on ai model and related device
CN122530785A