Underwater sonar image small target detection method
The underwater sonar image is preprocessed through sparse decomposition and ROI extraction operations, and the YOLOv11 neural network is pruned and optimized, which solves the background noise interference and model complexity of small object detection of underwater sonar image in the prior art, and achieves high-precision and lightweight small object detection effect.
Patent Information
- Application Number
- CN202510695212.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing underwater sonar image small object detection method has problems in background noise interference, large model parameters, and high computational complexity, which is difficult to meet the dual requirements of real-time and lightweight.
The underwater sonar image is decomposed foreground and background by sparse decomposition, combined with ROI extraction operations, and the initialized foreground image is obtained; then the YOLOv11 neural network is pruned and optimized, using only the 2-4 layers of the prediction layer and adding a deformable convolutional layer to obtain a lightweight YOLOv11 neural network.
It achieves higher accuracy of small object detection of underwater sonar images, overcomes background noise interference problems, and reduces model parameters and calculation complexity, meeting the requirements of real-time and lightweight.
Smart Images

Figure CN120219940A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target recognition, and particularly to a method for detecting small targets in underwater sonar images. Background Art
[0002] Currently, there are mainly key problems in the detection of small targets in underwater sonar images, such as background noise interference, large number of parameters in deep neural networks, and high computational complexity. Due to the complexity of the underwater environment, bubbles and suspended matter result in a low signal-to-noise ratio of the acquired underwater sonar images. In particular, the foreground (small target features) is easily submerged; low resolution and weak contrast also lead to blurred edges of small targets, making it difficult for current methods to effectively separate the background and the foreground. At the same time, existing convolutional neural networks need to stack deep neural networks to capture the details of small targets, resulting in an explosion of the number of parameters, exacerbating the risk of overfitting. Moreover, small target detection still requires retaining fine-grained features and multi-scale feature fusion, forcing the network to use more channels or a larger input size, increasing the memory and computational costs. The inference speed of its complex model is slow and it is difficult to meet the actual needs.
[0003] Sparse decomposition technology aims to separate the target signal and the background from complex observation data and achieve the separation of low-rank background and sparse foreground through a mathematical model. Currently, related research has been applied to multiple scenarios such as image processing and underwater acoustics, but there are still problems with limited suppression effect on the scattering background. The YOLO series of algorithms is one of the mainstream algorithms for target detection. However, the existing YOLO algorithms still have problems such as insufficient accuracy, large number of model parameters, and slow detection speed when dealing with underwater small targets and blurred targets, and it is difficult to meet the dual requirements of real-time and lightweight for underwater sonar image small target detection.
[0004] Therefore, designing a method for detecting small targets in underwater sonar images with high accuracy and lightweight is one of the key problems that need to be solved urgently at present. Summary of the Invention
[0005] In view of the above problems, the purpose of the present invention is to provide a method for detecting small targets in underwater sonar images. First, considering the background interference and its low-rank property, while imposing constraints on the sparsity and connectivity of the foreground, the underwater sonar image is sparsely decomposed into foreground, background, and noise; then the foreground (small target) information is extracted through ROI (Region of Interest); on the basis of the above preprocessing, lightweight processing such as pruning and optimizing the YOLOv11 network is performed to solve key problems such as complex model structure and computational resource requirements, so as to obtain a higher accuracy for detecting small targets in underwater sonar images.
[0006] To solve the above technical problems, the present invention provides the following technical solutions:
[0007] On the one hand, a method for detecting small targets in underwater sonar images is provided, including the following steps:
[0008] S1. Input the underwater sonar image as the initial input data.
[0009] S2. Decompose the underwater sonar image into foreground and background through sparse decomposition, impose low-rank constraints on the background, and at the same time impose sparsity constraints and connectivity constraints on the foreground; then perform an operation of extracting the region of interest (ROI) on the foreground to obtain an initialized foreground image, that is, an initialized small target image.
[0010] S3. Prune and optimize the YOLOv11 neural network, including: differentiating using the Focal loss function and the DIoU loss function, only using the 2nd - 4th layers of the prediction layer and adding a deformable convolutional layer to obtain a lightweight YOLOv11 neural network.
[0011] S4. Input the initialized foreground image into the lightweight YOLOv11 neural network to obtain the detection result of small targets in the underwater sonar image.
[0012] Optionally, the step S2 specifically includes:
[0013] Considering the low rank of the background, as well as the smoothness, sparsity, and connectivity of the foreground, design a total objective function L, which consists of a background fitting term L fidelity , a foreground sparsity term L sparse and a foreground connectivity term L connect These three objective function terms; the specific operations are as follows:
[0014] (1)
[0015] The background fitting term is constrained by the L2 norm to obtain:
[0016] (2)
[0017] According to the matrix operation rules, formula (2) can be rewritten as:
[0018] (3)
[0019] Among them, is the support vector corresponding to the image sequence vector d, s i is the i-th element of the vector s, then represents the probability possibility that the i-th pixel in the image belongs to the background; d i is the i-th element of d, U i represents the i-th row of the matrix U; among them Denotes the dot product operation, that is, the element-by-element multiplication of corresponding elements; m is the number of pixels in the image, and w is the coefficient vector;
[0020] The foreground sparse term is constrained by the L0 norm to obtain:
[0021] (4)
[0022] The foreground connectivity term is constrained by the L0 norm to obtain:
[0023] (5)
[0024] Among them, C is the cyclic difference matrix;
[0025] Combining formulas (2), (4), and (5), the total objective function is obtained and solved by the augmented Lagrangian method and the alternating direction multiplier method to obtain s; then the region of interest ROI extraction operation is performed on s to obtain the initialized foreground image, that is, the initialized small target image.
[0026] Optionally, the step S3 specifically includes:
[0027] First, use the Focal loss function and the DIoU loss function for differentiation:
[0028] (6)
[0029] (7)
[0030] (8)
[0031] Among them, α t is the class balance factor, γ is the focal factor, p t is the predicted probability of the object detection task, ρ represents the Euclidean distance between the centers of the two boxes, c is the diagonal length of the smallest closed region containing the two boxes, and IoU is the intersection over union; L DIoU represents the specific calculation method of the DIoU loss, and ρ(b, b gt ) represents the Euclidean distance between the centers of the predicted box b and the ground truth box b gt ;
[0032] Secondly, only use the 2nd - 4th layers of the prediction layer, that is, the P2 - P4 layers, and at the same time add the deformable convolutional layer DCN to the cross - stage partial network C3k2 with a convolutional kernel of 2 to obtain the lightweight YOLOv11 neural network.
[0033] Optionally, the lightweight YOLOv11 neural network includes: the backbone network Backbone, the neck network Neck, and the head network Head;
[0034] Among them, the CBS module in the backbone network Backbone is used for preliminary feature extraction. The CBS module includes: a convolutional layer Conv, a batch normalization layer BN, and an activation function SiLU; C3k2-DCN introduces a deformable convolutional layer, and the spatial pyramid pooling fusion module SPPF fuses global context information through multi-scale pooling; the cross-stage partial pyramid squeeze attention module C2PSA combines position-sensitive attention to strengthen the weight distribution of key spatial regions;
[0035] The C3k2-DCN in the neck network Neck improves the multi-scale feature alignment ability. The feature concatenation Contact and the upsampling Upsample fuse the shallow and deep layer features to construct the feature pyramid network FPN and the path aggregation network PAN structures;
[0036] In the head network Head, the three prediction layers P2, P3, and P4 correspond to feature maps of different scales respectively, realizing multi-task decoupling.
[0037] Optionally, the step S4 specifically includes:
[0038] Input the initialized foreground image obtained by preprocessing in step S2, that is, the initialized small target image, into the lightweight YOLOv11 neural network obtained in step S3 to obtain the final underwater sonar image small target detection result.
[0039] On the other hand, an electronic device is provided. The electronic device includes:
[0040] A processor;
[0041] A memory, on which computer-readable instructions are stored. When the computer-readable instructions are loaded and executed by the processor, the steps of the underwater sonar image small target detection method as described above are implemented.
[0042] On the other hand, a computer-readable storage medium is provided. Program codes are stored in the computer-readable storage medium, and the program codes can be called by the processor to execute the steps of the underwater sonar image small target detection method as described above.
[0043] The beneficial effects brought by the technical solution provided by the present invention at least include:
[0044] The present invention provides an underwater sonar image small target detection method based on sparse decomposition and lightweight YOLOv11. Sparse decomposition is used to remove background interference and extract preliminary foreground information (small targets), overcoming the problem of background noise interference during small target detection, which is more conducive to the subsequent accurate extraction of small targets; at the same time, the lightweight optimization of the YOLOv11 neural network is carried out to solve problems such as too large network model parameters and high computational complexity, so as to obtain a higher accuracy of underwater sonar image small target detection. Brief Description of the Drawings
[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0046] Figure 1 is a flowchart of the underwater sonar image small target detection method based on sparse decomposition and lightweight YOLOv11 provided by the embodiments of the present invention;
[0047] Figure 2 is a structural diagram of the lightweight YOLOv11 network model provided by the embodiments of the present invention. Detailed Embodiments
[0048] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions of the embodiments of the present invention with reference to the drawings of the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0049] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, the use of the word "example" is intended to present concepts in a specific manner. In addition, in the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, their intended meanings are the same. In the embodiments of the present invention, sometimes subscripts such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, their intended meanings are the same.
[0050] The embodiments of the present invention provide an underwater sonar image small target detection method, which can be implemented by an electronic device, and the electronic device can be a terminal or a server. Refer to Figure 1 As shown, the processing flow of this method can include the following steps:
[0051] S1. Input an underwater sonar image as the initial input data.
[0052] S2. Perform foreground and background decomposition on the underwater sonar image through sparse decomposition, impose low-rank constraint on the background, and at the same time impose sparsity constraint and connectivity constraint on the foreground; then perform Region of Interest (ROI) extraction operation on the foreground to obtain an initialized foreground image, that is, an initialized small target image.
[0053] First, perform sparse decomposition operation on the underwater sonar image input in step S1 to obtain foreground, background, and noise. Considering the low-rank property of the background, as well as the smoothness, sparsity, and connectivity of the foreground, design a total objective function L, which consists of a background fitting term L fidelity , a foreground sparsity term L sparse , and a foreground connectivity term L connect . The specific operations are as follows:
[0054] (1)
[0055] The background fitting term is constrained by the L2 norm to obtain:
[0056] (2)
[0057] According to the matrix operation rules, formula (2) can be rewritten as:
[0058] (3)
[0059] Among them, is the support vector corresponding to the image sequence vector d, s i is the i-th element of the vector s, then represents the probability possibility that the i-th pixel in the image belongs to the background, that is, when s i gets closer to 1, the more likely this pixel belongs to the background, and vice versa. d i is the i-th element of d, U i represents the i-th row of the matrix U. m is the number of pixels in the image, and w is the coefficient vector. Among them represents the dot product operation, that is, the element-by-element multiplication of the corresponding elements.
[0060] The foreground sparsity term is constrained by the L0 norm to obtain:
[0061] (4)
[0062] The foreground connectivity term is constrained by the L0 norm to obtain:
[0063] (5)
[0064] Among them, C is the cyclic difference matrix.
[0065] Therefore, by integrating formulas (2), (4), and (5), the overall objective function is obtained. This overall objective function is efficiently solved using the augmented Lagrangian method and the alternating direction method of multipliers to obtain s. Then, a region of interest (ROI) extraction operation is performed on s to obtain the initialized foreground image, that is, the initialized small target image.
[0066] S3. Prune and optimize the YOLOv11 neural network, including: taking the difference between the Focal loss function and the DIoU (Distance-IoU) loss function, only using the 2nd - 4th layers of the prediction layer and adding a deformable convolutional layer to obtain a lightweight YOLOv11 neural network.
[0067] Reference Figure 2 As shown, in this step, LAMP (Layer-Adaptive Magnitude-based Pruning) model pruning and lightweight processing are performed on the YOLOv11 neural network.
[0068] First, take the difference between the Focal loss function and the DIoU loss function:
[0069] (6)
[0070] (7)
[0071] (8)
[0072] Among them, α t is the class balance factor, γ is the focus factor, p t is the predicted probability of the object detection task, ρ represents the Euclidean distance between the centers of the two boxes, c is the diagonal length of the smallest closed region containing the two boxes, and IoU is the intersection over union; L DIoU represents the specific calculation method of the DIoU loss, and ρ(b, b gt ) represents the Euclidean distance between the centers of the predicted box b and the ground truth box b gt .
[0073] Secondly, only use the 2nd - 4th layers of the prediction layer, that is, the P2 - P4 layers, and at the same time add a deformable convolutional layer DCN (Deformable Convolution) to the cross-stage partial network C3k2 (Cross Stage Partial with kernel size 2) with a convolutional kernel of 2 to obtain a lightweight YOLOv11 neural network.
[0074] The lightweight YOLOv11 neural network includes: a backbone network Backbone, a neck network Neck, and a head network Head.
[0075] Among them, the CBS module in the backbone network Backbone is used for preliminary feature extraction. The CBS module includes: a convolutional layer Conv, a batch normalization layer BN, and an activation function SiLU. C3k2-DCN (C3k2 with Deformable Convolution Network) introduces a deformable convolutional layer, and the spatial pyramid pooling fusion module SPPF (Spatial Pyramid Pooling-Fast) fuses global context information through multi-scale pooling; the cross-stage partial pyramid squeeze attention module C2PSA (Cross Stage Partial with Pyramid Squeeze Attention) combines position-sensitive attention to strengthen the weight distribution of spatial key regions.
[0076] The C3k2-DCN in the neck network Neck enhances the multi-scale feature alignment ability, and the feature concatenation Contact and upsampling Upsample fuse shallow and deep layer features to construct the feature pyramid network FPN (Feature Pyramid Networks) and the path aggregation network PAN (Path Aggregation Network) structures.
[0077] The three prediction layers P2, P3, P4 (Prediction Layer 2, Prediction Layer 3, Prediction Layer 4) in the head network Head correspond to feature maps of different scales respectively, realizing multi-task decoupling.
[0078] S4. Input the initialized foreground image into the lightweight YOLOv11 neural network to obtain the detection result of small targets in the underwater sonar image.
[0079] Specifically, input the initialized foreground image obtained by preprocessing in step S2, that is, the initialized small target image, into the lightweight YOLOv11 neural network obtained in step S3 to obtain the final detection result of small targets in the underwater sonar image.
[0080] In an embodiment of the present invention, a method for detecting small targets in underwater sonar images based on sparse decomposition and lightweight YOLOv11 is proposed. This method is divided into two stages: the first stage is sparse decomposition. Considering the low rankness of the background, as well as the smoothness, sparsity, and connectivity of the foreground, a foreground segmentation algorithm based on sparsity and connectivity is adopted to decompose the underwater sonar image into the foreground (small targets), background, and noise, and then ROI (region of interest) extraction processing is performed to extract the foreground information after removing the background. The second stage is the lightweight YOLOv11 neural network, which detects the foreground image preprocessed in the first stage, that is, pruning and optimizing YOLOv11 first. First, the Focal loss function and the DIoU loss function are used for differentiation to improve the accuracy of small target detection in the YOLOv11 network; secondly, only the 2nd - 4th layers of the prediction layer are used, and a deformable convolutional layer is added to obtain a lightweight YOLOv11 neural network, and finally, the detection of small targets in underwater sonar images is realized.
[0081] The method of the present invention removes background interference by adopting sparse decomposition, extracts preliminary foreground information (small targets), overcomes the problem of background noise interference during small target detection, and is more conducive to the accurate extraction of subsequent small targets. At the same time, lightweight optimization of the YOLOv11 neural network solves problems such as too large number of network model parameters and high computational complexity, thereby obtaining a higher accuracy rate for detecting small targets in underwater sonar images.
[0082] In an exemplary embodiment, the present invention also provides an electronic device, which includes:
[0083] A processor;
[0084] A memory, on which computer-readable instructions are stored. When the computer-readable instructions are loaded and executed by the processor, the steps of the method for detecting small targets in underwater sonar images as described above are implemented.
[0085] In an exemplary embodiment, the present invention also provides a computer-readable storage medium, in which at least one instruction is stored. The at least one instruction is loaded and executed by the processor to implement the steps of the method for detecting small targets in underwater sonar images as described above. For example, the computer-readable storage medium can be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0086] It should be noted that in this text, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or terminal device comprising a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or terminal device comprising said element.
[0087] References in the specification to "one embodiment", "an embodiment", "exemplary embodiments", "some embodiments", etc. indicate that the described embodiments may include a particular feature, structure, or characteristic, but not necessarily every embodiment includes that particular feature, structure, or characteristic. Additionally, when combining embodiments to describe a particular feature, structure, or characteristic, implementing such feature, structure, or characteristic in combination with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the relevant art.
[0088] It should be understood that the term "and / or" in this text is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B may be singular or plural. Additionally, the character " / " in this text generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context before and after.
[0089] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single item(s) or plural item(s). For example, at least one of a, b, or c may represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c may be single or multiple.
[0090] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not impose any limitation on the implementation process of the embodiments of the present invention.
[0091] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0092] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0093] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0094] If the described functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or this part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, and other media that can store program codes.
[0095] The present invention covers any substitutions, modifications, equivalent methods, and solutions made within the spirit and scope of the present invention. For the public to have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention. However, those skilled in the art can fully understand the present invention without these detailed descriptions. In addition, well-known methods, processes, procedures, components, and circuits are not described in detail to avoid unnecessary confusion to the essence of the present invention.
[0096] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. An underwater sonar image small target detection method, characterized in that, It includes the following steps: S1. Input an underwater sonar image as the initial input data; S2. Perform foreground and background decomposition on the underwater sonar image through sparse decomposition, impose a low-rank constraint on the background, and at the same time impose a sparsity constraint and a connectivity constraint on the foreground; then perform an operation of extracting the region of interest (ROI) on the foreground to obtain an initialized foreground image, that is, an initialized small target image; S3. Prune and optimize the YOLOv11 neural network, including: taking the difference between the Focal loss function and the DIoU loss function, only using the second to fourth layers of the prediction layer and adding a deformable convolutional layer to obtain a lightweight YOLOv11 neural network; S4. Input the initialized foreground image into the lightweight YOLOv11 neural network to obtain the detection result of small targets in the underwater sonar image.
2. The underwater sonar image small target detection method according to claim 1, wherein The specific steps of step S2 include: Considering the low rank of the background, as well as the smoothness, sparsity, and connectivity of the foreground, design a total objective function \(L\), which consists of a background fitting term \(L\) fidelity , a foreground sparsity term \(L\) sparse , and a foreground connectivity term \(L\) connect , which are composed of three objective function terms; the specific operations are as follows: (1) The background fitting term is constrained by the L2 norm to obtain: (2) According to the matrix operation rules, formula (2) can be rewritten as: (3) Among them, is the support vector corresponding to the image sequence vector d, and s i is the i-th element of the vector s, then represents the probability that the i-th pixel in the image belongs to the background; d i is the i-th element of d, and U i represents the i-th row of the matrix U; among them represents the dot product operation, that is, the element-by-element multiplication; m is the number of pixels in the image, and w is the coefficient vector; The foreground sparsity term is constrained by the L0 norm to obtain: (4) The foreground connectivity term is constrained by the L0 norm to obtain: (5) Among them, C is a cyclic difference matrix; Combining formulas (2), (4), and (5), the total objective function is obtained, and it is solved by the augmented Lagrangian method and the alternating direction multiplier method to obtain s; then an operation of extracting the region of interest (ROI) is performed on s to obtain an initialized foreground image, that is, an initialized small target image.
3. The method for detecting small targets in underwater sonar images according to claim 1, wherein, The specific steps of step S3 include: First, take the difference between the Focal loss function and the DIoU loss function: (6) (7) (8) Among them, α t is the class balance factor, γ is the focal factor, and p t is the predicted probability of the object detection task, representing the Euclidean distance between the center points of two boxes, c is the diagonal length of the smallest closed region containing the two boxes, and IoU is the intersection over union; L DIoU represents the specific calculation method of the DIoU loss, and ρ(b, b gt ) represents the Euclidean distance between the center points of the predicted box b and the ground truth box b gt ; Secondly, only use the second to fourth layers of the prediction layer, that is, the P2 - P4 layers, and at the same time add a deformable convolutional layer (DCN) to the cross-stage partial network C3k2 with a convolutional kernel of 2 to obtain a lightweight YOLOv11 neural network.
4. The underwater sonar image small target detection method according to claim 3, wherein, The lightweight YOLOv11 neural network includes: a backbone network Backbone, a neck network Neck, and a head network Head; Among them, the CBS module in the backbone network Backbone is used for preliminary feature extraction, and the CBS module includes: a convolutional layer Conv, a batch normalization layer BN, and an activation function SiLU; C3k2 - DCN introduces a deformable convolutional layer, and the spatial pyramid pooling fusion module SPPF fuses global context information through multi-scale pooling; the cross-stage partial pyramid squeeze attention module C2PSA combines position-sensitive attention to strengthen the weight distribution of key spatial regions; The C3k2 - DCN in the neck network Neck improves the multi-scale feature alignment ability, and the feature concatenation Contact and upsampling Upsample fuse the shallow and deep layer features to construct a feature pyramid network (FPN) and a path aggregation network (PAN) structure; In the head network Head, the three prediction layers P2, P3, and P4 correspond to feature maps of different scales respectively to achieve multi-task decoupling.
5. The underwater sonar image small target detection method according to claim 1, characterized in that, The specific steps of step S4 include: Input the initialized foreground image, that is, the initialized small target image, obtained by preprocessing in step S2 into the lightweight YOLOv11 neural network obtained in step S3 to obtain the final detection result of small targets in the underwater sonar image.
6. An electronic device, characterized in that, The electronic device includes: A processor; A memory storing computer-readable instructions, which, when loaded and executed by the processor, implement the method according to any one of claims 1 to 5.
7. A computer-readable storage medium, characterized in that, Program code is stored in the computer-readable storage medium, and the program code can be called by the processor to execute the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Sea-surface infrared small object detection method
CN105931264A
Infrared weak and small target detection method based on deep sparse low-rank neural network
CN115510660A
Traffic target detection method based on improved YOLOv7
CN117315614A
Deep learning-based crop disease and insect pest real-time detection method and device, and computer storage medium
CN117351353A
Medical image small target anomaly detection method based on multi-step iterative optimization
CN118396943A