Ship detection method based on edge semantic fusion and contrast feature aggregation
By adopting edge semantic fusion and contrast feature aggregation methods in SAR ship detection, the problem of difficult to balance detection accuracy, computing efficiency and generalization capabilities in the prior art is solved, and higher detection accuracy and more robust generalization capabilities are achieved.
Patent Information
- Application Number
- CN202510662317.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-22
AI Technical Summary
The existing SAR ship detection methods are difficult to balance the detection accuracy, computing efficiency and generalization capabilities, especially in complex maritime scenarios, where the resolution of the target and the background is insufficient, resulting in high false alarm rates and high missed detection rates.
The ship detection method based on edge semantic fusion and contrast feature aggregation is adopted. Through edge semantic fusion backbone network, contrast-driven feature aggregation module and dynamic fusion head, the resolution of the target and background is improved, and the cross-model generalization ability of the model is enhanced.
It effectively solves the semantic fuzzy problem caused by poor input of SAR texture, enhances the resilience of the target and background, significantly improves the detection accuracy and computing efficiency, and shows robust generalization ability in various maritime scenarios.
Smart Images

Figure CN120198682A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of ship detection, and provides a ship detection method based on edge semantic fusion and contrast feature aggregation. Background Art
[0002] Synthetic Aperture Radar (SAR) is a high-resolution imaging technology that can provide detailed radar images under various weather conditions, without being affected by cloud cover or lighting. However, due to factors such as low target-background contrast, speckle noise, and complex background clutter, ship detection in SAR images remains quite challenging. For example, in coastal areas, clutter generated by sea waves and land reflections complicates target recognition, and at the same time, false alarms are likely to occur when wave crests or buoys are misclassified as ships. In crowded waterways, overlapping radar echo signals of multiple ships impede precise separation. In addition, it is difficult to detect small or distant ships in the open sea because weak radar echoes are easily masked by sea clutter, often resulting in missed detections. Existing SAR ship detection methods are mainly divided into two paradigms: traditional Constant False Alarm Rate (CFAR)-based methods and deep learning-driven detectors. However, it is difficult for both of these paradigms to achieve a balance among detection accuracy, computational efficiency, and generalization ability in various maritime scenarios.
[0003] Traditional CFAR detectors, such as methods based on the generalized gamma distribution and bilateral CFAR, rely on statistical models to distinguish targets from background clutter. Although effective in controlled environments, due to their reliance on handcrafted features and strict assumptions about clutter distribution, their adaptability to real-world scenarios is limited. In addition, they cannot cope with the inherent challenges of SAR images, such as discontinuous target boundaries and spatially correlated speckle noise, which are more severe in dynamic sea conditions.
[0004] The emergence of deep learning has changed SAR ship detection. Frameworks such as YOLO, Faster R-CNN (Faster Region-based Convolutional Neural Network), and anchor-free detectors have achieved success in the visible light field. Despite these advancements, traditional deep learning architectures still have key flaws when applied to SAR images. Specifically, SAR images are usually single-channel and lack texture information, requiring precise edge localization to accurately depict targets, and this ability has not received sufficient attention in networks optimized for multi-spectral inputs. In addition, popular attention mechanisms (such as CBAM, RAM) tend to overemphasize local details, inadvertently amplifying background noise and exacerbating semantic ambiguity in complex scenarios.
[0005] Edge detection, as the cornerstone of computer vision, plays a crucial role in SAR ship detection by separating the target contour from the cluttered background. Traditional edge operators (such as Sobel, Canny) and deep learning-based methods (such as HED, DexiNed) have been widely applied. However, these methods are mainly designed for natural images with obvious foreground and background differentiation and are not applicable to the unique features of SAR images. SAR images usually resemble a collection of background objects with subtle intensity gradients, and traditional edge detectors are difficult to distinguish the true ship boundaries from the artifacts generated by noise. Summary of the Invention
[0006] The present invention aims to solve at least one of the technical problems existing in the related art. For this purpose, the present invention provides a ship detection method based on edge semantic fusion and contrast feature aggregation, which solves the semantic ambiguity problem caused by the texture-poor input of SAR, enhances the distinguishability between the target and the background, and highlights the robust cross-model generalization ability of the present invention.
[0007] The present invention provides a ship detection method based on edge semantic fusion and contrast feature aggregation, including the following steps: S1: Obtain and preprocess the image to obtain the preprocessed image; S2: Establish a ship detection model based on the base model, where the ship detection model includes an edge semantic fusion backbone network, a contrast-driven feature aggregation module, and a dynamic fusion head; S3: Set the training parameters, and use the preprocessed image to train the ship detection model to obtain a ship detection training model; S4: Input the detection image to be detected into the ship detection training model to obtain the detection result.
[0008] According to a ship detection method based on edge semantic fusion and contrast feature aggregation provided by the present invention, the base model includes the YOLO model and the RT-DETR model (Real-Time Detection Transformer).
[0009] According to a ship detection method based on edge semantic fusion and contrast feature aggregation provided by the present invention, the edge semantic fusion backbone network includes a basic feature extractor and a multi-scale edge feature generator.
[0010] According to a ship detection method based on edge semantic fusion and contrast feature aggregation provided by the present invention, the working process of the multi-scale edge feature generator is as follows: S11: Convolve the input image data to obtain an input feature map; S12: Perform edge information fusion on the input feature map to obtain the first edge information feature; S13: Perform edge information fusion on the first edge information feature to obtain the second edge information feature; S14: Perform edge information fusion on the second edge information feature to obtain the third edge information feature.
[0011] According to a ship detection method based on edge semantic fusion and contrast feature aggregation provided by the present invention, the process of the edge information fusion is as follows: S101: Define the Sobel horizontal direction operator and the Sobel vertical direction operator : is the transpose of the matrix; S102: Perform horizontal Sobel convolution on the input data according to the Sobel horizontal direction operator to obtain the horizontal Sobel result; S103: Perform vertical Sobel convolution on the input feature map according to the Sobel vertical direction operator to obtain the vertical Sobel result; S104: Multiply the horizontal Sobel result and the vertical Sobel result bit by bit to obtain the edge information feature.
[0012] According to a ship detection method based on edge semantic fusion and contrast feature aggregation provided by the present invention, the steps of the basic feature extractor are as follows: S111: Input the input feature map into the C3k2 module for feature extraction to obtain the first C3k2 result, and perform edge-guided feature fusion on the first C3k2 result and the first edge information feature to obtain the first feature map; S112: Input the first feature map into the C3k2 module for feature extraction to obtain the second C3k2 result, and perform edge-guided feature fusion on the second C3k2 result and the second edge information feature to obtain the second feature map; S112: Input the second feature map into the C3k2 module for feature extraction to obtain the third C3k2 result, and perform edge-guided feature fusion on the third C3k2 result and the third edge information feature to obtain the third feature map.
[0013] According to a ship detection method based on edge semantic fusion and contrast feature aggregation provided by the present invention, the process of the edge-guided feature fusion is as follows: S121: Perform pooling on the input feature map to obtain the first pooling result; S122: Pool the first pooling result and then perform convolution to obtain the second pooling result; S123: Pool the second pooling result and then perform convolution to obtain the third pooling result; S124: Bitwise multiply the edge feature, the first pooling result, the second pooling result, the third pooling result, and the C3k2 result to obtain an edge-guided intermediate value; S125: Perform convolution on the edge-guided intermediate value to obtain a feature map.
[0014] According to a ship detection method based on edge semantic fusion and contrast feature aggregation provided by the present invention, the process of the contrast-driven feature aggregation module is as follows: S21: Perform Haar wavelet transform on the third feature map to obtain a decomposed feature map; S22: Input the decomposed feature map into the CDFA image processing module to obtain a weighted feature map.
[0015] According to a ship detection method based on edge semantic fusion and contrast feature aggregation provided by the present invention, the dynamic fusion head integrates a bidirectional feature pyramid for multi-scale feature fusion, and the bidirectional feature pyramid includes a feature pyramid network and a path aggregation network.
[0016] According to a ship detection method based on edge semantic fusion and contrast feature aggregation provided by the present invention, the dynamic fusion head uses an oriented bounding box detector to replace the axis-aligned bounding box.
[0017] One or more of the above technical solutions in the embodiments of the present invention have at least one of the following technical effects: The ship detection method based on edge semantic fusion and contrast feature aggregation provided by the present invention, through a new framework that fuses the edge semantic fusion and contrast-driven feature fusion modules, solves the semantic ambiguity problem caused by the SAR texture-poor input, enhances the distinguishability between the target and the background, and highlights the robust cross-model generalization ability of the present invention.
[0018] The additional aspects and advantages of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. Brief Description of the Drawings
[0019] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0020] Figure 1 It is a schematic flowchart of the ship detection method based on edge semantic fusion and contrast feature aggregation provided by the present invention.
[0021] Figure 2 It is the ablation experiment result of the ship detection method based on edge semantic fusion and contrast feature aggregation provided by the present invention under different datasets.
[0022] Figure 3 It is the visualization experiment comparison chart of the ship detection method based on edge semantic fusion and contrast feature aggregation provided by the present invention. Detailed implementation manners
[0023] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention. The following embodiments are used to illustrate the present invention, but cannot be used to limit the scope of the present invention.
[0024] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the embodiments of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0025] The following will describe the present invention in conjunction with Figures 1 to 3 Describe the present invention.
[0026] Embodiment As Figure 1 shown, the embodiment of the present invention provides a ship detection method based on edge semantic fusion and contrast feature aggregation, including the following steps: S1: Obtain and preprocess the image to get the preprocessed image; S2: Establish a ship detection model based on the base model. The ship detection model includes an edge semantic fusion backbone network, a contrast-driven feature aggregation module, and a dynamic fusion head; S3: Set the training parameters, and use the preprocessed image to train the ship detection model to obtain a ship detection training model; S4: Input the detection image to be detected into the ship detection training model to obtain the detection result.
[0027] Specifically, the base model includes the YOLO model and the RT-DETR model.
[0028] Specifically, the edge semantic fusion backbone network consists of a base feature extractor and a multi-scale edge feature generator to form a dual-path structure. The base feature extractor is constructed with the improved C3k2 module as an example. This path goes through four downsampling stages to generate hierarchical semantic features. Each module uses depthwise separable convolutions to balance computational efficiency and feature richness. If in other models such as RT-DETR, the feature extractor can be replaced. The multi-scale edge generator uses cascaded Sobel kernels and adaptive pooling operations to extract noise-resistant edge features from the shallow convolutional layers. It outputs edge feature maps of three scales, and then fuses these feature maps with the semantic features of the backbone network.
[0029] The Multi-Scale Edge Generator (MSEG) aims to enhance the network's ability to capture edge information. This module first uses Sobel convolution (SobelConv) to extract edge information from the input feature map, and then performs progressive downsampling through max pooling to generate multi-scale edge feature maps. Then these feature maps will be adjusted through 1×1 convolution to ensure compatibility with the backbone network. This multi-scale edge information makes up for the deficiency of traditional convolutional neural networks in extracting fine edge details.
[0030] Specifically, the steps of the multi-scale edge feature generator are as follows: S11: Convolve the input image data to obtain the input feature map; S12: Perform edge information fusion on the input feature map to obtain the first edge information feature; S13: Perform edge information fusion on the first edge information feature to obtain the second edge information feature; S14: Perform edge information fusion on the second edge information feature to obtain the third edge information feature.
[0031] Specifically, the process of the above edge information fusion is as follows: S101: Define the Sobel horizontal direction operator and the Sobel vertical direction operator : is the input image, is the convolution operation, is the transpose of the matrix; Similarly, the gradient magnitude and the gradient angle Strong edges characterized by high gradient magnitudes usually appear in regions with significant color or texture changes, while weak edges with lower magnitudes indicate smoother regions. The output of the Sobel convolution combines information from both directions to highlight the edge response in the image.
[0032] S102: Perform horizontal Sobel convolution on the input data according to the Sobel horizontal direction operator to obtain the horizontal Sobel result; S103: Perform vertical Sobel convolution on the input feature map according to the Sobel vertical direction operator to obtain the vertical Sobel result; The Sobel kernels in the vertical and horizontal directions are defined as [3,3] matrices. Then these kernels are extended to three dimensions to meet the requirements of the PyTorch convolution layer. It should be noted that the Sobel kernels are fixed and do not participate in backpropagation updates. The weights of the convolution layer are respectively assigned to the Sobel_X and Sobel_Y kernels for horizontal and vertical edge detection. The Sobel_X kernel highlights the brightness changes in the horizontal direction through the positive and negative alternating coefficients in the middle row, enhancing the horizontal edge response. The Sobel_Y kernel detects the brightness gradient in the vertical direction through the positive and negative alternating coefficients in the middle column, strengthening the vertical edge features. Using depthwise separable convolution allows each input channel to independently apply the Sobel kernel without sharing convolution weights, reducing the computational amount.
[0033] S104: Multiply the horizontal Sobel result and the vertical Sobel result bitwise to obtain the edge information feature.
[0034] Specifically, the working process of the basic feature extractor is as follows: S111: Input the input feature map into the C3k2 module for feature extraction to obtain the first C3k2 result, and perform edge-guided feature fusion on the first C3k2 result and the first edge information feature to obtain the first feature map; S112: Input the first feature map into the C3k2 module for feature extraction to obtain the second C3k2 result, and perform edge-guided feature fusion on the second C3k2 result and the second edge information feature to obtain the second feature map; S112: Input the second feature map into the C3k2 module for feature extraction to obtain the third C3k2 result, and perform edge-guided feature fusion on the third C3k2 result and the third edge information feature to obtain the third feature map.
[0035] The feature map after Sobel edge extraction will undergo consecutive max pooling operations. The size of the feature map is halved after each pooling, generating edge features of different scales, and the edge information is retained at different resolutions through the multi-scale feature maps.
[0036] Specifically, the process of the edge-guided feature fusion is as follows: S121: Pool the input features to obtain the first pooling result; S122: Pool the first pooling result and then perform convolution to obtain the second pooling result; S123: Pool the second pooling result and then perform convolution to obtain the third pooling result; S124: Multiply the edge features, the first pooling result, the second pooling result, the third pooling result, and the C3k2 result bitwise to obtain the edge-guided intermediate value; S125: Perform convolution on the edge-guided intermediate value to obtain the feature map.
[0037] For the selection of downsampling, it needs to be more cautious. The goal of the present invention is to retain and enhance the edge information while performing downsampling. MaxPool (max pooling layer) can retain the strongest features in the local area and better reflect the edge information. AvgPool (average pooling layer) is more suitable for scenarios where features need to be smoothed or homogenized, but its performance in retaining details and edge information is not as good as MaxPool. Different feature maps adjust the dimensions of the feature maps of different scales to be consistent through 1x1 convolution, and perform cross-channel fusion of edge information and ordinary convolution features in a cross-channel manner. Subsequently, the fused features are further extracted through 3x3 convolution to enhance the model's ability to capture local details, and the output feature dimension is adjusted through 1x1 convolution.
[0038] Specifically, the process of the contrast-driven feature aggregation module is as follows: S21: Perform Haar wavelet transform on the third feature map to obtain the decomposed feature map.
[0039] A single edge detection method often performs poorly in complex scenarios. By combining the Haar wavelet transform and the Contrast-driven Feature Aggregation (CDFA) image processing module, the foreground and background information are compared through an attention mechanism. The attention mechanism can dynamically adjust the weights of features and enhance the expression ability of important regions.
[0040] First, four filters are defined, corresponding to the low-frequency (LL), horizontal (LH), vertical (HL), and diagonal (HH) directions respectively. Among them, LL retains low-frequency information through a double low-pass filter. LH / HL respectively focus on single-direction high-frequency signals, and the filter design reflects the directional differences. HH retains high-frequency signals in both directions and responds to sharp cross features. The filter weights are initialized by hard coding and are not involved in training by default. Convolve the input feature map to obtain feature maps in four directions, and then decompose them into low-frequency and high-frequency features. Among them, the high-frequency features capture detailed information, apply the attention weights to the values to obtain the weighted foreground features. The low-frequency features capture the overall structural information and apply the attention weights to the weighted foreground features to further enhance the features.
[0041] S22: Input the decomposed feature map into the CDFA image processing module to obtain a weighted feature map.
[0042] The attention mechanism in CDFA runs implicitly through the convolutional layer, directly deriving query and key information from the decomposed feature map (the high-frequency signal is represented by fg and the low-frequency signal is represented by bg), rather than through an explicit linear layer. First, high-frequency and low-frequency feature maps are randomly generated, and then they are concatenated to obtain a fused feature map. Then, two convolutional layers, the query convolution and the key convolution, are defined to generate queries and keys respectively. This method avoids the computational overhead of the traditional attention mechanism while maintaining the ability to dynamically adjust the feature weights. The downsampled feature map is projected into the attention weight space through a linear layer, enabling the weights to be directly generated from the feature map. This design conforms to the multi-scale characteristics of images and reduces the computational complexity.
[0043] Perform local feature extraction on the feature map after Haar wavelet transform and attention-weighted aggregation, and then gradually fuse higher-level semantic information through stacked convolutions while keeping the spatial dimensions of the input and output consistent. The finally output feature map combines low-frequency and high-frequency information, and at the same time weights important regions through the attention mechanism, enabling better capture of the global structure and local details of the image. Combining the previous edge information, the CDFA module can improve the adaptability of the model to complex scenarios and make the detection results more accurate.
[0044] Specifically, the dynamic fusion head integrates a bidirectional feature pyramid for multi-scale feature fusion, and the bidirectional feature pyramid includes a Feature Pyramid Network and a Path Aggregation Network.
[0045] Specifically, the dynamic fusion head replaces the axis-aligned bounding box with an oriented bounding box detector. Dynamic fusion weights features according to their importance, significantly enhancing the separation effect between the target and the background. In addition, to address the problem of arbitrary orientations of ships in synthetic aperture radar (SAR) images, the oriented bounding box (OBB) detector is used to replace the traditional axis-aligned bounding box (AABB). The OBB detector can predict the rotation angle and coordinates, accurately locate rotating targets such as ships, and can fit the ship contour more closely and reduce background interference, thereby improving the detection accuracy.
[0046] To verify the performance of the proposed object detection network, we conducted ablation experiments on three datasets, namely SSDD (SAR Ship Detection Dataset), HRSID (High Resolution SAR Images Dataset), and RSDD-SAR (Rotated Ship Detection Dataset in SAR Images), using three benchmark models, RT-DETR, YOLOv11, and YOLOv8, respectively, and compared them with existing SAR image ship detection networks.
[0047] In the embodiment of the present invention, the dataset is divided into 928 training images and 232 test images according to a ratio of 8:2. This division ensures the balanced distribution of data and facilitates fair comparison with other methods. The dataset covers ship targets in nearshore and offshore areas, presenting diverse scenarios, which poses challenges to the robustness of the detection algorithm. This experiment is based on an NVIDIA GeForce RTX 3080Ti graphics card, adopting a cosine learning rate decay strategy, configuring the learning rate and momentum coefficient to 0.001 and 0.937 respectively, setting the initial learning rate to 0.001, and defining the minimum learning rate as 0.01 times the initial learning rate.
[0048] The performance of the experimental method is mainly evaluated by the mean average precision while using precision and recall as auxiliary evaluation indicators. Among them, precision represents the proportion of correctly detected positive samples among all samples detected as positive, and recall represents the proportion of correctly detected positive samples among all true positive samples. The calculation formulas are as follows: Among them, is precision, is recall, The number of correctly detected positive samples; Indicates the number of background regions misdetected as targets; Indicates the number of real targets not detected. The average precision is obtained by integrating the precision-recall curve, and its calculation formula is: where, Indicates the precision at recall The higher the value, the better the performance of the target detection algorithm. Since the dataset used in this study only contains a single category (ship-shaped targets), so the value is the value. is the mean average precision calculated at different Intersection over Union (IoU) thresholds, and its calculation formula is: where ranges from 0.50 to 0.95, increasing in steps of 0.05. is the step ordinal number. Compared with using a single IoU threshold as the evaluation criterion, calculate the values at 10 different IoU thresholds and then take the average. This evaluation method places higher requirements on the localization accuracy of the algorithm.
[0049] In the embodiments of the present invention, RT-DETR, YOLOv11, and YOLOv8 are used as the base models respectively to conduct systematic ablation experiments to evaluate the performance of the proposed multi-scale edge feature generator (MSEG) and contrast-driven feature aggregation module (CDFA), and verify their effectiveness for the target detection task. As Figure 2 shown, Figure 2 shows the detailed experimental results. In the experiment, is used as the main evaluation index to measure the detection accuracy of the model. In the figure, the blue bars are for adding only the CDFA module, MEF is for adding only the MSEG module, and COMP is for adding both modules. Figure 2 In (a), RT-DETR is used as the base model, Figure 2 In (b), YOLOv11 is used as the base model, Figure 2 In (c), YOLOv8 is used as the base model. Among them, RT-DETR does not adopt the detection head for OBB boxes. MSEG has a significant improvement effect on the multi-object scenario in the open sea. On the RT-DETR model, the mAP gain of the Inshore dataset is 4.10%, which is the largest gain among all models for a scenario, indicating that edge features are crucial for the low-contrast inshore ship scenario, and the generated multi-scale edge features can effectively enhance the distinguishability between small targets (inshore ships) and complex backgrounds (waves, docks). The CDFA module shows a certain performance improvement in different base algorithms and scenarios, demonstrating its broad application potential and adaptability. The experimental results show that both the MSEG and CDFA modules can significantly improve the model performance on different datasets, especially on the RT-DETR model. The MSEG module significantly improves the model's detection ability through the enhancement of multi-scale edge features, while the CDFA module is more suitable for scenarios with dense targets and more background interference. When both the MSEG and CDFA are added simultaneously, the performance improvement is more obvious, proving the synergistic effect of these two modules in multi-scale feature fusion and edge feature extraction.
[0050] Figure 3 The comparison of the detection effects of different module combinations in typical scenarios of the SSDD dataset is shown, and the improvement brought by the modules can be more intuitively felt from the image details. From left to right are the detection results of the base model (Base), adding only the MSEG module (MSEG), adding only the CDFA module (CDFA), and joint use (COMP). Each row corresponds to the inshore (Inshore) and offshore (Offshore) scenarios in the SSDD dataset, covering the detection challenges of small targets, large targets, and dense targets. Among them, the red circles indicate false detections, the yellow circles indicate missed detection targets, and the green circles indicate detection boxes with a large overlapping area. In the inshore and offshore small target scenarios of the third and fifth rows, many small targets are severely missed under the base model. After adding the MSEG module, the small target missed detection rate is reduced, and the blue detection points increase significantly. In the dense inshore target separation scenario of the second row, it is difficult for the base model to accurately distinguish the densely arranged targets. After adding the CDFA module, the detection boxes can fit more closely to the contours of the arranged targets, and the coincidence degree is significantly improved. In the inshore large target localization scenario of the first row, the detection boxes of the MSEG module will have the problem of capturing more true targets but increasing false alarms, and the CDFA module can effectively suppress false alarms. Generally speaking, MSEG significantly enhances edge sensitivity, especially suitable for small targets and low-contrast scenarios, but may introduce false alarms. CDFA optimizes the detection confidence through dynamic feature weighting, effectively suppressing false alarms and improving the boundary fitting degree. When the two modules are used comprehensively, the detection accuracy and integrity can be greatly improved while avoiding missed detections.
[0051] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.
[0052] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0053] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
[0054] It should be noted that the embodiments of the present disclosure can be implemented through hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic: the software part can be stored in a memory and executed by a suitable instruction execution system, such as a microprocessor or dedicated design hardware. Those skilled in the art can understand that the above devices and methods can be implemented using computer-executable instructions and / or included in the processor control code, for example, such code is provided on a programmable memory or a data carrier such as an optical or electronic signal carrier.
[0055] Moreover, although the operations of the methods of the present disclosure are depicted in the figures in a particular order, this is not required or implied to perform these operations in that particular order, or that all of the illustrated operations must be performed to achieve the desired result. On the contrary, the steps depicted in the flowchart may be changed in their order of execution. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step and executed, and / or one step may be decomposed into multiple steps and executed. It should also be noted that the features and functions of two or more devices according to the present disclosure may be embodied in one device. Conversely, the features and functions of one device described above may be further divided and embodied by multiple devices.
[0056] Although the present disclosure has been described with reference to several specific embodiments, it should be understood that the present disclosure is not limited to the specific embodiments disclosed. The present disclosure is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.
Claims
1. A ship detection method based on edge semantic fusion and contrast feature aggregation, characterized in that, It includes the following steps: S1: Obtain and preprocess the image to obtain the preprocessed image; S2: Establish a ship detection model based on the basic model. The ship detection model includes an edge semantic fusion backbone network, a contrast-driven feature aggregation module, and a dynamic fusion head; S3: Set the training parameters, and use the preprocessed image to train the ship detection model to obtain a ship detection training model; S4: Input the detection image to be detected into the ship detection training model to obtain the detection result.
2. The ship detection method based on edge semantic fusion and contrast feature aggregation according to claim 1, characterized in that The basic model includes the YOLO model and the RT-DETR model.
3. A ship detection method based on edge semantic fusion and contrast feature aggregation according to claim 1, characterized in that The edge semantic fusion backbone network includes a basic feature extractor and a multi-scale edge feature generator.
4. A ship detection method based on edge semantic fusion and contrast feature aggregation according to claim 3, characterized in that The working process of the multi-scale edge feature generator is as follows: S11: Convolve the input image data to obtain an input feature map; S12: Perform edge information fusion on the input feature map to obtain a first edge information feature; S13: Perform edge information fusion on the first edge information feature to obtain a second edge information feature; S14: Perform edge information fusion on the second edge information feature to obtain a third edge information feature.
5. A ship detection method based on edge semantic fusion and contrast feature aggregation according to claim 4, characterized in that The working process of the edge information fusion is as follows: S101: Define the Sobel horizontal direction operator and the Sobel vertical direction operator : is the transpose of the matrix; S102: Perform horizontal Sobel convolution on the input data according to the Sobel horizontal direction operator to obtain a horizontal Sobel result; S103: Perform vertical Sobel convolution on the input feature map according to the Sobel vertical direction operator to obtain a vertical Sobel result; S104: Multiply the horizontal Sobel result and the vertical Sobel result bit by bit to obtain an edge information feature.
6. The ship detection method based on edge semantic fusion and contrast feature aggregation according to claim 5, characterized in that, The steps of the basic feature extractor are as follows: S111: Input the input feature map into the C3k2 module for feature extraction to obtain a first C3k2 result, and perform edge-guided feature fusion on the first C3k2 result and the first edge information feature to obtain a first feature map; S112: Input the first feature map into the C3k2 module for feature extraction to obtain a second C3k2 result, and perform edge-guided feature fusion on the second C3k2 result and the second edge information feature to obtain a second feature map; S113: Input the second feature map into the C3k2 module for feature extraction to obtain a third C3k2 result, and perform edge-guided feature fusion on the third C3k2 result and the third edge information feature to obtain a third feature map.
7. The ship detection method based on edge semantic fusion and contrast feature aggregation according to claim 6, wherein The edge-guided feature fusion process is as follows: S121: Pool the input feature map to obtain a first pooling result; S122: Pool the first pooling result and then perform convolution to obtain a second pooling result; S123: Pool the second pooling result and then perform convolution to obtain the third pooling result; S124: Multiply the edge feature, the first pooling result, the second pooling result, the third pooling result, and the C3k2 result bit by bit to obtain an edge-guided intermediate value; S125: Convolve the edge-guided intermediate value to obtain a feature map.
8. The ship detection method based on edge semantic fusion and contrast feature aggregation according to claim 6, characterized in that The process of the contrast-driven feature aggregation module is as follows: S21: Perform Haar wavelet transform on the third feature map to obtain a decomposed feature map; S22: Input the decomposed feature map into the CDFA image processing module to obtain a weight feature map.
9. A ship detection method based on edge semantic fusion and contrast feature aggregation according to claim 1, characterized in that The dynamic fusion head integrates a bidirectional feature pyramid for multi-scale feature fusion, and the bidirectional feature pyramid includes a feature pyramid network and a path aggregation network.
10. The ship detection method based on edge semantic fusion and contrast feature aggregation according to claim 1, wherein, The dynamic fusion head replaces the axis-aligned bounding box with an oriented bounding box detection head.
Citation Information
Patent Citations
Pooling unit design method of convolutional neural network
CN108805285A
Automatic quantitative analysis method and system for lung digital pathological image
CN113222012A
Data processing system, method of operating same, and computing system
CN115860072A
Semiconductor silicon wafer detection method and device, computer equipment and medium
CN116433674A
High-resolution remote sensing image building extraction method
CN118279762A