Coastline target intelligent identification method based on improved SOLOv2
By improving the SOLOv2 model and combining deformable convolution and attention modules, the problems of pixel-level masking and small target recognition in coastal litter monitoring are solved, achieving high-precision litter quantity statistics and spatial visualization, which is suitable for real-time monitoring by UAVs and satellite platforms.
Patent Information
- Application Number
- CN202511144752.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies struggle to achieve accurate pixel-level mask output and individual object differentiation in coastal debris monitoring, especially exhibiting low recognition accuracy under complex lighting and dynamic backgrounds, and insufficient detection capability for small targets.
An improved SOLOv2 model is adopted, which embeds a deformable convolutional backbone network (ResNet50vd-DCN) and a SimAM attention module, and combines a feature pyramid network to perform residual feature enhancement and adaptive spatial fusion. Parallel branches are used to predict the target category and generate instance masks, and a heatmap of enriched regions is generated by kernel density estimation.
It improves the accuracy of coastline target identification and adaptability to complex scenarios, can accurately count the amount of garbage and provide spatial visualization, and is suitable for real-time coastline garbage monitoring using drones and satellite platforms.
Smart Images

Figure CN120976803A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image analysis technology, and in particular to an intelligent method for recognizing coastline targets based on an improved SOLOv2. Background Technology
[0002] In recent years, with the development of UAV remote sensing technology and deep learning, target identification in coastal areas has gradually shifted from traditional manual patrols to intelligent monitoring. Current mainstream methods mainly include deep learning models based on target detection and semantic segmentation. For example, Mask R-CNN achieves pixel-level target recognition by introducing an instance segmentation branch, while YOLACT improves detection efficiency while maintaining real-time performance. In addition, some improved convolutional neural network structures have been used to enhance feature extraction capabilities, such as introducing attention mechanisms and multi-scale fusion strategies. These methods have improved the recognition accuracy in complex scenarios to a certain extent, providing technical support for marine environmental protection.
[0003] However, existing technologies still have many shortcomings and cannot meet the actual needs of coastline litter monitoring. On the one hand, traditional object detection methods can only output bounding boxes and cannot provide accurate pixel-level masks, resulting in inaccurate statistics on litter coverage area. On the other hand, although semantic segmentation methods can achieve pixel-level classification, they lack the ability to distinguish individual objects, making it difficult to count the number of litter items, especially when dealing with scattered small targets (such as cigarette butts and plastic fragments), where the recognition accuracy is low. In addition, most models have high false detection rates and poor robustness under complex lighting and dynamic backgrounds (such as waves, beaches, and vegetation). Existing models (such as MaskR-CNN) rely on the accuracy of detection boxes, and inaccurate localization can lead to segmentation shifts; although YOLACT has high real-time performance, its segmentation performance is limited (the mask AP on the COCO dataset is only 34.6%), and its ability to detect small targets is insufficient.
[0004] Therefore, there is an urgent need for an intelligent target recognition method based on the improved SOLOv2, which can take into account both target detection and instance segmentation and improve recognition accuracy. Summary of the Invention
[0005] To address the aforementioned technical problems, this invention provides an intelligent target recognition method for coastlines based on an improved SOLOv2, which can enhance target detection accuracy, small target recognition capability, and adaptability to complex scenarios.
[0006] This invention provides a method for intelligent identification of coastline targets based on an improved SOLOv2, comprising the following steps: S1. Acquire images of the coastline area to be monitored using UAV imagery and preprocess them to generate training samples; S2. Input the training samples into a deformable convolutional backbone network to extract multi-scale feature maps; input the multi-scale feature maps into a feature pyramid network for residual feature enhancement and adaptive spatial fusion to output a refined feature map; S3. The refined feature map is processed through parallel branching to predict the target category and generate an instance mask, respectively. S4. Perform nonmaximum suppression and morphological processing on the instance mask, and count the number and location of each type of target; and generate an enriched area heat map based on kernel density estimation according to the location information, and bind the heat map with geographic coordinates and output it to the GIS system.
[0007] Furthermore, the preprocessing in S1 includes: The image of the coastline area to be monitored is normalized to obtain an image of a preset size; The image of the preset size is subjected to multi-scale blending enhancement to generate training samples; wherein, the multi-scale blending enhancement includes random scaling, illumination perturbation and noise simulation.
[0008] Furthermore, in S2, the backbone network is a ResNet50vd-DCN structure; A dilated convolution with a dilation rate of 2 is introduced in the Stage 4 layer of the backbone network. A SimAM attention module is embedded at the end of the residual module of the backbone network, and the SimAM attention module calculates the attention weights through an energy function; The backbone network parameters were compressed to 3.2M after channel pruning.
[0009] Furthermore, the implementation of the SimAM attention module includes the following sub-steps: Calculate the energy function E(x) of the feature map of the input training samples. i ); The attention weight w at each position is obtained through normalization. i ; The attention weight w i The feature map is multiplied element-wise with the feature map of the input training sample to obtain the attention-enhanced feature map.
[0010] Furthermore, the formula for the energy function is: ; Wherein, E(x) i ) represents the spatial and channel attention weights of the current location features; x i Let x be the feature value at the i-th position in the feature map. j Let be the feature value of the j-th position within the neighborhood of the i-th position; N is the total number of neighborhood positions.
[0011] Furthermore, in step S2, the multi-scale feature map is input into a feature pyramid network for residual feature enhancement and adaptive spatial fusion, outputting a refined feature map, specifically including: The high-level feature C5 in the multi-scale feature map is adjusted to 7×7 through adaptive pooling, and then enhanced by 1×1 convolution dimensionality reduction and upsampling. The enhanced feature is fused with the dimensionality-reduced feature M5, and then the feature P5 with residual enhancement is output by 3×3 convolution. Features P2, P3, P4, and P5 are concatenated and then normalized sequentially using 1×1 convolution for dimensionality reduction, 3×3 convolution, and the Sigmoid function to generate a spatial weight map. The spatial weight map is then weighted and aggregated with the original features P2, P3, P4, and P5 to output the refined feature map.
[0012] Furthermore, in S3, the parallel branch includes: The refined feature map is compressed by a 1×1 convolutional layer to output a target semantic category probability map and identify the number of target categories. The kernel offset is dynamically learned by a 3×3 deformable convolutional layer to perform coordinate regression and generate an instance mask.
[0013] Furthermore, in step S5, the enrichment region heatmap is generated through kernel density estimation, and the kernel density estimation formula is: Where y is the coordinate of any point on the heatmap, y a Let be the coordinates of the a-th target, m be the total number of targets, h be the bandwidth parameter, and K be the Gaussian kernel function.
[0014] Furthermore, the bandwidth parameter h is adaptively adjusted according to the target density: h=0.3 when the target density is ≥50 targets / km², and h=0.7 when the target density is <50 targets / km².
[0015] Furthermore, during the training phase, adversarial training strategies are employed, including: Adversarial examples are generated using the FGSM algorithm; The adversarial loss function is minimized as follows: ,in, For adversarial loss function; Here, x is the mathematical expectation operator; x is the original training sample, x adv For adversarial examples, It is a discriminator.
[0016] The present invention has the following technical effects: 1. By combining image preprocessing with multi-scale hybrid enhancement, including size normalization, random scaling, illumination perturbation, and noise simulation, the model's generalization ability is improved based on data augmentation theory. Multi-scale input enhances the model's adaptability to targets of different sizes, increases the diversity of training samples, enhances the model's robustness under complex lighting and backgrounds, and reduces overfitting.
[0017] 2. By embedding a deformable convolution backbone network (ResNet50vd-DCN) and introducing dilated convolution and SimAM attention modules, deformable convolution enables dynamic shifting of spatial sampling points, enhancing the perception of irregular targets. The SimAM module calculates attention weights based on energy functions, enhancing key feature regions and significantly improving the recognition accuracy of complex-shaped targets such as curved fishing nets and irregular garbage, thereby enhancing the model's ability to capture detailed features.
[0018] 3. By using residual enhancement and adaptive spatial fusion strategies in the feature pyramid network, the semantic expression capability is enhanced by multi-scale feature complementarity mechanism, the spatial weight map guides the feature fusion direction, improves the utilization rate of low-level feature information, enhances the recognition capability of small targets (such as cigarette butts and debris), and improves the overall detection accuracy.
[0019] 4. By predicting categories and generating instance masks through parallel branches, and using deformable convolution to dynamically learn the convolution kernel offset to adapt to target deformation, pixel-level instance segmentation is achieved, accurately distinguishing individual target objects and solving the problem that existing methods cannot count the number of targets.
[0020] 5. Based on kernel density estimation, heat maps of enriched areas are generated and bound to geographic coordinates for output to the GIS system. Kernel density estimation reflects the distribution density of the target, and combined with the GIS system, spatial visualization is achieved, intuitively presenting waste-rich areas and providing a scientific basis for marine environmental protection decisions. Attached Figure Description
[0021] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0022] Figure 1 This is a flowchart of an intelligent coastal target recognition method based on an improved SOLOv2 provided by an embodiment of the present invention. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0024] Figure 1 This is a flowchart illustrating an intelligent coastline target recognition method based on an improved SOLOv2, provided in an embodiment of the present invention. See also... Figure 1 This invention provides a method for intelligent identification of coastline targets based on an improved SOLOv2, comprising the following steps: S1. Acquire images of the coastline area to be monitored using UAV imagery and perform preprocessing to generate training samples.
[0025] In some embodiments, the preprocessing in S1 includes: The image of the coastline area to be monitored is normalized to obtain an image of a preset size; The image of the preset size is subjected to multi-scale blending enhancement to generate training samples; wherein, the multi-scale blending enhancement includes random scaling, illumination perturbation and noise simulation.
[0026] Specifically, drones equipped with high-resolution cameras are used to collect images of the target coastline area. The drones should fly along a predetermined path to ensure coverage of the entire monitored area. The acquired raw images may have different sizes and resolutions, therefore size normalization is required: determine a preset size, and select a suitable preset size (e.g., 512x512 pixels) based on the needs of subsequent processing steps and the limitations of computing resources; adjust the image size, using image processing software or libraries to adjust all raw images to the aforementioned preset size. Then, to increase the diversity of training samples and improve the model's generalization ability and robustness, this invention employs a multi-scale hybrid enhancement technique to further process the normalized images, specifically including: random scaling, randomly scaling the image size within a certain range (e.g., between 50% and 150% of the original image size) to simulate the effect of shooting at different distances; and illumination perturbation, simulating shooting scenarios under different lighting conditions by adjusting parameters such as image brightness, contrast, and saturation. For example, the brightness value can be randomly changed by ±30%, and the contrast by ±20%; different types of noise (such as Gaussian noise and salt-and-pepper noise) can be added to the image to enhance the model's resistance to various interference factors that may occur in the real world. The noise intensity can be set according to the actual situation.
[0027] After the above steps, a diverse training sample set is generated. Each sample not only contains information about the original target coastline area but also covers various factors that may affect the recognition effect (such as size variations, lighting differences, noise interference, etc.). Such a training sample set helps the model learn target features better, thus exhibiting higher accuracy and stability in practical applications.
[0028] S2. Input the training samples into the deformable convolutional backbone network to extract multi-scale feature maps; input the multi-scale feature maps into the feature pyramid network for residual feature enhancement and adaptive spatial fusion to output a refined feature map.
[0029] In some embodiments, in step S2, the backbone network is a ResNet50vd-DCN structure; A dilated convolution with a dilation rate of 2 is introduced in the Stage 4 layer of the backbone network. A SimAM attention module is embedded at the end of the residual module of the backbone network, and the SimAM attention module calculates the attention weights through an energy function; The backbone network parameters were compressed to 3.2M after channel pruning.
[0030] In some embodiments, the implementation of the SimAM attention module includes the following sub-steps: Calculate the energy function E(x) of the feature map of the input training samples. i ); The attention weight w at each position is obtained through normalization. i ; The attention weight w i The feature map is multiplied element-wise with the feature map of the input training sample to obtain the attention-enhanced feature map.
[0031] In some embodiments, the formula for the energy function is: ; Wherein, E(x) i ) represents the spatial and channel attention weights of the current location features; x i Let x be the feature value at the i-th position in the feature map. j Let be the feature value of the j-th position within the neighborhood of the i-th position; N is the total number of neighborhood positions.
[0032] In some embodiments, in step S2, the multi-scale feature map is input into a feature pyramid network for residual feature enhancement and adaptive spatial fusion to output a refined feature map, specifically including: The high-level feature C5 in the multi-scale feature map is adjusted to 7×7 through adaptive pooling, and then enhanced by 1×1 convolution dimensionality reduction and upsampling. The enhanced feature is fused with the dimensionality-reduced feature M5, and then the feature P5 with residual enhancement is output by 3×3 convolution. Features P2, P3, P4, and P5 are concatenated and then normalized sequentially using 1×1 convolution for dimensionality reduction, 3×3 convolution, and the Sigmoid function to generate a spatial weight map. The spatial weight map is then weighted and aggregated with the original features P2, P3, P4, and P5 to output the refined feature map.
[0033] Specifically, in a preferred embodiment of the present invention, step S2 includes the following sub-steps: 1. Construct a backbone network with embedded deformable convolutions (ResNet50vd-DCN). First, an improved backbone network, ResNet50vd-DCN, is constructed. This network is based on the classic ResNet50 architecture and optimized. The main improvements include: Introducing Deformable Convolution (DCN): In some convolutional layers of Stage 3 and Stage 4, standard convolutions are replaced with deformable convolutions, enabling the model to dynamically learn the offset of sampling points, thereby better perceiving irregularly shaped targets (such as curved fishing nets and irregularly shaped garbage) and improving the ability to fit the edges of targets.
[0034] Add dilated convolution to Stage 4: Change the convolution operation of the last residual block in Stage 4 to a dilated convolution with a dilation rate of 2 to expand the receptive field and improve the ability to capture a wide range of contextual information, while keeping the feature map resolution unchanged.
[0035] Embedding a SimAM attention module at the end of the residual module: SimAM is a parameterless attention module that calculates the spatial and channel attention weights at each location using an energy function. Its specific implementation steps are as follows: Input the feature map of the training sample; for the i-th location x in the feature map... i Select several neighboring locations x within its neighborhood. j And calculate the energy function E(x) according to the formula. i The attention weights w are obtained by normalizing the energy values at all locations using the Softmax function. i ; assign attention weights w i Multiply the original feature map element by element to output the attention-enhanced feature map.
[0036] 2. To further improve model efficiency, channel pruning is performed on the backbone network after training, specifically including: Redundant channels are evaluated based on their importance scores (such as L1 norm or gradient information). Remove the channels with lower scores and fine-tune the model to restore accuracy; After pruning, the number of backbone network parameters is reduced to 3.2M, significantly reducing computing resource consumption and making it easier to deploy on drones or edge devices.
[0037] 3. The multi-scale feature maps C2, C3, C4, and C5 output from the backbone network are input into the feature pyramid network for further processing: (1) Enhanced residual characteristics of high-rise buildings Adaptive pooling is applied to the highest layer feature C5, and its size is adjusted to a fixed size (e.g., 7×7). Next, 1×1 convolution is used for dimensionality reduction to reduce the number of channels and thus reduce the amount of subsequent computation. The feature map size was restored to its original size after upsampling. The generated enhanced features are added to the intermediate features M5, and then a 3×3 convolution is used to extract richer semantic information, finally outputting the residual-enhanced features P5.
[0038] (2) Adaptive spatial fusion The enhanced features P2, P3, P4, and P5 are concatenated along the channel dimension; 1×1 convolution is used to reduce channel dimensionality and reduce redundant information; spatial features are then extracted using 3×3 convolution; finally, the spatial weight map is generated by normalization using the Sigmoid function; the original features P2-P5 are then weighted and aggregated using this weight map to output the final refined feature map.
[0039] The SimAM attention module effectively enhances the representation of key region features without requiring additional parameters, improving the model's ability to distinguish complex backgrounds. This strategy effectively improves the fusion quality of low-level detail features and high-level semantic features, enhancing the model's ability to recognize small targets and complex backgrounds.
[0040] S3. The refined feature map is processed through parallel branching to predict the target category and generate an instance mask.
[0041] In some embodiments, in S3, the parallel branch includes: The refined feature map is compressed by a 1×1 convolutional layer to output a target semantic category probability map and identify the number of target categories. The kernel offset is dynamically learned by a 3×3 deformable convolutional layer to perform coordinate regression and generate an instance mask.
[0042] Specifically, in a preferred embodiment of the present invention, step S3 includes the following sub-steps: 1. Parallel Branch Design: After obtaining the refined feature map, two parallel processing branches are designed for target category prediction and instance mask generation, respectively. These two branches share the refined feature map as input but employ different network layer structures and operations to meet their respective task requirements.
[0043] 2. Target semantic category probabilistic graph prediction branch: First, a 1×1 convolutional layer is used to compress the refined feature map along the channel dimension, mapping the original high-dimensional features to a low-dimensional space. This process not only helps reduce computation but also removes redundant information and enhances the representation of key features.
[0044] Subsequently, the feature vector at each location is normalized using the Softmax function, converting it into a probability value corresponding to each target category, resulting in the final target semantic category probability map. This probability map clearly identifies the probability of each pixel belonging to different target categories, providing a basis for subsequent accurate classification.
[0045] 3. Instance mask generation branch: For instance mask generation, a 3×3 deformable convolutional layer is used to dynamically learn the convolutional kernel offset. Unlike traditional convolution, deformable convolution can dynamically adjust the position of the convolutional kernel sampling points according to the input features, thus capturing detailed information such as target boundaries more flexibly and accurately. Specifically, in this layer, in addition to standard convolution operations, an additional set of offset parameters is learned to guide the movement direction and distance of the sampling points.
[0046] Based on the features extracted by the deformable convolution, a coordinate regression task is further performed to predict the precise contour coordinates of each target instance. Then, these coordinate information are combined to generate a binarized instance mask. This step effectively solves problems such as mutual occlusion or irregular shapes between targets, improving the accuracy of instance segmentation.
[0047] During parallel branch processing, the instance segmentation algorithm exhibits a unique dual capability. When the refined feature map is input into the parallel processing channel, the category prediction branch generates a semantic distribution map through a pixel-by-pixel classification mechanism. This pixel-level resolution capability can accurately distinguish the boundary transition regions of the target, and can completely outline the contour even when the target edge is semi-transparent. At the same time, the mask generation branch utilizes the dynamic sampling characteristics of deformable convolution to automatically adjust the receptive field range when locating micro-targets in the crevices of reefs, so that targets less than five pixels wide can still generate independent mask entities. This collaborative mechanism enables micro-targets (such as shell fragments and barnacle communities scattered in the intertidal zone) to be identified as different categories of entities and to be counted as independent individuals.
[0048] When dealing with high-density target clusters, the algorithm avoids merging of adjacent targets through a spatial offset learning mechanism. For example, when multiple beverage bottles are pushed into a cluster by waves, the offset vector field generated by the coordinate regression branch will cause adjacent bottle caps to produce displacement vectors in opposite directions, ensuring that overlapping targets are separated into independent instances during the mask decoding stage. This characteristic is particularly crucial when statistically analyzing the distribution of small targets (such as cigarette butts on a beach), even if small targets are partially buried in sand and interspersed, they can still be accurately separated and counted. It is worth noting that this process completely abandons the traditional mode that relies on rectangular detection boxes, directly achieving target localization through pixel-level instance encoding, fundamentally eliminating the mask fragmentation problem caused by detection box drift.
[0049] By designing the two parallel branches, the model directly predicts the target class probability after channel compression using a 1×1 convolutional layer, simplifying the model structure while ensuring high classification accuracy. The model's ability to fit complex target edges is enhanced by dynamically adjusting the sampling point offset using a 3×3 deformable convolutional layer, ensuring high-quality instance mask output. Since the two branches are relatively independent, the specific configuration of each branch can be flexibly adjusted or new functional modules can be introduced according to the needs of actual application scenarios.
[0050] S4. Perform nonmaximum suppression and morphological processing on the instance mask, and count the number and location of each type of target; and generate an enriched area heat map based on kernel density estimation according to the location information, and bind the heat map with geographic coordinates and output it to the GIS system.
[0051] After obtaining the instance mask for each target, non-maximum suppression is first applied to eliminate redundant predictions. For multiple masks with high overlap, the mask with the highest confidence score is retained, while other masks are removed. This step effectively reduces duplicate detections and improves the accuracy of the final result. Then, to further improve the quality of the instance masks, a series of morphological operations, such as dilation and erosion, are performed to fill small holes or remove noise. Specifically, dilation helps connect adjacent target regions, while erosion can be used to remove isolated small noise points. By adjusting the size and shape of the structuring elements, the processing effect can be optimized to ensure that the instance masks clearly and accurately represent each target entity.
[0052] After completing the above processing, the number of all targets in the image can be counted, and their position coordinates can be recorded. Assume a set of target position coordinates is obtained. Where p is the total number of targets, y d Let d represent the two-dimensional planar coordinates of the d-th target. Based on the collected target location information, a heat map of the enriched region is generated using the kernel density estimation method.
[0053] Finally, the generated enriched area heatmap is bound to the corresponding geographic coordinates to form a spatial dataset. Then, this dataset is imported into the GIS system using standard data exchange formats (such as GeoJSON, Shapefile, etc.) so that users can intuitively view and analyze the distribution of targets and their enriched areas.
[0054] In some embodiments, in step S4, the enrichment region heatmap is generated by kernel density estimation, and the kernel density estimation formula is: Where y is the coordinate of any point on the heatmap, y a Let be the coordinates of the a-th target, m be the total number of targets, h be the bandwidth parameter, and K be the Gaussian kernel function.
[0055] Furthermore, to better reflect the characteristics of regions with different densities, the bandwidth parameter h is adaptively adjusted according to the target density: h = 0.3 when the target density is ≥ 50 units / km², and h = 0.7 when the target density is < 50 units / km². This strategy ensures good visualization even in sparse regions, avoiding overfitting or underfitting.
[0056] In a preferred embodiment of the present invention, an adversarial training strategy is introduced during the model training phase to enhance the model's robustness to interference factors such as noise and illumination disturbances in the input image, thereby improving the generalization ability and stability of the coastline target recognition method in complex environments. Specifically: 1. Adversarial examples are generated using the FGSM algorithm; Fast Gradient Sign Method (FGSM) is a simple white-box attack method used to generate adversarial examples. Its basic idea is to add small perturbations in the direction of increase of the loss function based on the gradient information of the input image, thereby generating adversarial examples that mislead the model.
[0057] 2. The adversarial loss function is minimized as follows: ,in, For adversarial loss function; Here, x is the mathematical expectation operator; x is the original training sample, x adv For adversarial examples, It is a discriminator.
[0058] In practical applications, this invention's instance segmentation algorithm can simultaneously output a garbage density heatmap (rich areas) and independent masks (scattered areas), balancing efficiency and accuracy. Compared to traditional methods that separate object detection and semantic segmentation, it reduces redundant computation and improves overall system efficiency. On a self-built coastal garbage dataset, the algorithm improves SOLOv2's mask AP, enhancing segmentation accuracy. Through dynamic feature extraction and attention mechanisms, the model can stably identify small targets (such as plastic fragments) even under strong light and wave interference, improving robustness in complex scenes. Adversarial training strategies significantly reduce false detection rates (e.g., misidentifying shadows as garbage). The algorithm supports real-time processing of large-scale coastline data and is adaptable to multiple platforms such as drones and satellites.
[0059] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.
Claims
1. A method for intelligent recognition of coastline targets based on an improved SOLOv2, characterized in that, Includes the following steps: S1. Acquire images of the coastline area to be monitored using UAV imagery and preprocess them to generate training samples; S2. Input the training samples into a deformable convolutional backbone network to extract multi-scale feature maps; input the multi-scale feature maps into a feature pyramid network for residual feature enhancement and adaptive spatial fusion to output a refined feature map; S3. The refined feature map is processed through parallel branching to predict the target category and generate an instance mask, respectively. S4. Perform nonmaximum suppression and morphological processing on the instance mask, and count the number and location of each type of target; and generate an enriched area heat map based on kernel density estimation according to the location information, and bind the heat map with geographic coordinates and output it to the GIS system.
2. The method for intelligent identification of coastline targets based on improved SOLOv2 according to claim 1, characterized in that, The preprocessing in S1 includes: The image of the coastline area to be monitored is normalized to obtain an image of a preset size; The image of the preset size is enhanced by multi-scale blending to generate training samples; wherein, the multi-scale blending enhancement includes random scaling, illumination perturbation and noise simulation.
3. The method for intelligent identification of coastline targets based on improved SOLOv2 according to claim 1, characterized in that, In S2, the backbone network is a ResNet50vd-DCN structure; A dilated convolution with a dilation rate of 2 is introduced in the Stage 4 layer of the backbone network. A SimAM attention module is embedded at the end of the residual module of the backbone network, and the SimAM attention module calculates the attention weights through an energy function; The backbone network parameters were compressed to 3.2M after channel pruning.
4. The method for intelligent identification of coastline targets based on improved SOLOv2 according to claim 3, characterized in that, The implementation of the SimAM attention module includes the following sub-steps: Calculate the energy function E(x) of the feature map of the input training samples. i ); The attention weight w at each position is obtained through normalization. i ; The attention weight w i The feature map is multiplied element-wise with the feature map of the input training sample to obtain the attention-enhanced feature map.
5. The method for intelligent identification of coastline targets based on the improved SOLOv2 according to claim 4, characterized in that, The formula for the energy function is: ; Wherein, E(x) i ) represents the spatial and channel attention weights of the current location features; x i Let x be the feature value at the i-th position in the feature map. j Let be the feature value of the j-th position within the neighborhood of the i-th position; N is the total number of neighborhood positions.
6. The method for intelligent identification of coastline targets based on improved SOLOv2 according to claim 1, characterized in that, In step S2, the multi-scale feature map is input into a feature pyramid network for residual feature enhancement and adaptive spatial fusion, outputting a refined feature map, specifically including: The high-level feature C5 in the multi-scale feature map is adjusted to 7×7 through adaptive pooling, and then enhanced by 1×1 convolution dimensionality reduction and upsampling. The enhanced feature is fused with the dimensionality-reduced feature M5, and then the feature P5 with residual enhancement is output by 3×3 convolution. Features P2, P3, P4, and P5 are concatenated and then normalized sequentially using 1×1 convolution for dimensionality reduction, 3×3 convolution, and the Sigmoid function to generate a spatial weight map. The spatial weight map is then weighted and aggregated with the original features P2, P3, P4, and P5 to output the refined feature map.
7. The method for intelligent identification of coastline targets based on improved SOLOv2 according to claim 1, characterized in that, In S3, the parallel branches include: The refined feature map is compressed by a 1×1 convolutional layer to output a target semantic category probability map and identify the number of target categories. The kernel offset is dynamically learned by a 3×3 deformable convolutional layer to perform coordinate regression and generate an instance mask.
8. The method for intelligent identification of coastline targets based on improved SOLOv2 according to claim 1, characterized in that, In step S4, the heat map of the enriched region is generated through kernel density estimation, and the kernel density estimation formula is as follows: Where y is the coordinate of any point on the heatmap, y a Let be the coordinates of the a-th target, m be the total number of targets, h be the bandwidth parameter, and K be the Gaussian kernel function.
9. The method for intelligent identification of coastline targets based on improved SOLOv2 according to claim 8, characterized in that, The bandwidth parameter h is adaptively adjusted according to the target density: h=0.3 when the target density is ≥50 targets / km², and h=0.7 when the target density is <50 targets / km².
10. The method for intelligent identification of coastline targets based on improved SOLOv2 according to claim 3, characterized in that, During the training phase, adversarial training strategies are employed, including: Adversarial examples are generated using the FGSM algorithm; The adversarial loss function is minimized as follows: ,in, For adversarial loss function; Here, x is the mathematical expectation operator; x is the original training sample, x adv For adversarial examples, It is a discriminator.