Image segmentation system and method for unmanned aerial vehicle seaweed bed monitoring
By using Gaussian energy prototype-driven grouping and mapping module and channel spatial feature fusion module in seaweed bed monitoring, dynamic prototype grouping and two-dimensional feature fusion, the problem of low monitoring accuracy of seaweed bed in traditional methods is solved, and high precision and high generalization ability of seaweed bed image segmentation is achieved.
Patent Information
- Application Number
- CN202510368493.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-27
AI Technical Summary
Traditional seaweed bed monitoring methods have problems such as time-consuming and labor-intensive, limited coverage, and subjectivity and error in data acquisition and analysis, especially in the morphology and color diversity of seaweed beds and the segmentation accuracy is not high under complex backgrounds.
The grouping and mapping module (GEPGMM) and channel spatial feature fusion module (CSFFM) driven by Gaussian energy prototypes improve the segmentation performance of the model through dynamic prototype grouping and two-dimensional feature fusion.
It significantly improves the segmentation accuracy and generalization ability of the model, can effectively deal with the diversity of morphology and color of seaweed beds, reduce feature loss and domain mismatch problems, and improve the accuracy of cross-domain small sample segmentation.
Smart Images

Figure CN120219749A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image segmentation based on computer vision, and particularly relates to an image segmentation system and method for monitoring seaweed beds by unmanned aerial vehicles. Background Art
[0002] With the degradation of the global marine ecosystem and the gradual reduction of seaweed beds, the monitoring and protection of seaweed beds have become particularly important. Seaweed beds are a key component of the marine ecosystem, providing habitats and food sources for numerous marine organisms. Seaweed beds not only play an important role in maintaining marine biodiversity, but also play a key role in the carbon cycle, climate regulation, and marine ecological balance. However, due to factors such as overfishing, pollution, and climate change, the distribution and health status of seaweed beds are facing serious threats. According to relevant studies, the coverage area of global seaweed beds has decreased significantly in the past few decades, resulting in many marine organisms losing important habitats and food sources. Traditional seaweed bed monitoring methods mainly rely on manual diving surveys and underwater cameras. These methods are not only time-consuming and laborious, but also have limited coverage, making it difficult to meet the needs of large-scale monitoring. Manual diving surveys require professional personnel to conduct underwater operations, which are not only costly, but also greatly affected by weather and sea conditions, making it difficult to conduct in harsh environments. Although underwater cameras can provide relatively intuitive image data, their coverage is limited, making it difficult to quickly monitor large areas of seaweed beds. In addition, traditional methods have large subjectivity and errors in the process of data acquisition and analysis, making it difficult to achieve high-precision seaweed bed segmentation and identification.
[0003] The rapid development of drone technology has provided a new solution for seaweed bed monitoring. Drones can take high-resolution images from the air, covering large areas of the sea, providing comprehensive data support for the distribution and health status of seaweed beds. The images taken by drones not only have high resolution but also can quickly obtain data over large areas, greatly improving the monitoring efficiency. In addition, the cost of drone technology is relatively low, the operation is simple, and it can operate under various weather and sea conditions, with high flexibility and adaptability. However, the image segmentation and recognition of seaweed beds also face many challenges. First, the morphology and color of seaweed beds vary significantly in different sea areas and different seasons, which makes the segmentation effect based on traditional image processing methods poor. The morphology and color of seaweed beds are affected by many factors, such as water quality, light, temperature, etc., resulting in large variations in different sea areas and different seasons. Second, the images taken by drones often contain complex background information, such as sea water, sandy bottom, and other marine organisms, which will interfere with the segmentation of seaweed beds. The contrast between seaweed beds and the surrounding environment is relatively low, making it particularly difficult to accurately segment seaweed beds in complex backgrounds. In addition, the image data of seaweed beds is usually scarce, especially in some remote sea areas, and it is difficult to obtain a large amount of labeled data to train deep learning models. The problem of scarce data limits the generalization ability and segmentation accuracy of the model, making its application in different sea areas and different conditions restricted.
[0004] Existing few-shot segmentation methods have certain limitations when dealing with seaweed bed images. On the one hand, these methods usually rely on generating prototype features to match the target objects in the query images. However, in seaweed bed images, due to the diversity and complexity of the target objects, a single prototype feature is difficult to accurately capture all the features of seaweed beds, resulting in a decrease in segmentation accuracy. On the other hand, when dealing with cross-domain data, existing methods often cannot effectively cope with the feature distribution differences between the source domain and the target domain, leading to a decrease in segmentation accuracy. There are significant differences in the feature distributions of seaweed bed images in different sea areas, and existing methods are difficult to maintain a high segmentation accuracy in cross-domain tasks. Summary of the Invention
[0005] In view of the above problems, the present invention proposes an image segmentation system and method for drone seaweed bed monitoring. By using a Gaussian energy prototype-driven grouping and mapping module (GEPGMM) and a channel space feature fusion module (CSFFM), the segmentation performance of the model is improved. The model effectively captures the diversity and complexity of seaweed beds using the GEPGMM module, generating more representative prototype features. At the same time, the CSFFM module further enhances the feature representation ability through channel and spatial attention mechanisms, significantly improving the segmentation accuracy, thereby achieving high-precision seaweed bed image segmentation in few-shot segmentation tasks.
[0006] In the first aspect of the present invention, an image segmentation system for monitoring seaweed beds by drones is provided, which is improved based on the Iterative Few-Shot Adaptor (IFA), and takes the seaweed bed images captured by drones as input; it includes a feature grouping and mapping part, a feature fusion part, and a segmentation prediction part; The feature grouping and mapping part includes a decoder and a Grouping Mapping Module Driven by Gaussian Energy Prototype (GEPGMM), which is used for feature extraction and strengthening the attention to key features; the decoder adopts a ResNet-50 model, removes the fully connected layer and the last residual block of the original network, and uses the last three network levels with different depths to extract feature information from the input image; The feature fusion part adopts a Channel-Spatial Feature Fusion Module (CSFFM) to perform feature fusion in the channel and spatial dimensions on the features extracted by the feature grouping and mapping part. The CSFFM module includes a channel attention unit and a spatial attention unit. Among them, the channel attention unit generates channel weights through global pooling operations and convolutions, and the spatial attention unit generates spatial weights through convolution operations; the channel weights and spatial weights are added together to weight the input feature map to obtain an optimized feature map; The segmentation prediction part includes a decoder, and the decoder adopts three segmentation heads, which are responsible for converting the optimized feature map into a final segmentation mask.
[0007] Preferably, the Grouping Mapping Module GEPGMM includes a seed point selection unit and a feature vector extraction unit. The seed point selection unit first uses the Canny edge detection algorithm to extract the edge features of the input image, and then randomly samples to select representative seed points; the feature vector extraction unit extracts the corresponding feature vectors for each seed point, then calculates the feature distance and spatial distance, and updates the feature point center in combination with the Gaussian energy function to finally generate a multi-scale prototype.
[0008] Preferably, for the Channel-Spatial Feature Fusion Module CSFFM, for the input feature map, where the number of channels, height, and width of the feature map are C, H, and W respectively, the channel attention module obtains aggregated features through global average pooling and global maximum pooling, and then uses a shared 1D convolution to generate channel weights ; the spatial attention module generates spatial weights through 3×3 convolution , and finally the channel weights and spatial weights are added together to weight the input feature map to obtain an optimized feature map.
[0009] Preferably, the decoder adopts a ResNet-50 model, removes the fully connected layer and the last residual block Stage5, and retains the first four stages Stage1 - Stage4; the resolution of the feature map output by Stage1 is about 1 / 4 of the input image and contains rich detail information; the resolution of the feature map output by Stage2 is about 1 / 8 of the input image and contains medium-scale features; the resolution of the feature map output by Stage3 is about 1 / 16 of the input image and contains deeper semantic information; the resolution of the feature map output by Stage4 is about 1 / 32 of the input image and contains the deepest semantic information; then, in the feature extraction process, the normalized input image is fed into the ResNet-50 model. The input image first passes through the first convolutional layer and the max-pooling layer, and the image size is halved to extract shallow detail features, obtaining the feature map of Stage1. Then, the middle-level features are further extracted through the second residual block, and the resolution is halved again to obtain the feature map of Stage2. Next, deeper semantic features are extracted through the third residual block, and the resolution is halved again to obtain the feature map of Stage3. Finally, the deepest semantic features are extracted through the fourth residual block, and the resolution is further halved to obtain the feature map of Stage4. Finally, the feature grouping and mapping part outputs the feature map; appropriate feature maps are selected from Stage1 - Stage4 as the input to the grouping and mapping module GEPGMM.
[0010] Preferably, the specific data processing process of the grouping and mapping module GEPGMM is as follows: S1, Use the Canny edge detection algorithm to extract the edge features of the support image, generate an edge map to mark the positions of the edge points, and randomly sample representative seed points from the edge points as the starting points for subsequent feature grouping. S2, For each selected seed point, extract the corresponding feature vector from the feature map Fs of the support image. The feature vector is obtained from the feature map extracted from a certain stage of the ResNet-50 model. S3, Calculate the feature distance Dfc and the spatial distance Dsc between each pixel in the support image feature map Fs and the seed point; the feature distance Dfc is measured by cosine similarity, while the spatial distance Dsc is calculated by Euclidean distance. Combine the smoothing term and the energy function to update the feature point center. S4, Through an iterative optimization process, gradually update the feature point center, and finally generate 5 groups of multi-scale prototypes that can cover target features of different sizes. S5, Expand and map the generated prototypes into the query image for segmentation, and dynamically select the most appropriate prototype by calculating the similarity between the expanded feature points and the query features. S6, generate a response map respmap and a guide map guidemap; the response map respmap represents the maximum similarity between each pixel in the query image and the prototype, and the guide map guidemap represents the prototype index corresponding to each pixel. These two maps are spliced with the original query feature map Fq to generate the optimized query feature a2.
[0011] Preferably, the specific data processing process of the channel space feature fusion module CSFFM is: First, the channel attention mechanism is used to enhance the features of important channels in the feature map, while suppressing unimportant channels. This is achieved through global average pooling GAP and global maximum pooling GMP, respectively, to obtain two feature vectors A and M, whose shapes are (C, 1, 1), where C is the number of channels in the feature map. These two vectors are then concatenated and passed to a 1×1 convolutional layer to generate channel attention weights Wc, and the Sigmoid activation function σ is used to normalize the weights to the range of (0, 1) to highlight important channels and suppress unimportant channels. Next, the spatial attention mechanism is used to enhance the features of important spatial positions in the feature map, while suppressing unimportant positions. First, the mean μ and maximum value max of the input feature map FGEPGMM in the channel dimension are calculated, and its shape is (1, H, W), where H and W are the height and width of the feature map respectively. Then, μ and max are concatenated and passed to a 3×3 convolutional layer to generate the spatial attention weight Ws, which is also normalized using the Sigmoid activation function σ. Finally, the channel attention weight Wc and the spatial attention weight Ws are added to obtain the final combined attention weight, and the input feature map FGEPGMM is weighted to generate the optimized feature map F′.
[0012] A second aspect of the present invention provides an image segmentation method for unmanned aerial vehicle seaweed bed monitoring, comprising the following process: Obtain image data of seaweed beds in target areas through drone photography; Inputting the image data into the image segmentation system as described in the first aspect to segment the target area of the input image; By performing real-time online analysis on image data and outputting segmentation results, the target area can be accurately segmented.
[0013] Compared with the prior art, the present invention has the following beneficial effects: In the cross-domain small sample segmentation task, the present invention effectively solves the problems of feature loss and domain mismatch in traditional methods through dynamic prototype grouping and two-dimensional feature fusion, and significantly improves the segmentation accuracy and generalization ability of the model.
[0014] The GEPGMM module uses Gaussian energy for similarity measurement and can dynamically generate multiple prototypes according to the feature distribution of the support images. These prototypes can not only better capture the features of foreground objects but also effectively avoid the feature loss problem caused by a single prototype in traditional methods. By dynamically selecting the most suitable prototype, the model can more accurately match the foreground objects in the query image, thereby improving the segmentation accuracy. In the images of seagrass beds from the perspective of an unmanned aerial vehicle, this method can effectively handle the diversity of seagrass bed morphology and color, generate more representative prototypes, and improve the segmentation accuracy.
[0015] The CSFFM module fuses the features generated by the GEPGMM through channel and spatial attention mechanisms. The channel attention mechanism enables the model to focus on the most important channel features, while the spatial attention mechanism enables the model to focus on the key spatial positions in the image. This two-dimensional feature fusion not only enhances the feature representation ability but also effectively retains and enhances the important information in the feature map, further improving the segmentation accuracy. In the images of seagrass beds captured by an unmanned aerial vehicle, this method can effectively capture the fine features of seagrass beds and improve the segmentation accuracy.
[0016] Through dynamic prototype grouping and two-dimensional feature fusion, the method of the present invention not only performs excellently on specific datasets but also can maintain a high segmentation accuracy on datasets in different fields and under different conditions, indicating that the method of the present invention has strong generalization ability and can adapt to various complex practical application scenarios. In the images of seagrass beds captured by an unmanned aerial vehicle, this method can adapt to different sea areas, different lighting conditions, and different seagrass bed morphologies, demonstrating strong generalization ability and practical application value.
[0017] The present invention effectively solves the problems of feature loss and domain mismatch in traditional cross-domain few-shot segmentation methods through an innovative method of dynamic prototype grouping and two-dimensional feature fusion, significantly improving the segmentation accuracy and generalization ability of the model, and providing an efficient and practical solution for cross-domain few-shot segmentation tasks. In the segmentation of seagrass bed images from the perspective of an unmanned aerial vehicle, the method of the present invention can significantly improve the segmentation accuracy and enhance the generalization ability of the model, providing strong technical support for the monitoring and protection of seagrass beds. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following description is only one embodiment of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 It is a block diagram of the overall structure of the image segmentation system of the present invention.
[0020] Figure 2 This is the network structure diagram of the improved ResNet-50 in the present invention.
[0021] Figure 3 This is the schematic structural diagram of the group mapping module GEPGMM of the present invention.
[0022] Figure 4 This is the network structure diagram of the channel spatial feature fusion module CSFFM of the present invention.
[0023] Figure 5 This is the schematic diagram of the segmentation result of the present invention on the UAV seaweed bed image. Detailed implementation manners
[0024] The present invention will be further described below in conjunction with specific embodiments.
[0025] The present invention proposes an image segmentation system and method for UAV seaweed bed monitoring. The system structure diagram is as Figure 1 shown. It is improved based on the framework of the Iterative Few-shot Adaptor, and a segmentation network model combining channel spatial features is built. It can be divided into three parts, namely, the feature group mapping part, the feature fusion part, and the segmentation prediction part. The seaweed bed image captured by the UAV is used as the input; the seaweed bed image data captured by the UAV is divided into a support image set and a query image set. The support image set contains labeled images, while the query image set only contains unlabeled images. The data processing process includes: Feature group mapping part: First, the improved ResNet-50 model is used to extract features from the input image, and then the Gaussian energy prototype-driven grouping and mapping module (GEPGMM) is used to group and fuse the extracted features. ResNet-50, with its powerful feature extraction ability and residual structure design, can effectively process complex information in the image and avoid the problem of gradient disappearance in the training of deep networks. In the present invention, the ResNet-50 model is appropriately modified, removing the fully connected layer and the last residual block (Stage5) in the original network, and retaining the first four stages (Stage1-Stage4) to extract multi-scale feature maps. These feature maps not only contain the shallow detail information of the image but also cover the deep semantic information, providing a rich feature basis for subsequent feature fusion and segmentation tasks.
[0026] Feature Fusion Part: The Channel-Spatial Feature Fusion Module (CSFFM) is used to group and fuse the features obtained from the feature grouping and mapping part. The CSFFM module further fuses the features generated by the GEPGMM through channel and spatial attention mechanisms to enhance the feature representation ability. The channel attention mechanism enables the model to focus on the most important channel features, while the spatial attention mechanism enables the model to focus on the key spatial positions in the image. In this way, the model can more accurately identify and focus on the key feature regions in the image, thereby improving the accuracy and robustness of segmentation.
[0027] Segmentation Prediction Part: Finally, the segmentation prediction part includes a decoder, which uses three segmentation heads and is responsible for converting the fused feature map into the final segmentation result. The segmentation heads generate a segmentation mask with the same resolution as the input image through convolution operations and activation functions. During the training process, the output of the segmentation heads is compared with the ground truth mask labels, and the cross-entropy loss function is used to calculate the difference between the prediction result and the ground truth labels, thereby guiding the model to optimize the segmentation result. At the same time, a cross-domain matching loss is introduced to constrain the alignment of the support domain and query domain features, ensuring that the model has good generalization ability in cross-domain tasks. During the training process, the model is first pre-trained on the source domain dataset to learn general feature representations and segmentation capabilities. Then, it is fine-tuned on the target domain dataset to adapt to the specific feature distribution of the target domain. Through this two-stage training strategy, the model can achieve effective knowledge transfer between the source domain and the target domain, significantly improving the accuracy and robustness of cross-domain few-shot segmentation.
[0028] I. Regarding the Improved ResNet-50 Model In the present invention, the feature grouping and mapping part uses an encoder, namely the ResNet-50 model, for feature extraction because of its powerful feature extraction ability and good generalization performance. The structure of the ResNet-50 model is modified by removing the fully connected layer and the last residual block (Stage5), and retaining the first four stages (Stage1 - Stage4). Through this modification, the model can extract multi-scale feature maps while reducing the computational amount and the number of parameters. The structure is as Figure 2As shown. Specifically, the resolution of the feature map output by Stage1 is about 1 / 4 of the input image, containing rich detail information; the resolution of the feature map output by Stage2 is about 1 / 8 of the input image, containing medium-scale features; the resolution of the feature map output by Stage3 is about 1 / 16 of the input image, containing deeper semantic information; the resolution of the feature map output by Stage4 is about 1 / 32 of the input image, containing the deepest semantic information. Then, the feature extraction process is carried out. The normalized input image (size 512×512) is fed into the ResNet-50 model. The input image first passes through the first convolutional layer and the max-pooling layer, and the image size is halved, extracting shallow detail features to obtain the feature map of Stage1. Then, the middle-level features are further extracted through the second residual block, and the resolution is halved again to obtain the feature map of Stage2. Next, deeper semantic features are extracted through the third residual block, and the resolution is halved again to obtain the feature map of Stage3. Finally, the deepest semantic features are extracted through the fourth residual block, and the resolution is further halved to obtain the feature map of Stage4. Finally, the feature grouping and mapping part outputs the feature map. These feature maps cover rich information from shallow details to deep semantics, providing a solid foundation for subsequent feature fusion and segmentation tasks. To prepare for feature fusion, appropriate feature maps (such as P3 and P4) are selected from Stage1 - Stage4 as the input to the grouping and mapping module driven by the Gaussian energy prototype. These feature maps not only contain rich detail information but also cover deep semantic information, and can provide high-quality feature representations for subsequent feature fusion and segmentation tasks.
[0029] II. Regarding the Grouping and Mapping Module GEPGMM The Gaussian Energy Prototype-Driven Grouping and Mapping Module GEPGMM (Gaussian Energy Prototype-Driven Grouping and Mapping Module) is a key module in the present invention, and its structure is as Figure 3 shown, aiming to measure similarity through Gaussian energy, group and map similar features, generate more representative prototypes, and thus capture and optimize foreground object features more effectively. The specific steps are as follows: First, the Canny edge detection algorithm is used to extract the edge features of the support image, generating an edge map to mark the positions of the edge points. Representative seed points are randomly sampled from these edge points. For example, 50 edge points are randomly selected as the initial seed points. These seed points will serve as the starting points for subsequent feature grouping, ensuring the diversity and representativeness of the grouping process.
[0030] Next, for each selected seed point, the corresponding feature vectors are extracted from the feature map Fs of the support image. These feature vectors contain rich semantic information and can provide a basis for subsequent feature grouping and prototype generation. The feature vectors are obtained from the feature maps extracted at a certain stage (such as Stage4) of the ResNet-50 model, and these feature maps have high semantic expression ability.
[0031] Then, calculate the feature distance Dfc and the spatial distance Dsc between each pixel in the support image feature map Fs and the seed points. The feature distance Dfc can be measured by cosine similarity, while the spatial distance Dsc can be calculated by Euclidean distance. Combining the smoothing term and the energy function, update the feature point center. The specific formula is as follows:
[0032] where, Dfc ( p , Ci ) = 1 - cos( θp , Ci ) and θp , Ci is the angle between the feature vectors Fs ( p ) and Ci ; Dsc ( p , Ci ) = ∥ p - Ci ∥2. λ is the weight coefficient used to balance the influence of the feature distance and the spatial distance.
[0033] Through an iterative optimization process (such as performing 5 iterations), gradually update the feature point center, and finally generate 5 groups of multi-scale prototypes. These prototypes can cover target features of different sizes and provide a richer feature representation for subsequent segmentation tasks.
[0034] Subsequently, expand and map the generated prototypes into the query image for segmentation. By calculating the similarity between the expanded feature points and the query features, dynamically select the most appropriate prototype. The similarity calculation formula is as follows:
[0035] where, is the feature of pixel p in the query image, is the center of the i-th feature point, is the standard deviation of the Gaussian function, which is used to control the attenuation rate of the similarity.
[0036] Finally, the response map respmap and the guidance map guidemap are generated. The response map respmap represents the maximum similarity between each pixel in the query image and the prototypes, while the guidance map guidemap represents the prototype index corresponding to each pixel. These two maps are concatenated with the original query feature map Fq to generate the optimized query feature a2. This process provides richer feature information for subsequent segmentation tasks and enhances the model's ability to recognize foreground objects.
[0037] Through the above steps, the GEPGMM module can effectively group and map similar features, generate more representative prototypes, and thus better capture and optimize foreground object features. This not only improves the segmentation accuracy of the model in cross-domain few-shot segmentation tasks but also enhances its generalization ability, enabling it to maintain high segmentation performance under different domains and conditions.
[0038] III. Regarding the Channel-Space Feature Fusion Module CSFFM The Channel-Space Feature Fusion Module CSFFM (Channel-Space Feature Fusion Module) is a key module in the present invention, aiming to fuse feature maps through channel and spatial attention mechanisms to enhance the feature representation ability and improve the segmentation accuracy. Its structure is as Figure 4 shown. Specifically, the CSFFM module starts from the feature maps output by the Gaussian energy prototype-driven grouping and mapping module (GEPGMM). Although these feature maps contain rich semantic and spatial information, they need to be further optimized to improve the segmentation performance.
[0039] First, the module enhances the features of important channels in the feature map through the channel attention mechanism while suppressing unimportant channels. This process is achieved through global average pooling (GAP) and global max pooling (GMP), obtaining two feature vectors A and M with the shape of (C, 1, 1), where C is the number of channels of the feature map. These two vectors are then concatenated and passed to a 1×1 convolutional layer to generate the channel attention weight Wc, and the Sigmoid activation function σ is used to normalize the weight to the range of (0, 1) to highlight important channels and suppress unimportant channels. The specific formula is as follows:
[0040]
[0041]
[0042] Subsequently, the module enhances the features of important spatial positions in the feature map through the spatial attention mechanism while suppressing unimportant positions. This process first calculates the mean μ and the maximum max of the input feature map FGEPGMM in the channel dimension, with shapes of (1, H, W), where H and W are the height and width of the feature map, respectively. Then, μ and max are concatenated and passed to a 3×3 convolutional layer to generate the spatial attention weight Ws, which is also normalized using the Sigmoid activation function σ. The specific formula is as follows:
[0043]
[0044] Finally, the channel attention weight Wc and the spatial attention weight Ws are added to obtain the final combined attention weight, and the input feature map FGEPGMM is weighted to generate the optimized feature map F′. This process is achieved through element-wise multiplication ⊗, and the specific formula is as follows:
[0045] Through this two-dimensional feature fusion, the CSFFM module not only improves the model's ability to recognize foreground objects but also significantly enhances the segmentation accuracy. The channel attention mechanism highlights the features of important channels and suppresses unimportant channels; the spatial attention mechanism highlights the features of important spatial positions and suppresses unimportant positions. This fusion method provides richer feature information for subsequent segmentation tasks and enhances the overall performance of the model.
[0046] IV. Explanation of Experimental Results In the experiments on the seaweed bed image segmentation task, the network model proposed in the present invention demonstrated excellent performance, and the visualization results are as Figure 5 shown. On the seaweed bed dataset, the model achieved significant performance improvements. As shown in Table 1, the mean intersection over union (mIoU) in the 1-shot and 5-shot scenarios reached 72.5% and 76.8% respectively, representing improvements of 3.0% and 3.8% compared to the baseline method. This result not only proves the efficiency and accuracy of CDGCNet in processing the seaweed bed image segmentation task but also highlights its strong generalization ability in the few-shot segmentation scenario. By introducing the Gaussian energy prototype-driven grouping mapping module (GEPGMM) and the channel-spatial feature fusion module (CSFFM), the model can effectively capture and optimize the key features of the seaweed bed area while reducing background interference. These innovative designs give CDGCNet significant advantages in the field of seaweed bed image segmentation and provide strong technical support for marine ecological monitoring and protection work.
[0047] Table 1 Segmentation Results on the Self-made Dataset
[0048] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and changes can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
[0049] Although the specific implementation manners of the present invention have been described above, they do not limit the protection scope of the present invention. Those skilled in the art should understand that, based on the technical solutions of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.
Claims
1. An image segmentation system for monitoring seaweed beds using an unmanned aerial vehicle, characterized in that: The improvement is based on the iterative small sample adapter IFA, which takes the seaweed bed images taken by drones as input; include Feature grouping mapping part, feature fusion part and segmentation prediction part; The feature group mapping part includes a decoder and a Gaussian energy prototype driven group mapping module GEPGMM, which is used for feature extraction and strengthening the focus on key features; the decoder adopts the ResNet-50 model, removes the fully connected layer and the last residual block of the original network, and uses the last three network layers of different depths to extract feature information from the input image; The feature fusion part adopts a channel space feature fusion module CSFFM to fuse the features extracted by the feature grouping mapping part in channel and spatial dimensions. The CSFFM module includes a channel attention unit and a spatial attention unit, wherein the channel attention unit generates channel weights through global pooling operations and convolutions, and the spatial attention unit generates spatial weights through convolution operations; Add the channel weight and the spatial weight to weight the input feature map to obtain the optimized feature map; The segmentation prediction part includes a decoder, which uses three segmentation heads and is responsible for converting the optimized feature map into a final segmentation mask.
2. The image segmentation system for monitoring seaweed beds by an unmanned aerial vehicle according to claim 1, characterized in that: The group mapping module GEPGMM includes a seed point selection unit and a feature vector extraction unit. The seed point selection unit first uses the Canny edge detection algorithm to extract the edge features of the input image, and then selects representative seed points by random sampling; the feature vector extraction unit extracts the corresponding feature vector for each seed point, then calculates the feature distance and spatial distance, and updates the feature point center in combination with the Gaussian energy function to finally generate a multi-scale prototype.
3. The image segmentation system for monitoring seaweed beds by an unmanned aerial vehicle according to claim 1, characterized in that: The channel space feature fusion module CSFFM, for the input feature map, where the number of channels, height and width of the feature map are C, H and W respectively, the channel attention module obtains the aggregated features through global average pooling and global maximum pooling, and then uses the shared 1D convolution to generate the channel weight ; The spatial attention module generates spatial weights through 3×3 convolution ,Finally, the channel weights and spatial weights are added together to weight the input feature map to obtain the optimized feature map.
4. The image segmentation system for monitoring seaweed beds by an unmanned aerial vehicle according to claim 1, characterized in that: The decoder adopts the ResNet-50 model, removes the fully connected layer and the last residual block Stage5, and retains the first four stages Stage1-Stage4; the resolution of the feature map output by Stage1 is about 1 / 4 of the input image, containing rich detail information; the resolution of the feature map output by Stage2 is about 1 / 8 of the input image, containing medium-scale features; the resolution of the feature map output by Stage3 is about 1 / 16 of the input image, containing deeper semantic information; the resolution of the feature map output by Stage4 is about 1 / 32 of the input image, containing the deepest semantic information; then, the feature extraction process is carried out, and the normalized input image is sent to ResNet-50 Model, the input image first passes through the first convolutional layer and the maximum pooling layer, the image size is halved, and the shallow detail features are extracted to obtain the feature map of Stage1. Then, the middle-level features are further extracted through the second residual block, and the resolution is halved again to obtain the feature map of Stage2. Then, the deeper semantic features are extracted through the third residual block, and the resolution is halved again to obtain the feature map of Stage3. Finally, the deepest semantic features are extracted through the fourth residual block, and the resolution is further halved to obtain the feature map of Stage4. Finally, the feature grouping and mapping part outputs the feature map; select the appropriate feature map from Stage1-Stage4 as the input of the grouping and mapping module GEPGMM.
5. The image segmentation system for monitoring seaweed beds by an unmanned aerial vehicle according to claim 2, characterized in that: The specific data processing process of the group mapping module GEPGMM is as follows: S1, use the Canny edge detection algorithm to extract the edge features of the support image, generate an edge map to mark the position of the edge points, and randomly sample representative seed points from the edge points as the starting point for subsequent feature grouping; S2, for each selected seed point, extract the corresponding feature vector from the feature map Fs of the support image. The feature vector is obtained from the feature map extracted from a certain stage of the ResNet-50 model; S3, calculate the feature distance Dfc and spatial distance Dsc between each pixel and the seed point in the support image feature map Fs; the feature distance Dfc is measured by cosine similarity, while the spatial distance Dsc is calculated by Euclidean distance, combined with smoothing terms and energy functions, to update the feature point center; S4, through the iterative optimization process, the feature point centers are gradually updated, and finally 5 groups of multi-scale prototypes are generated, which can cover target features of different sizes; S5, expands and maps the generated prototypes to the query image for segmentation, and dynamically selects the most appropriate prototype by calculating the similarity between the expanded feature points and the query features; S6, generate a response map respmap and a guide map guidemap; the response map respmap represents the maximum similarity between each pixel in the query image and the prototype, and the guide map guidemap represents the prototype index corresponding to each pixel. These two maps are spliced with the original query feature map Fq to generate the optimized query feature a2.
6. The image segmentation system for monitoring seaweed beds by an unmanned aerial vehicle according to claim 3, characterized in that: The specific data processing process of the channel space feature fusion module CSFFM is as follows: First, the channel attention mechanism is used to enhance the features of important channels in the feature map, while suppressing unimportant channels. This is achieved through global average pooling GAP and global maximum pooling GMP, respectively, to obtain two feature vectors A and M, whose shapes are (C, 1, 1), where C is the number of channels in the feature map. These two vectors are then concatenated and passed to a 1×1 convolutional layer to generate channel attention weights Wc, and the Sigmoid activation function σ is used to normalize the weights to the range of (0, 1) to highlight important channels and suppress unimportant channels. Next, the spatial attention mechanism is used to enhance the features of important spatial positions in the feature map, while suppressing unimportant positions. First, the mean μ and maximum value max of the input feature map FGEPGMM in the channel dimension are calculated, and its shape is (1, H, W), where H and W are the height and width of the feature map respectively. Then, μ and max are concatenated and passed to a 3×3 convolutional layer to generate the spatial attention weight Ws, which is also normalized using the Sigmoid activation function σ. Finally, the channel attention weight Wc and the spatial attention weight Ws are added to obtain the final combined attention weight, and the input feature map FGEPGMM is weighted to generate the optimized feature map F′.
7. An image segmentation method for unmanned aerial vehicle seaweed bed monitoring, characterized in that: The process includes: Obtain image data of seaweed beds in target areas through drone photography; Inputting image data into the image segmentation system according to any one of claims 1 to 6 for segmenting a target area of the input image; By performing real-time online analysis on image data and outputting segmentation results, the target area can be accurately segmented.
Citation Information
Patent Citations
Few-sample document layout analysis method based on metric learning
CN112069961A
Small sample remote sensing image target detection method based on multi-task optimization
CN115049944A
Small sample segmentation method and device based on attention mechanism, terminal and medium
CN116258937A
High-resolution remote sensing image semantic segmentation passive domain adaptive method and device
CN119445113A
Few-shot point cloud semantic segmentation method, and network, storage medium and processor
WO2024113169A1