Deep learning-based enteromorpha detection method and system for remote sensing images
By using the EpM-UNet network for edge enhancement and multi-scale texture extraction, combined with depthwise separable convolution and state-space attention mechanisms, the problems of inaccurate boundary recognition and insufficient long-range dependency modeling in Ulva prolifera remote sensing image detection are solved, achieving efficient and accurate Ulva prolifera detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)
- Filing Date
- 2025-05-20
- Publication Date
- 2026-05-05
AI Technical Summary
Existing remote sensing image detection technologies for Ulva prolifera suffer from inaccurate boundary identification, insufficient long-range dependency modeling, high computational resource consumption, and inadequate multi-scale feature integration, making it difficult to meet the accurate detection requirements of Ulva prolifera in high-resolution remote sensing images.
By employing the EpM-UNet network, edge feature extraction and multi-scale feature fusion are achieved through edge enhancement, small target sensitivity detection, multi-scale texture extraction, and spatial context modeling. By utilizing edge gradient extraction and small target enhancement techniques, combined with depthwise separable convolution and state space attention mechanisms, edge feature extraction and multi-scale feature fusion are realized.
It significantly improved the accuracy, real-time performance, and reliability of remote sensing monitoring of Ulva prolifera, enhanced boundary identification accuracy, reduced positioning errors, and optimized detection performance.
Smart Images

Figure CN120510528B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, specifically to a method and system for detecting seaweed in remote sensing images based on deep learning. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] With increasing awareness of marine ecological environment protection, timely and accurate monitoring of marine plankton, especially outbreaks of Enteromorpha prolifera, has become a crucial issue in marine environmental management and disaster early warning. Large-scale outbreaks of Enteromorpha prolifera not only damage marine ecosystems but also severely impact coastal economic activities such as fisheries and tourism. In recent years, the widespread application of high spatial resolution remote sensing images (HSRIs) has provided abundant data support for large-scale, fine-grained detection of Enteromorpha prolifera. However, due to the complexity of Enteromorpha prolifera's spatial distribution, spectral characteristics, and scale variations, achieving accurate detection still faces numerous technical challenges.
[0004] Existing methods for detecting Ulva prolifera can be broadly categorized into two types: traditional methods based on manual feature extraction, such as those using color features, texture features, threshold segmentation, and classifiers (e.g., SVM, decision trees); and methods based on convolutional neural networks (CNNs), which have emerged in recent years and utilize deep learning techniques to automatically learn image features and segment Ulva prolifera. While CNN-based detection methods offer significant advantages over traditional methods in terms of automation and generalization ability in feature extraction, they still have limitations when applied to high-resolution remote sensing images:
[0005] (1) Inaccurate boundary recognition: Existing CNN models have difficulty accurately capturing the blurry and scattered boundary features of the seaweed area, resulting in blurry boundaries and large positioning errors in the detection results;
[0006] (2) Insufficient long-range dependency modeling: Due to the limited receptive field size, traditional CNN structures are difficult to effectively model the spatial distribution patterns and long-range correlations of Ulva prolifera in large water areas, which affects the overall segmentation accuracy;
[0007] (3) High computational resource consumption: In order to improve performance, existing CNN models often improve detection capabilities by increasing the number of network layers and parameters, which leads to a decrease in model inference speed and makes it difficult to meet the real-time and low-energy consumption requirements in practical applications.
[0008] (4) Insufficient integration of multi-scale features: Ulva prolifera exhibits obvious multi-scale features in remote sensing images. Existing methods have failed to fully integrate the semantic information of images at different scales, and to take into account both global recognition and local details.
[0009] It is evident that existing deep learning methods struggle to address the challenges of multi-scale feature extraction, edge blurring, and small-area target recognition in Ulva prolifera remote sensing images, and cannot effectively handle the noise and cloud interference in complex marine environments. Summary of the Invention
[0010] To address the aforementioned problems, this invention proposes a deep learning-based remote sensing image detection method and system for *Ulva prolifera*. It introduces the EpM-UNet network, which, through edge enhancement, sensitive small target detection, multi-scale texture extraction, and spatial context modeling, achieves accurate boundary segmentation and long-range dependency modeling, thereby improving the accuracy, real-time performance, and reliability of *Ulva prolifera* remote sensing monitoring.
[0011] To achieve the above objectives, the present invention adopts the following technical solution:
[0012] One or more embodiments provide a deep learning-based remote sensing image detection method for *Ulva prolifera*, comprising the following steps:
[0013] For the acquired remote sensing image to be detected, edge gradient extraction and small target enhancement processing are performed to obtain the first feature map after edge enhancement and small target feature enhancement.
[0014] The first feature map is transmitted to a U-shaped backbone network composed of cascaded multi-level Ep-VSS block modules for multi-level feature extraction. The detection result of the Ulva prolifera remote sensing image is obtained based on the feature map output by the last-level Ep-VSS block module.
[0015] Each Ep-VSS block module includes two processing branches. One branch extracts multi-scale texture features from the input feature map based on a deep separable convolutional network. The other branch models the spatial context information of the medium-to-long-range dependencies of the input feature map through a state-space attention mechanism. The outputs of the two branches are superimposed on the input feature map of the current Ep-VSS block module to obtain the output feature map of the current Ep-VSS block module, which is then input to the next Ep-VSS block module for further feature extraction.
[0016] One or more embodiments provide a deep learning-based remote sensing image detection system for *Ulva prolifera*, comprising:
[0017] The GEDM module is configured to perform edge gradient extraction and small target enhancement processing on the acquired remote sensing image to be detected, and obtain the first feature map after edge enhancement and small target feature enhancement.
[0018] Detection module: It is configured to transmit the first feature map to a U-shaped backbone network composed of cascaded multi-level Ep-VSS block modules, perform multi-level feature extraction, and obtain the detection result of the Ulva prolifera remote sensing image based on the feature map output by the last-level Ep-VSS block module;
[0019] Each Ep-VSS block module includes two processing branches. One branch extracts multi-scale texture features from the input feature map based on a deep separable convolutional network. The other branch models the spatial context information of the medium-to-long-range dependencies of the input feature map through a state-space attention mechanism. The outputs of the two branches are superimposed on the input feature map of the current Ep-VSS block module to obtain the output feature map of the current Ep-VSS block module, which is then input to the next Ep-VSS block module for further feature extraction.
[0020] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0021] This implementation significantly improves the accuracy of *Ulva prolifera* region boundary recognition and reduces localization errors caused by blurred boundaries by introducing edge gradient extraction and small target enhancement techniques. Simultaneously, a dual-branch feature extraction method combining depthwise separable convolution and state-space attention mechanisms not only enhances the ability to capture multi-scale textures in images but also overcomes the receptive field limitations of traditional CNN models, enabling effective modeling of long-range dependent features of *Ulva prolifera* in vast water areas, thereby significantly improving the overall detection and segmentation accuracy. Furthermore, the modularly cascaded U-shaped backbone network design provides excellent multi-level information fusion capabilities during feature extraction, further optimizing detection performance.
[0022] The advantages of the present invention, as well as its additional advantages, will be described in detail in the following specific embodiments. Attached Figure Description
[0023] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute a limitation thereof.
[0024] Figure 1 This is a schematic diagram of the EpM-UNet network structure in Embodiment 1 of the present invention;
[0025] Figure 2 This is a schematic diagram of the GEDM module according to Embodiment 1 of the present invention;
[0026] Figure 3 This is a schematic diagram of the Ep-SS2D scan expansion operation and scan merging operation in Embodiment 1 of the present invention;
[0027] Figure 4 This is a schematic diagram of the scanning direction in Embodiment 1 of the present invention;
[0028] Figure 5 This is a flowchart of the method for detecting Ulva prolifera remote sensing images according to Embodiment 1 of the present invention;
[0029] Figure 6 This is a segmentation effect diagram of the distribution feature map of Ulva prolifera in different environments in the experiment of Embodiment 1 of the present invention;
[0030] Figure 7 This is a segmentation effect diagram of the distribution feature map of Ulva prolifera in different environments obtained by removing different modules in the ablation experiment of EpM-UNet in Embodiment 1 of the present invention. Detailed Implementation
[0031] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0032] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0033] It should be noted that the terminology used herein is for describing particular embodiments only and is not intended to limit the exemplary embodiments of the present invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof. It should be noted that, without conflict, the various embodiments and features within those embodiments can be combined with each other. The embodiments will now be described in detail with reference to the accompanying drawings.
[0034] Example 1
[0035] In one or more of the technical solutions disclosed in the embodiments, such as Figures 1 to 7 As shown, a deep learning-based remote sensing image detection method for *Ulva prolifera* includes the following steps:
[0036] Step 1: For the acquired remote sensing image to be detected, perform edge gradient extraction and small target enhancement processing to obtain the first feature map after edge enhancement and small target feature enhancement;
[0037] Step 2: Transmit the first feature map to a U-shaped backbone network composed of cascaded multi-level Ep-VSS block modules for multi-level feature extraction. Based on the feature map output by the last-level Ep-VSS block module, obtain the detection result of the Ulva prolifera remote sensing image.
[0038] Each Ep-VSS block module includes two processing branches. One branch extracts multi-scale texture features from the input feature map based on a deep separable convolutional network. The other branch models the spatial context information of the input feature map based on the state space attention mechanism (Ep-SS2D). The outputs of the two branches are superimposed on the input feature map of the current Ep-VSS block module to obtain the output feature map of the current Ep-VSS block module, which is then input to the next Ep-VSS block module for further feature extraction.
[0039] In this embodiment, the above-mentioned multi-scale texture features of the input feature map are extracted based on the depthwise separable convolutional network, and the spatial context information of the medium- and long-distance dependence of the input feature map is modeled. The input feature map is the input feature map to the current level Ep-VSS block module.
[0040] In this embodiment, the method first processes the input remote sensing image using an edge gradient extraction algorithm to effectively enhance the boundary features of the *Ulva prolifera* region in the image. Simultaneously, a small target enhancement technique is employed to improve the detection sensitivity for small, sparsely distributed *Ulva prolifera* targets, generating a first feature map. Subsequently, the first feature map is input into a U-shaped backbone network, which consists of multiple Ep-VSS block modules connected in series. Each Ep-VSS block module includes two parallel feature processing branches. One branch, based on a depthwise separable convolutional network, is responsible for extracting image texture features from different scales, enhancing the recognition ability of boundary details and target textures. The other branch employs the state-space attention mechanism Ep-SS2D, specifically modeling long-distance spatial context dependencies in the image, thereby overcoming the problem of limited receptive fields in traditional convolutional neural networks. The features extracted by the two branches are fused with the current input feature map and output to the next-level Ep-VSS block module, progressively extracting more complex feature information. Finally, the feature map output by the last-level Ep-VSS block module is used for the detection and localization of the *Ulva prolifera* region.
[0041] This implementation significantly improves the accuracy of *Ulva prolifera* region boundary recognition and reduces localization errors caused by blurred boundaries by introducing edge gradient extraction and small target enhancement techniques. Simultaneously, a dual-branch feature extraction method combining depthwise separable convolution and state-space attention mechanisms not only enhances the ability to capture multi-scale textures in images but also overcomes the receptive field limitations of traditional CNN models, enabling effective modeling of long-range dependent features of *Ulva prolifera* in vast water areas, thereby significantly improving the overall detection and segmentation accuracy. Furthermore, the modularly cascaded U-shaped backbone network design provides excellent multi-level information fusion capabilities during feature extraction, further optimizing detection performance.
[0042] Step 1, the edge gradient extraction method, includes the following steps:
[0043] Step 11.1: Using a convolutional layer with a fixed kernel, calculate the image gradient features of the input remote sensing image to be detected by performing convolution operations in the vertical and horizontal directions respectively.
[0044] Step 11.2: Fuse the obtained image gradient features into the original remote sensing image to be detected to obtain the feature map after edge enhancement;
[0045] Step 1, the method for enhancing small targets, includes the following steps:
[0046] Step 12.1: Extract multi-scale features from the enhanced feature map using multi-dilation rate convolution at different resolutions to obtain texture and edge information at different scales, and then activate it using the ReLU function;
[0047] Step 12.2: Concatenate the features activated by the ReLU function along the channel dimension to obtain a multi-scale feature map;
[0048] Step 12.3: The obtained multi-scale feature maps are weighted and fused using an attention mechanism, and then the first feature map is output after convolution operation;
[0049] In step 1, edge gradient extraction and small target enhancement processing can be implemented using the constructed feature extractor module GEDM (Gradient and Enhancement Detection Module), such as... Figure 1 As shown;
[0050] To address two key challenges in semantic segmentation of Ulva prolifera remote sensing images—the extraction of Ulva prolifera edge features and the difficulty in effectively detecting small, scattered Ulva prolifera patches—the GEDM module was constructed, which includes a gradient feature extraction submodule and a small target enhancement submodule.
[0051] The gradient feature extraction submodule uses fixed convolution kernels in the vertical and horizontal directions to calculate the gradient of the input feature map, enhancing the edge information of the seaweed to identify and preserve the boundary information between the seaweed and the seawater.
[0052] The small target enhancement submodule uses multi-dilation rate convolution to extract multi-scale features and combines channel attention mechanism to weight and enhance channel features to improve the detection capability of small-area seaweed patches.
[0053] The GEDM module enables the original features, gradient features, and small target enhancement features to complement each other, resulting in a first feature map that comprehensively represents the information of Ulva prolifera.
[0054] like Figure 2 As shown, the gradient feature extraction submodule includes two parallel convolutional layers ConvX and ConvY with fixed kernels connected in sequence, as well as the first fusion module Concat;
[0055] In the gradient feature extraction submodule, two parallel convolutional layers with fixed kernels, ConvX and ConvY, calculate the image gradient through convolution operations in the vertical and horizontal directions, respectively; the first fusion module, Concat, fuses the obtained image gradient into the original remote sensing image to be detected.
[0056] Since seaweed often exhibits complex edge structures in marine environments, image gradients are calculated through convolution operations in both the vertical and horizontal directions to enable the network to focus on the edge information of seaweed.
[0057] Optionally, the convolutional layers ConvX and ConvY use the following 3×3 Sobel convolution kernels in the horizontal and vertical directions respectively: a horizontal kernel (for detecting vertical edges) and a vertical kernel (for detecting horizontal edges). The Sobel convolution kernel formula is as follows:
[0058] Horizontal core:
[0059] (1)
[0060] Vertical core:
[0061] (2)
[0062] By convolving the input feature map with the aforementioned convolutional kernel, the gradient components in the horizontal and vertical directions are obtained, respectively. and Then, the square root of the sum of the squares of these two components is calculated to obtain the gradient magnitude. The gradient calculation formula is as follows:
[0063] (3)
[0064] Where G is the gradient magnitude, i.e., the gradient result, and ϵ is a small constant (1e-6) set to ensure numerical stability.
[0065] In the above scheme of this embodiment, gradient calculation is applied to each channel of the input feature, which preserves the structural information of the multi-channel feature. It can effectively capture the boundary information between the seaweed and the sea surface, improve the model's sensitivity to the edge of the seaweed, and solve the problem that existing CNN models have difficulty accurately capturing the boundary features of the seaweed area, resulting in blurred detection results and large positioning errors.
[0066] like Figure 2 As shown, the Small Object Enhancement submodule includes multiple parallel multi-scale dilated convolutional layers connected in sequence, a second fusion module Concat, and a first attention module (CoordAttention).
[0067] Multi-scale dilated convolutional layers perform dilated convolution operations on the input features (output features of the gradient feature extraction submodule) with different dilation rates to obtain different scales to capture texture and edge information at different scales, and are activated by the ReLU function;
[0068] The second fusion module (Concat) is configured to concatenate the features activated by the ReLU function along the channel dimension to obtain a multi-scale feature map.
[0069] The first attention module (Coord Attention) performs weighted fusion of the obtained multi-scale feature maps through an attention mechanism, and then outputs the result after convolution.
[0070] Multi-scale hollow convolutional layers, specifically, such as Figure 2 As shown, four dilated convolution branches are first applied in parallel to the input feature x, with dilation rates of 1, 2, 4, and 6, respectively, to capture texture and edge information at different scales; each branch is followed by ReLU activation. Subsequently, the features from these four branches are concatenated along the channel dimension to obtain a multi-scale feature map.
[0071] To further preserve spatial location information and enhance channel representation, the first attention module (CoordAttention) introduces the Coordinate Attention mechanism. First, 1D global pooling is performed along the vertical H and horizontal W directions, as shown in Equation 4. Then, a shared 1×1 convolution is used to reduce the dimensionality and activate the feature map. Finally, a branch generates a direction-aware attention map, which weights the original feature map in the H and W dimensions, thereby taking into account both long-range dependencies and accurate location information recognition.
[0072] (4)
[0073] Finally, the weighted features are integrated and output through a 1×1 convolutional layer (conv) based on the output features of the first attention module, resulting in the final first feature map after small target enhancement.
[0074] In this embodiment, to address the issue of feature loss in small-area Ulva prolifera patches during deep learning downsampling pooling in remote sensing images, an enhancement module combining multi-scale dilated convolution and coordinate attention is designed. This module effectively preserves the features of small Ulva prolifera patches, enabling small target detection. The GEDM module, through "gradient feature extraction + multi-scale dilated convolution + coordinate attention," achieves fine-grained edge detection and multi-scale context fusion, allowing the model to maintain high sensitivity to local details while achieving global perception, significantly improving the accuracy of Ulva prolifera boundary recognition and small target detection.
[0075] The gradient feature extraction submodule uses a fixed Sobel convolution kernel to calculate gradients on the input feature map in both the horizontal and vertical directions, thereby highlighting boundary features in the image. The extracted gradient information clearly reflects the transition area between the seaweed and the surrounding seawater, especially the complex edge structures, helping the model accurately identify and distinguish the boundaries of the seaweed while improving sensitivity to local details.
[0076] The small target enhancement submodule processes the input features through multiple parallel, multi-scale dilated convolutional layers. Different dilation rates enable the module to simultaneously extract texture and edge information at both small and large scales, effectively capturing features of different-sized *Ulva prolifera* patches. The features after dilated convolution are concatenated along the channel dimension to form a multi-scale feature map. Subsequently, a coordinate attention mechanism is used to weight and fuse the multi-scale features. This mechanism extracts orientation-aware contextual information through global pooling in the vertical and horizontal directions and weights it in the spatial dimension, balancing long-range dependencies with precise location information recognition. Ultimately, the fused feature map not only retains the detailed information of small *Ulva prolifera* patches but also possesses strong spatial awareness, significantly improving the model's ability to detect small targets.
[0077] It is evident that the gradient feature extraction submodule and small target enhancement submodule set in the GEDM module effectively integrate global and local features, enhancing the model's response capability to small-area, scattered seaweed patches.
[0078] To address the multi-scale and diffuse distribution characteristics of *Ulva prolifera* in remote sensing image segmentation tasks, this embodiment proposes a novel feature extraction module, the Ep-VSS Block module in step 2, whose architecture is as follows: Figure 1 As shown in small figure (b) in the figure. The Ep-VSS Block module effectively integrates the global modeling capability of the state-space model with the local feature extraction advantage of the convolutional neural network through parallel processing paths, while capturing global contextual information and local detailed features, resulting in good extraction performance.
[0079] In some embodiments, in step 2, the first feature map is transmitted to a U-shaped backbone network composed of cascaded multi-level Ep-VSS block modules for multi-level feature extraction. The detection result of the *Ulva prolifera* remote sensing image is obtained based on the feature map output by the last-level Ep-VSS block module. Figure 1 As shown, the U-shaped backbone network includes an encoder and a decoder; the Ep-VSS block modules of each layer at the encoder end are connected by skip connections, and the output features are concatted and transmitted to the corresponding Ep-VSS block modules at the decoder end.
[0080] Figure 1 In this context, Patch Embedding represents the input image... The image is divided into non-overlapping 4×4 patches, and then the dimension of the image is mapped to C, indicated by the red arrow.
[0081] Patch Merging represents a downsampling operation, indicated by a brown arrow;
[0082] Patch Expanding indicates an upsampling operation, represented by a yellow arrow;
[0083] Final Projection represents the final projection layer used to recover the size of the features to match the segmentation target, indicated by the blue arrow;
[0084] Skip Connection, indicated by a black dashed line, effectively reduces gradient vanishing and network degradation problems.
[0085] Element-wise addition represents the addition operation of corresponding elements in two arrays or tensors with the same shape;
[0086] The Ep-VSS block module is represented by a blue-gray box;
[0087] Furthermore, the Ep-VSS block module includes:
[0088] 1) First layer normalization module (LN): It is configured to preprocess the input feature map of the input Ep-VSS block module using layer normalization to obtain the preprocessed feature F1;
[0089] In this embodiment, the Ep-VSS block module has a layer normalization module (LN) at the very front. Normalization helps improve the stability and convergence speed of subsequent linear layers and activation layers. It can effectively stabilize the training process and accelerate convergence.
[0090] Layer normalization can be represented by formula (5):
[0091] (5)
[0092] in, and For learnable scaling and translation parameters, It is a set small constant. This is represented as the feature vector of the i-th input. Represented as the dimension of the feature vector;
[0093] The Ep-VSS block module adopts a residual connection design, which is beneficial for gradient flow and deep network training.
[0094] 2) Local-enhancement branch: Extracts multi-scale texture features from the input image based on a deep separable convolutional network; it may include one or more local feature enhancement units, each of which includes a first deep separable convolutional network (DwConv), a first batch normalization (BN), a point convolution (Conv 1×1), a second batch normalization (BN), and a SiLU activation module connected in sequence. It performs deep convolution, batch normalization, point convolution, batch normalization, and activation operations on the preprocessed feature map F1 to obtain multi-scale texture features;
[0095] Specifically, the local-enhance branch employs a two-stage depthwise separable convolutional structure, with each stage containing depthwise convolutional operations and pointwise convolutional operations, supplemented by batch normalization and SiLU activation functions, to efficiently extract local texture features with multi-scale information while maintaining low computational complexity, thereby enhancing the model's ability to perceive the boundary region of Ulva prolifera.
[0096] The formula for depthwise separable convolution (DwConv) in the above process is as follows:
[0097] (6)
[0098] in, The spatial location of the preprocessed feature map in channel c is given by the input. The value at that location; The kernel is a 3×3 convolution kernel for channel c (radius r=1). It is the relative offset of the convolution kernel; On channel c, spatial position The depth can separate the intermediate outputs of the convolution;
[0099] Point convolution (Conv 1×1) linearly combines all channels at each spatial location. The point convolution is represented as follows:
[0100] (7)
[0101] in, It is a 1×1 convolution kernel used to fuse information from different channels. In the passage Spatial location The final output of point convolution;
[0102] 3) The remote modeling branch, also known as the Adaptive SSM branch, is used to model long-distance dependent spatial context information in images through the state-space attention mechanism (Ep-SS2D).
[0103] The remote modeling branch consists of a first linear layer, a SiLU activation layer, a second depthwise separable convolutional network (DwConv), an adaptive Ep-SS2D state space scanning module, a second-layer normalization module (LN), and a bypass branch, all connected in sequence.
[0104] The first linear layer is used to perform a linear transformation on the preprocessed feature map F1.
[0105] The SiLU activation layer performs a nonlinear mapping on the features after linear transformation.
[0106] The second depthwise separable convolutional network is used to perform depthwise separable convolution operations on the features after nonlinear mapping to extract local spatial information, including edge and texture information.
[0107] The Adaptive Ep-SS2D module for the state space of Ulva prolifera employs a dynamically adjusted 2D scanning module. It introduces adaptive attention in the spatial dimension to perform adaptive scanning, obtaining the feature map F2 for modeling long-range spatial dependencies, thereby enhancing the model's ability to perceive long-distance dependencies.
[0108] The second normalization module (LN) normalizes the feature map F2 after adaptive scanning again to obtain the feature map F3, in order to maintain numerical stability.
[0109] The output of the first-layer normalization module (LN) is also connected to a bypass branch, which includes a second linear layer and an activation layer. The second linear layer performs layer normalization processing on the preprocessed feature F1 output by the first-layer normalization module (LN) and then performs an activation operation to obtain the feature map F4.
[0110] The third linear layer linearly maps the features obtained by fusing feature map F3 and feature map F4 to obtain feature map F5.
[0111] The DropOut layer is used to randomly discard features from feature map F5, thus obtaining the output of the remote modeling branch.
[0112] The feature map F5 is randomly discarded using the DropOut layer, as shown in the following formula:
[0113]
[0114] Specifically, when When =1, the output is ;when When =0, the output is 0.
[0115] 4) The third fusion module is used to perform weighted fusion of the two branch outputs and the input feature maps of the Ep-VSS block module to obtain the output feature map of the Ep-VSS block module;
[0116] In the above scheme of this embodiment, for the spatial distribution characteristics of Ulva prolifera in remote sensing images, the remote modeling branch constructs an adaptive Ulva prolifera state space two-dimensional scanning module (Adaptive Ep-SS2D), which significantly reduces the computational complexity while maintaining the global receptive field through the state space model.
[0117] Ulva prolifera exhibits a large-scale spatial distribution across multiple directions on the sea surface, including irregular edge contours and diverse arrangement patterns. Traditional SS2D scanning methods only model global features in the horizontal and vertical directions, failing to effectively capture key spatial information of Ulva prolifera in diagonal and other directions, thus limiting segmentation accuracy. To address these issues, this embodiment proposes an adaptive Ulva prolifera state space 2D scanning module, namely the AdaptiveEp-SS2D module, suitable for the distribution characteristics of Ulva prolifera in remote sensing images. This module can identify irregular edge contours and diverse arrangement patterns, comprehensively improving the model's ability to model global features of Ulva prolifera in remote sensing images.
[0118] Furthermore, the Adaptive Ep-SS2D state space scanning module for *Ulva prolifera*, such as... Figure 1As shown in section c, it is configured to perform the following procedure:
[0119] Step 21: Perform scan expansion operation: For the two-dimensional feature map input to the Adaptive Ep-SS2D module, perform scan flattening in multiple set directions to obtain multiple scan-expanded one-dimensional sequences;
[0120] Specifically, the scanning in the set direction includes horizontal scanning, vertical scanning, diagonal scanning, and anti-diagonal scanning; the two-dimensional feature map input to the Adaptive Ep-SS2D module is expanded in eight directions along the horizontal, vertical, main diagonal, and anti-diagonal directions;
[0121] The scan expansion operation scans the input two-dimensional feature map according to an eight-directional scanning strategy, such as... Figure 4 The horizontal, vertical, and diagonal scanning methods shown represent horizontal scanning, vertical scanning, diagonal scanning, anti-diagonal scanning, and their reverse order, respectively; flattened, they form a one-dimensional sequence of length L = H × W. k∈[1,2,3,4,5,6,7,8], such as Figure 3 As shown, the one-dimensional sequences obtained after scanning in eight directions are defined as follows:
[0122] (8)
[0123] in, The feature maps obtained from the convolution and activation layers, In ascending order horizontally, Horizontal reverse order, It is in vertical ascending order. It is a vertical reverse order. and Main diagonal and reverse. and For the opposite diagonal and the reverse, For batch size, For the height and width of the feature map, The length of the flattened sequence. The index matrix for the main diagonal scan. For the anti-diagonal scan index matrix, This is a diagonal index in the direction from "bottom right to top left". This is the anti-diagonal index in the "bottom left to top right" direction.
[0124] In this embodiment, the Adaptive Ep-SS2D module introduces diagonal and anti-diagonal scanning directions and performs weighted processing on the directional sequences with obvious spatial features of Ulva prolifera, which can ensure that the model can effectively identify and process the spatial features of Ulva prolifera in various directions.
[0125] Step 22, State-space sequence processing: The one-dimensional sequence in each direction after the scan expansion is modeled using a state-space model. Based on the current observation and the latent state information of the previous time step, dynamic fusion is performed to obtain the output at each time step, which is then decomposed into an output sequence containing delay terms, state parameters, and filter coefficients. ;
[0126] Specifically, the one-dimensional sequence in each direction after the scan expansion is input into the state-space model (SSM Block) for processing. Dynamic fusion is performed using the current observation and the latent state information from the previous time step to capture long-range dependencies and dynamic changes in the sequence. For the first... In one direction, SSM will input the sequence After a series of parameterized linear transformations, the system is decomposed into the various sub-components required for state updates (such as delay terms, state parameters, and filter coefficients), and the output sequence is generated recursively in sequence. ;
[0127] Step 23, Scan and merge operation: Perform a scan on the resulting output sequence. Using a scan-merging mechanism to output multi-directional sequences Merge into a unified global feature map;
[0128] Step 231: Perform global average pooling on the output sequence in each direction to extract the global feature vector of Ulva prolifera in each direction; these feature vectors can reflect the salience and continuity of Ulva prolifera targets in each direction.
[0129] Step 232: Normalize the global feature vectors in each direction to generate a set of adaptive direction weights. ;
[0130] Specifically, a lightweight multilayer perceptron (MLP) and a softmax layer are used to normalize the features in each direction, generating a set of adaptive directional weights. , k∈[1,2,3,4,5,6,7,8].
[0131] Step 232: For the output sequences in each direction, perform a weighted summation according to the corresponding adaptive direction weights, such as... Figure 3 As shown on the right side (b), a comprehensive global feature representation is formed:
[0132] (9)
[0133] Through the above-mentioned processing stages, the Adaptive Ep-SS2D module can not only fully capture the global spatial features of Ulva prolifera using multi-directional scanning, but also efficiently model long-range dependencies through a state-space model and use an adaptive weighting mechanism to focus on key directional features. This allows the model to automatically focus on those directions that are more sensitive to the spatial features of Ulva prolifera, thereby more accurately segmenting the target area while suppressing the interference of background noise and irrelevant information.
[0134] In traditional convolutional neural networks (CNNs), expanding the receptive field typically relies on stacking convolutional layers. However, a single convolutional operation can only cover a local area. To obtain the global receptive field (i.e., acquire information from the entire image), a large number of convolutional layers or dilated convolutions and attention mechanisms are required, leading to a surge in the number of parameters and a significant increase in computational complexity. To address this, the Adaptive Ep-SS2D proposed in this embodiment employs an innovative method that flattens two-dimensional feature maps into a one-dimensional sequence through multi-directional scanning (e.g., diagonal, anti-diagonal, horizontal, and vertical directions). The length of each sequence is equal to the dimension of the image in the corresponding direction, naturally covering all spatial locations along the entire scanning path, thus achieving long-range dependency perception. This method does not rely on multi-layer stacking to expand the receptive field; instead, it directly obtains the implicit global receptive field through sequence-level modeling.
[0135] Traditional global dependency modeling methods, such as self-attention, require calculating the correlation between all position pairs in the input sequence, resulting in a computational complexity of O(L²) (where L is the sequence length). This leads to significant computational and memory overhead when processing long sequences or large images. To address this issue, this embodiment employs a state-space model (SSM) to process the sequence obtained from multi-directional scanning. SSM dynamically updates the sequence using only the current input and the hidden state information from the previous time step through a recursive formula, achieving a computational complexity of O(L²). 2 The sequence length (L) significantly reduces the computational requirements. Furthermore, sequences from different scanning directions can be input into independent state-space models in parallel and processed simultaneously, further improving computational efficiency. Therefore, this design effectively models long-range dependencies while significantly reducing computational complexity, exhibiting excellent performance and scalability.
[0136] The EpM-UNet network constructed in this embodiment includes the aforementioned GEDM module and a U-shaped backbone network composed of cascaded multi-level Ep-VSS block modules. To illustrate the detection performance of the EpM-UNet network constructed in the above-mentioned Ulva prolifera remote sensing image detection method, experiments were conducted, and the details are as follows:
[0137] In this experiment, the FIO-EP dataset was used, a large-scale dataset specifically designed for detecting Ulva prolifera in high-resolution remote sensing imagery (HSRI). This dataset consists of images acquired by multiple satellite sensors (such as GF-1 B / C / D, GF-2, and ZY3-02), covering Ulva prolifera images with varying radiance, cloud cover, and distribution patterns, ensuring data diversity and richness. The dataset contains 1334 images, each with a size of 512×512 pixels. 800 images were randomly selected as the training set, 200 as the validation set, and 334 as the test set.
[0138] 3.3.1) Quantitative evaluation of the EpM-UNet network (ours) and different seaweed detection algorithms;
[0139] The EpM-UNet model is compared and evaluated with mainstream segmentation models such as UltraLight-VM-Unet, H-vmunet, VM-UNet, VM-UNet_V2, and UNet on the FIO-EP dataset. The results are shown in Table 1.
[0140] Table 1 shows the comparison results with existing models;
[0141]
[0142] In the table, F1-score: F1 score is a statistical metric used to measure the accuracy of binary classification (or multi-task binary classification) models.
[0143] IOU: Intersection over Union (IoU). IoU is the result of dividing the overlapping portion of two regions by the sum of the portions of the two regions. It is compared with this IoU result by setting a threshold.
[0144] Accuracy: The percentage of correctly predicted positive (TP) and negative (TN) examples out of the total number of examples.
[0145] Sensitivity: The proportion of all correctly predicted positive samples out of all actual positive samples.
[0146] Specificity: The proportion of all correctly predicted negative samples out of all actual negative samples.
[0147] For quantitative analysis, this embodiment selects three metrics—F1-score, IOU, and Accuracy—to objectively and quantitatively compare the performance of various semantic segmentation techniques in the field of *Ulva prolifera* recognition. These three metrics are presented with evaluation results ranging from 0 to 1: the closer the F1-score is to 1, the more accurate the algorithm's recognition result; the higher the IOU value, the greater the overlap between the algorithm's output and the manually labeled data; and an increase in the Accuracy value means improved accuracy in the classification task. A decrease in any of these metrics indicates a decline in the corresponding performance.
[0148] In addition, Sensitivity and Specificity are introduced as supplementary metrics for evaluating algorithm performance. Sensitivity measures the algorithm's ability to predict real-world *Ulva prolifera* regions, with values ranging from 0 to 1; higher values indicate a lower false positive rate. This metric reflects the completeness of the algorithm's correct identification of target regions. Specificity, on the other hand, quantifies the algorithm's ability to correctly exclude non-*Ulva prolifera* regions, also with values ranging from 0 to 1; higher values indicate a lower false positive rate. This metric evaluates the algorithm's accuracy in distinguishing between background and target. The specific calculation method uses the following mathematical formula for quantitative analysis.
[0149] (10)
[0150] (11)
[0151] (12)
[0152] (13)
[0153] (14)
[0154] In the formula, and These represent the number of correctly detected and incorrectly detected seaweed pixels in the image, respectively. and These represent the number of correctly detected background pixels and the number of incorrectly detected background pixels in the image, respectively.
[0155] In the table, UltraLight-VM-Unet is the model from the paper "Wu R, Liu Y, Liang P, et al. Ultralight vm-unet: Parallel vision mamba significantly reduces parameters for skin lesion segmentation[J]. arXiv preprint arXiv:2403.20035, 2024.";
[0156] H-vmunet: refers to the model in the paper "Wu R, Liu Y, Liang P, et al. H-vmunet: High-orderVision Mamba UNet for medical image segmentation[J].Neurocomputing, 2025, 624(000).DOI:10.1016 / j.neucom.2025.129447.";
[0157] VM-UNet: The model in the paper "Ruan J, Li J, Xiang S.VM-UNet: Vision Mamba UNet for Medical Image Segmentation[J].2024."
[0158] VM-UNet_V2: This is the model from the paper "Zhang M, Yu Y, Jin S, et al. VM-UNET-V2: rethinkingvision mamba UNet for medical image segmentation[C] / / International Symposium on Bioinformatics Research and Applications. Singapore: Springer NatureSingapore, 2024: 335-346.";
[0159] UNet: Ronneberger O, Fischer P, Brox T. U-net: Convolutional networks for biomedical image segmentation[C] / / Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference,Munich, Germany, October 5-9, 2015, proceedings, part III 18. Springerinternational publishing, 2015: 234-241.” model in;
[0160] As shown in Table 1, the method (ours) in this embodiment exhibits significantly superior performance compared to other models on various segmentation metrics on the test set. Experiments demonstrate that by fusing the global dependency modeling capability of the state-space model with the local feature extraction advantages of depthwise separable convolution, the performance of semantic segmentation of *Ulva prolifera* remote sensing images is effectively improved. The Ep-VSSBlock module significantly enhances the model's ability to perceive the complex spatial distribution characteristics of *Ulva prolifera* while maintaining computational efficiency. Notably, EpM-UNet achieves a miou coefficient of 0.8519 and an F1-score of 0.9201, significantly outperforming other models.
[0161] 3.3.2) Qualitative evaluation of the EpM-UNet network and different seaweed detection algorithms;
[0162] On the FIO-EP Ulva prolifera remote sensing high-resolution image dataset, the proposed EpM-UNet model was compared with several mainstream segmentation networks, including UltraLight-VM-UNet, H-vmunet, VM-UNet, VM-UNet_V2, and UNet, in multiple rounds of comparative experiments to comprehensively evaluate the performance of EpM-UNet. Representative image examples were selected based on the distribution characteristics of Ulva prolifera under different environments, and the segmentation effects of each model were visually demonstrated through comparison.
[0163] Figure 6 The differences in segmentation results from different models are marked with circular boxes, further highlighting the superiority of EpM-UNet. The original image (Image) is in the first column, the label image (Label) is in the second column, and the other columns show the results obtained using different detection algorithms;
[0164] from Figure 6 As can be seen from the selected areas in rows a and d of the graph, the EpM-UNet network (Ours) method in this embodiment performs excellently in the identification of seaweed boundaries, and is particularly accurate in the segmentation of continuous large areas; Figure 6 EpM-UNet can accurately detect and completely segment small areas of scattered Ulva prolifera in the b-row and f-row images; Figure 6 In the diffuse distribution region shown in the middle row of the graph, the EpM-UNet network significantly outperforms other models in recognizing fragmented and sparse targets; while... Figure 6 Even in the cloud-interference area indicated by the d-line icon, EpM-UNet maintained stable segmentation performance. In summary, the EpM-UNet network demonstrated superior seaweed segmentation capabilities in complex environments, including continuous, dispersed, diffuse, and cloud-covered conditions.
[0165] Visual comparison and evaluation results show that the EpM-UNet network outperforms the other five algorithms in detecting different types of *Ulva prolifera* in the FIO-EP dataset. Particularly for small and scattered *Ulva prolifera* patches, the EpM-UNet network's detection results are significantly better than the other five algorithms. Furthermore, the EpM-UNet model can accurately detect small, blurry, and diffusely distributed *Ulva prolifera* regions in different types of *HSRIs*.
[0166] 3.3.3) Ablation experiments of EpM-UNet;
[0167] To quantify the contribution of each module in the EpM-UNet network, systematic ablation experiments were conducted on the FIO-EP dataset using the EpM-UNet network after removing the Local-enhance branch, the GEDM module, and the Adaptive Ep-SS2D module, as shown in Table 2. Figure 7 As shown;
[0168] Table 2. Quantitative evaluation of the EpM-UNet framework;
[0169]
[0170] In the table, "No local-enhance" refers to the EpM-UNet network after removing the Local-enhance branch.
[0171] No GEDM refers to the EpM-UNet network after removing the GEDM module;
[0172] No Adaptive Ep-SS2D refers to the EpM-UNet network after removing the Adaptive Ep-SS2D module;
[0173] First, the introduction of the Local-enhance branch significantly improves the model's ability to capture subtle textures and details. Experiments show that this branch improves the IOU by 0.99% and the F1-score by 0.59%. Figure 7 The comparison of the selected regions clearly demonstrates the importance of local feature enhancement for fine-grained segmentation tasks. Secondly, the GEDM module, through effective fusion of global and local features, strengthens the model's response to small, scattered *Ulva prolifera* patches. After adding this module, the IOU improved by 0.66%, and the F1-score improved by 0.39%. Figure 7 The corresponding regions validated the crucial role of gradient extraction and small target enhancement mechanisms in capturing sparse targets. Furthermore, the AdaptiveEpSS2D module focuses on adaptive extraction of textures in different directions, further improving the model's recognition of directional *Ulva prolifera* features. Ablation results show that this module improved IOU by 0.25% and F1-score by 0.16%. Figure 7 The selected region in the figure demonstrates the effectiveness of the adaptive mechanism for directional features in the segmentation task.
[0174] When the three modules are integrated, EpM-UNet achieves optimal performance on the FIOEP dataset, fully demonstrating the synergistic gains among the components.
[0175] Example 2
[0176] Based on Example 1, this example provides a deep learning-based remote sensing image detection system for *Ulva prolifera*, including:
[0177] The GEDM module is configured to perform edge gradient extraction and small target enhancement processing on the acquired remote sensing image to be detected, and obtain the first feature map after edge enhancement and small target feature enhancement.
[0178] Detection module: It is configured to transmit the first feature map to a U-shaped backbone network composed of cascaded multi-level Ep-VSS block modules, perform multi-level feature extraction, and obtain the detection result of the Ulva prolifera remote sensing image based on the feature map output by the last-level Ep-VSS block module;
[0179] Each Ep-VSS block module includes two processing branches. One branch extracts multi-scale texture features from the input feature map based on a deep separable convolutional network. The other branch models the spatial context information of the medium-to-long-range dependencies of the input feature map through a state-space attention mechanism. The outputs of the two branches are superimposed on the input feature map of the current Ep-VSS block module to obtain the output feature map of the current Ep-VSS block module, which is then input to the next Ep-VSS block module for further feature extraction.
[0180] Furthermore, the GEDM module includes a gradient feature extraction submodule and a small target enhancement submodule;
[0181] The gradient feature extraction submodule includes two parallel convolutional layers ConvX and ConvY with fixed convolutional kernels connected in sequence, as well as the first fusion module Concat;
[0182] In the gradient feature extraction submodule, two parallel convolutional layers with fixed kernels, ConvX and ConvY, calculate the image gradient through convolution operations in the vertical and horizontal directions, respectively; the first fusion module, Concat, fuses the obtained image gradient into the original remote sensing image to be detected.
[0183] The small target enhancement submodule includes multiple parallel multi-scale dilated convolutional layers, a second fusion module, and a first attention module;
[0184] Multi-scale dilated convolutional layers perform dilated convolution operations on the input features (output features of the gradient feature extraction submodule) with different dilation rates to obtain different scales to capture texture and edge information at different scales, and are activated by the ReLU function;
[0185] The second fusion module is configured to concatenate the features activated by the ReLU function along the channel dimension to obtain a multi-scale feature map.
[0186] The first attention module performs weighted fusion of the obtained multi-scale feature maps through an attention mechanism, and then outputs the result after convolution.
[0187] Furthermore, the Ep-VSS block module includes:
[0188] The first normalization module is configured to preprocess the input feature map of the input Ep-VSS block module using layer normalization;
[0189] The local feature enhancement branch is configured to extract multi-scale texture features from the input image based on a depthwise separable convolutional network. It includes one or more local feature enhancement units, each of which includes a first depthwise separable convolutional network, a first batch normalization, a point convolution, a second batch normalization, and a SiLU activation module connected in sequence. The preprocessed features are subjected to depthwise convolution, batch normalization, point convolution, batch normalization, and activation operations to obtain multi-scale texture features.
[0190] The remote modeling branch is used to model long-distance dependent spatial context information in the image; it includes a first linear layer, a SiLU activation layer, a second depthwise separable convolutional network, an adaptive Ulva state space two-dimensional scanning module, a second layer normalization module, and a bypass branch connected in sequence.
[0191] The third fusion module is used to perform weighted fusion of the two branch outputs and the input feature maps of the Ep-VSS block module to obtain the output feature map of the Ep-VSS block module;
[0192] Alternatively, the adaptive Ulva prolifera state space 2D scanning module is configured to perform the following process:
[0193] For the input two-dimensional feature map, multiple scans are performed in multiple predetermined directions to flatten it, resulting in multiple scan-expanded one-dimensional sequences.
[0194] The one-dimensional sequence in each direction after the scan is expanded is modeled using a state-space model. Based on the current observation and the hidden state information of the previous time step, the output sequence at each time step is obtained.
[0195] The obtained output sequences are fused into a unified global feature map using a scanning merging mechanism.
[0196] It should be noted that each module in this embodiment corresponds one-to-one with each step in embodiment 1, and their specific implementation processes are the same. Each module corresponds to each module of the constructed EpM-UNet network, which will not be repeated here.
[0197] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
[0198] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A method for detecting *Ulva prolifera* in remote sensing images based on deep learning, characterized in that... Includes the following steps: For the acquired remote sensing image to be detected, edge gradient extraction and small target enhancement processing are performed to obtain the first feature map after edge enhancement and small target feature enhancement. The first feature map is transmitted to a U-shaped backbone network composed of cascaded multi-level Ep-VSS block modules for multi-level feature extraction. The detection result of the Ulva prolifera remote sensing image is obtained based on the feature map output by the last-level Ep-VSS block module. Each Ep-VSS block module includes two processing branches. One branch extracts multi-scale texture features from the input feature map based on a deep separable convolutional network. The other branch models the spatial context information of the medium- and long-range dependencies of the input feature map through a state-space attention mechanism. The outputs of the two branches are superimposed on the input feature map of the current Ep-VSS block module to obtain the output feature map of the current Ep-VSS block module, which is then input into the next Ep-VSS block module for further feature extraction. The Ep-VSS block module includes: The first normalization module is configured to preprocess the input feature map of the input Ep-VSS block module using layer normalization; The local feature enhancement branch is configured to extract multi-scale texture features from the input image based on a depthwise separable convolutional network. It includes one or more local feature enhancement units, each of which includes a first depthwise separable convolutional network, a first batch normalization, a point convolution, a second batch normalization, and a SiLU activation module connected in sequence. The preprocessed features are subjected to depthwise convolution, batch normalization, point convolution, batch normalization, and activation operations to obtain multi-scale texture features. The remote modeling branch is used to model long-distance dependent spatial context information in the image; it includes a first linear layer, a SiLU activation layer, a second depthwise separable convolutional network, an adaptive Ulva state space two-dimensional scanning module, a second layer normalization module, and a bypass branch connected in sequence. The third fusion module is used to perform weighted fusion of the output of the local feature enhancement branch and the remote modeling branch with the input feature map of the Ep-VSS block module to obtain the output feature map of the Ep-VSS block module; The adaptive Ulva prolifera state space two-dimensional scanning module is configured to perform the following process: For the input two-dimensional feature map, multiple scans are performed in multiple predetermined directions to flatten it, resulting in multiple scan-expanded one-dimensional sequences. The one-dimensional sequence in each direction after the scan is expanded is modeled using a state-space model. Based on the current observation and the hidden state information of the previous time step, the output sequence at each time step is obtained. The obtained output sequences are fused into a unified global feature map using a scan merging mechanism, including: Global average pooling is performed on the output sequence in each direction to extract the global feature vector of Ulva prolifera in each direction; The global feature vectors in each direction are normalized to generate a set of adaptive directional weights; The output sequences in each direction are weighted and summed according to the corresponding adaptive direction weights to form a comprehensive global feature representation.
2. The method for detecting *Ulva prolifera* in remote sensing images based on deep learning as described in claim 1, characterized in that, The method for edge gradient extraction includes the following steps: A convolutional layer with a fixed kernel is used to calculate the image gradient features of the input remote sensing image to be detected through convolution operations in the vertical and horizontal directions, respectively. The obtained image gradient features are fused into the original remote sensing image to be detected to obtain an edge-enhanced feature map.
3. The method for detecting *Ulva prolifera* in remote sensing images based on deep learning as described in claim 1, characterized in that... The method for enhancing small targets includes the following steps: The edge-enhanced feature map is used to extract multi-scale features using multi-dilation rate convolution at different resolutions to obtain texture and edge information at different scales, and then activated by the ReLU function. The features activated by the ReLU function are concatenated along the channel dimension to obtain a multi-scale feature map. The obtained multi-scale feature maps are weighted and fused using an attention mechanism, and then the first feature map is output after a convolution operation.
4. The method for detecting *Ulva prolifera* in remote sensing images based on deep learning as described in claim 1, characterized in that: The scanning in the set direction includes horizontal scanning, vertical scanning, diagonal scanning, and anti-diagonal scanning.
5. A deep learning-based remote sensing image detection system for *Ulva prolifera*, characterized in that, include: The GEDM module is configured to perform edge gradient extraction and small target enhancement processing on the acquired remote sensing image to be detected, and obtain the first feature map after edge enhancement and small target feature enhancement. Detection module: It is configured to transmit the first feature map to a U-shaped backbone network composed of cascaded multi-level Ep-VSS block modules, perform multi-level feature extraction, and obtain the detection result of the Ulva prolifera remote sensing image based on the feature map output by the last-level Ep-VSS block module; Each Ep-VSS block module includes two processing branches. One branch extracts multi-scale texture features from the input feature map based on a deep separable convolutional network. The other branch models the spatial context information of the medium- and long-range dependencies of the input feature map through a state-space attention mechanism. The outputs of the two branches are superimposed on the input feature map of the current Ep-VSS block module to obtain the output feature map of the current Ep-VSS block module, which is then input into the next Ep-VSS block module for further feature extraction. The Ep-VSS block module includes: The first normalization module is configured to preprocess the input feature map of the input Ep-VSS block module using layer normalization; The local feature enhancement branch is configured to extract multi-scale texture features from the input image based on a depthwise separable convolutional network. It includes one or more local feature enhancement units, each of which includes a first depthwise separable convolutional network, a first batch normalization, a point convolution, a second batch normalization, and a SiLU activation module connected in sequence. The preprocessed features are subjected to depthwise convolution, batch normalization, point convolution, batch normalization, and activation operations to obtain multi-scale texture features. The remote modeling branch is used to model long-distance dependent spatial context information in the image; it includes a first linear layer, a SiLU activation layer, a second depthwise separable convolutional network, an adaptive Ulva state space two-dimensional scanning module, a second layer normalization module, and a bypass branch connected in sequence. The third fusion module is used to perform weighted fusion of the output of the local feature enhancement branch and the remote modeling branch with the input feature map of the Ep-VSS block module to obtain the output feature map of the Ep-VSS block module; The adaptive Ulva prolifera state space two-dimensional scanning module is configured to perform the following process: For the input two-dimensional feature map, multiple scans are performed in multiple predetermined directions to flatten it, resulting in multiple scan-expanded one-dimensional sequences. The one-dimensional sequence in each direction after the scan is expanded is modeled using a state-space model. Based on the current observation and the hidden state information of the previous time step, the output sequence at each time step is obtained. The obtained output sequences are fused into a unified global feature map using a scan merging mechanism, including: Global average pooling is performed on the output sequence in each direction to extract the global feature vector of Ulva prolifera in each direction; The global feature vectors in each direction are normalized to generate a set of adaptive directional weights; The output sequences in each direction are weighted and summed according to the corresponding adaptive direction weights to form a comprehensive global feature representation.
6. The deep learning-based remote sensing image detection system for *Ulva prolifera* as described in claim 5, characterized in that: The GEDM module includes a gradient feature extraction submodule and a small target enhancement submodule; The gradient feature extraction submodule includes two parallel convolutional layers ConvX and ConvY with fixed convolutional kernels connected in sequence, as well as the first fusion module Concat; In the gradient feature extraction submodule, two parallel convolutional layers with fixed kernels, ConvX and ConvY, calculate the image gradient through convolution operations in the vertical and horizontal directions, respectively; the first fusion module, Concat, fuses the obtained image gradient into the original remote sensing image to be detected. The small target enhancement submodule includes multiple parallel multi-scale dilated convolutional layers, a second fusion module, and a first attention module; Multi-scale dilated convolutional layers perform dilated convolution operations on the input features with different dilation rates to obtain different scales, thereby capturing texture and edge information at different scales, and are activated by the ReLU function; The second fusion module is configured to concatenate the features activated by the ReLU function along the channel dimension to obtain a multi-scale feature map. The first attention module performs weighted fusion of the obtained multi-scale feature maps through an attention mechanism, and then outputs the result after convolution.
Citation Information
Patent Citations
Improved YOLOv4 plasmodium detection method based on enhanced feature fusion
CN117854111A
MR image segmentation method for fetus cerebellar earthworm part
CN119762509A