Multi-modal remote sensing image real-time water body segmentation method and system adaptive to unmanned aerial vehicle
By extracting spatial and spectral features from remote sensing images and fusing them with specific features, generating dedicated weights for feature fusion, the problem of low segmentation accuracy of water areas in remote sensing images is solved, and accurate segmentation under complex conditions is achieved.
Patent Information
- Application Number
- CN202510612531.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-10-03
AI Technical Summary
In the existing technology, the water area segmentation method of remote sensing images has low segmentation accuracy in multi-source image applications, especially under conditions of lighting changes, water quality differences or complex background interference, there are problems such as blurred edges and missed water areas.
A real-time water body segmentation method based on multimodal remote sensing images is adopted. The spatial features and spectral features are extracted separately, and shared fusion weights are generated for feature fusion. The water body segmentation prediction model is used to achieve accurate segmentation of water body areas.
It achieves accurate segmentation of water areas in remote sensing images under complex backgrounds and diverse water types, improves segmentation accuracy and robustness, and maintains the real-time performance and computational efficiency of the model.
Smart Images

Figure CN120747775A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing data processing, and in particular to a method and system for real-time water body segmentation of multimodal remote sensing images adapted to unmanned aerial vehicles. Background Art
[0002] With the improvement of remote sensing image resolution, water body extraction based on remote sensing images collected by drones has become a key task in applications such as ecological environment monitoring and water resources management.
[0003] In the existing technology, one type of method constructs water body indices based on fixed bands, such as NDWI and MNDWI, and identifies water areas by calculating the reflectivity differences between different bands. This has the advantage of simple calculations, but the segmentation accuracy is highly sensitive to image acquisition conditions and atmospheric correction effects, making it difficult to adapt to the widespread application needs of multi-source remote sensing images. Another type of method proposes lightweight segmentation models, but such models often lose some spatial resolution capabilities due to structural compression. In particular, under conditions of changing lighting, different water quality, or complex background interference, there are problems such as blurred edges and missed water areas. Therefore, all of the above existing solutions suffer from the problem of low accuracy in the segmentation results of water areas in remote sensing images. Summary of the Invention
[0004] The present invention provides a method, system, electronic device and storage medium for real-time water segmentation of multimodal remote sensing images adapted to drones, which are used to address the defects in the existing technology and achieve accurate segmentation of water areas in remote sensing images.
[0005] The present invention provides a real-time water body segmentation method for multimodal remote sensing images adapted to unmanned aerial vehicles, comprising the following steps: Determine spatial and spectral characteristics based on the remote sensing image to be processed; generating a shared fusion weight according to the spatial feature and the spectral feature, and performing feature fusion on the spatial feature and the spectral feature according to the shared fusion weight to obtain a fusion feature; The fusion feature is input into a water body segmentation prediction model to obtain a water body segmentation prediction result output by the water body segmentation prediction model; wherein the water body segmentation prediction model is trained based on the fusion feature samples and the true pixel labels.
[0006] According to a real-time water segmentation method for multimodal remote sensing images adapted to drones provided by the present invention, generating shared fusion weights based on the spatial features and the spectral features includes: Performing spatial attention calculations on the spatial features and the spectral features respectively to generate a spatial feature attention map and a spectral feature attention map respectively; Splicing the spatial feature attention map and the spectral feature attention map in the channel dimension to form a spliced feature map containing spatial information; A convolution operation is performed on the concatenated feature map and activated by Sigmoid to obtain the shared fusion weight.
[0007] According to a real-time water body segmentation method for multimodal remote sensing images adapted to drones provided by the present invention, the feature fusion of the spatial features and the spectral features according to the shared fusion weights to obtain fused features includes: The fusion feature is obtained by the following formula: ; Where, is the fusion feature, is the spatial feature, is the spectral feature, is the shared fusion weight.
[0008] According to the present invention, a real-time water body segmentation method for multimodal remote sensing images adapted to UAVs is provided, which also includes a training method for the water body segmentation prediction model: Inputting the fused feature samples into the main decoder and the self-distillation decoder of the initial model respectively, obtaining the main decoding branch feature map and the first prediction result output by the main decoder, and the self-distillation branch feature map and the second prediction result output by the self-distillation decoder; Binarizing the first prediction result and extracting the predicted edge to obtain a dynamic edge label; Constructing a feature consistency loss function according to the Euclidean distance between the main decoding branch feature map and the self-distillation branch feature map; constructing an edge loss function according to a difference between the second prediction result and the dynamic edge label; Constructing a subject loss function based on the difference between the first prediction result and the true label of the pixel; Summing the feature consistency loss function, the edge loss function, and the main loss function to obtain a total loss function, and optimizing the network parameters of the initial model according to the total loss function; After the training is completed, the self-distillation decoder is removed and the main decoder is retained to form the water body segmentation prediction model.
[0009] According to a real-time water segmentation method for multimodal remote sensing images adapted to drones provided by the present invention, constructing a feature consistency loss function based on the Euclidean distance between the main decoding branch feature map and the self-distillation branch feature map includes: The feature consistency loss function is obtained by the following formula: ; Where, is the feature consistency loss function, i is the current upsampling layer number, k is the total upsampling layer number, is linear interpolation upsampling, is the self-distillation branch characteristic diagram, represents the binary activation process, F is the main decoding branch feature map, Represents the XOR process, Representation and process, Represents the Euclidean distance metric constraint.
[0010] According to a real-time water segmentation method for multimodal remote sensing images adapted to drones provided by the present invention, constructing an edge loss function based on the difference between the second prediction result and the dynamic edge label includes: The marginal loss function is obtained by the following formula: ; Where, is the marginal loss function, Pixel is the probability of the edge of the target area, Pixel Dynamic edge labels.
[0011] According to a real-time water body segmentation method for multimodal remote sensing images adapted to drones provided by the present invention, determining spatial features and spectral features based on the remote sensing images to be processed includes: Performing band selection on the remote sensing image to be processed to obtain a visible light image, and obtaining expert prior knowledge based on the preset band calculation; Inputting the visible light image into a first feature extraction network to obtain spatial features output by the first feature extraction network; wherein the first feature extraction network includes an initial convolutional layer and multiple first stages, each of the first stages includes a downsampling unit and multiple basic units; The expert prior knowledge is input into a second feature extraction network to obtain spectral features output by the second feature extraction network; wherein the second feature extraction network includes an initial convolutional layer and multiple second stages, each of the second stages includes a downsampling unit and at least one basic unit.
[0012] The present invention also provides a multi-modal remote sensing image real-time water segmentation system adapted to UAVs, comprising the following modules: A first processing module is used to determine spatial features and spectral features based on the remote sensing image to be processed; a second processing module, configured to generate a shared fusion weight according to the spatial feature and the spectral feature, and perform feature fusion on the spatial feature and the spectral feature according to the shared fusion weight to obtain a fusion feature; The third processing module is used to input the fusion feature into the water body segmentation prediction model to obtain the water body segmentation prediction result output by the water body segmentation prediction model; wherein, the water body segmentation prediction model is trained based on the fusion feature samples and the true pixel labels.
[0013] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, it implements the real-time water body segmentation method of multimodal remote sensing images adapted to drones as described in any of the above.
[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the real-time water body segmentation method of multimodal remote sensing images adapted to drones as described in any of the above is implemented.
[0015] The present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the above-described methods for real-time water segmentation of multimodal remote sensing images adapted to drones.
[0016] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: By introducing the joint extraction of spatial and spectral features at the initial stage of the segmentation process, the morphological distribution and spectral response information of water areas in remote sensing images are fully captured, effectively compensating for the insufficient expression of traditional single feature extraction methods when dealing with complex backgrounds and diverse water types, and achieving the ability to both understand the whole picture and perceive local details. By constructing a shared fusion weight mechanism and performing weighted fusion of spatial and spectral features based on this weight, dynamic fusion of heterogeneous feature information is achieved, effectively highlighting important areas related to water targets and suppressing the interference of non-water areas in the feature expression process. By inputting the fused features into the water segmentation prediction model, the water segmentation prediction results are obtained, thereby achieving accurate segmentation of water areas in remote sensing images. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 This is one of the flow charts of the real-time water body segmentation method for multimodal remote sensing images adapted to drones provided by the present invention.
[0019] Figure 2 This is the second flow chart of the real-time water body segmentation method of multimodal remote sensing images adapted to drones provided by the present invention.
[0020] Figure 3 This is the third flow chart of the real-time water body segmentation method for multimodal remote sensing images adapted to drones provided by the present invention.
[0021] Figure 4 This is the fourth flow chart of the real-time water body segmentation method for multimodal remote sensing images adapted to drones provided by the present invention.
[0022] Figure 5 It is a structural diagram of the real-time water segmentation system of multimodal remote sensing images adapted to drones provided by the present invention.
[0023] Figure 6 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0024] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by ordinary technicians in this field without making creative efforts based on the embodiments of the present invention are within the scope of protection of the present invention.
[0025] It should be noted that, in the description of the present invention, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. The orientation or positional relationship indicated by the terms "upper" and "lower" is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the system or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to the specific circumstances.
[0026] The terms "first," "second," and so forth, used herein are used to distinguish similar objects, not to describe a specific order or precedence. It should be understood that such terms are interchangeable where appropriate, allowing embodiments of the present invention to be implemented in an order other than that illustrated or described herein. Furthermore, the terms "first," "second," and so forth generally distinguish objects of a single type, and do not limit the number of objects. For example, the first object may be one or more. Furthermore, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the connected objects.
[0027] The following combination Figures 1-6 The present invention describes a method, system, electronic device and storage medium for real-time water segmentation of multimodal remote sensing images adapted to drones.
[0028] Figure 1 This is one of the flow charts of the real-time water body segmentation method for multimodal remote sensing images adapted to UAVs provided by the present invention, such as Figure 1 As shown, including but not limited to the following steps: Step 101: Determine spatial features and spectral features based on the remote sensing image to be processed.
[0029] In this embodiment, step 101 is intended to lay a reliable feature foundation for subsequent water body segmentation prediction. The existing water body index method relies on precise atmospheric correction and is difficult to use in the real-time environment of drones. Although the general depth model has considerable accuracy, it cannot run in real time on drones due to the large number of parameters, and its generalization ability is insufficient when faced with complex changes such as illumination and water quality. In order to solve the core technical problem of "it is difficult to strike a balance between real-time and robustness", this solution first splits the remote sensing image into two information threads: spatial features and spectral features, and extracts two types of complementary features respectively through a lightweight twin network, which not only avoids the computational bottleneck of traditional methods, but also alleviates the dilemma of lightweight models' poor adaptability to complex scenes.
[0030] In one possible implementation, Figure 2 This is the second flow chart of the real-time water body segmentation method of multimodal remote sensing images adapted to UAVs provided by the present invention, such as Figure 2 As shown, step 101 specifically includes steps 201-203: Step 201: performing band selection on the remote sensing image to be processed to obtain a visible light image, and obtaining expert prior knowledge based on the preset band calculation.
[0031] Step 202: Input the visible light image into a first feature extraction network to obtain spatial features output by the first feature extraction network; wherein the first feature extraction network includes an initial convolutional layer and multiple first stages, each first stage includes a downsampling unit and multiple basic units.
[0032] Step 203: Input the expert prior knowledge into the second feature extraction network to obtain the spectral features output by the second feature extraction network; wherein the second feature extraction network includes an initial convolution layer and multiple second stages, each second stage includes a downsampling unit and at least one basic unit.
[0033] In this embodiment, steps 201-203 are executed around "band selection and dual-branch lightweight feature extraction," with the overarching goal of simultaneously capturing the morphological texture and spectral differences of water bodies with minimal computational effort. Traditional water indices, while easy to implement, rely on precise atmospheric corrections. While general deep models can automatically learn multi-scale representations, their large number of parameters makes them difficult to run in real time on drones. Therefore, this solution first constructs "expert prior knowledge" using band combinations, then extracts spatial and spectral features separately through an asymmetric twin network, fundamentally enhancing the lightweight model's adaptability to complex scenarios.
[0034] First, in step 201, the system performs band selection on the remote sensing image to be processed, retaining only visible light bands such as green and blue to generate a visible light image. Simultaneously, expert prior knowledge is calculated directly within the DN range according to a preset band combination formula (such as NDWI, MNDWI, or AWEI). This approach not only avoids atmospheric correction, which is difficult to perform in real-time drone processing, but also explicitly injects numerical values of the reflectance differences between water, vegetation, and soil in the near-infrared or short-wave infrared into the subsequent network, effectively reducing interference from non-water pixels.
[0035] Step 202 uses the visible light image as input and extracts spatial features using the first feature extraction network in the asymmetric lightweight twin architecture. This first feature extraction network first performs low-level texture capture and channel expansion through an initial 3×3 convolutional layer, then enters three progressive stages. Each stage begins with a downsampling unit using depthwise separable convolutions with a stride of 2, gradually halving the resolution and doubling the number of channels. The downsampling results are then subjected to a "channel splitting-independent convolution-channel rearrangement" process using multiple basic units, fully refining local and global morphological textures without increasing resolution. (The preferred number of basic units in the three stages of the first feature extraction network in this embodiment is 3, 5, and 3, respectively.) The outputs of these three stages, together with the first-level pooling results, form a multi-scale spatial feature pyramid. This pyramid preserves the fine boundaries of the shoreline while maintaining the global consistency of the overall shape of the water body. Channel reuse keeps the number of parameters and computational overhead extremely low. This design enables the model to perform inference at real-time frame rates on edge platforms and significantly improves spatial discrimination capabilities in complex terrain.
[0036] Step 203 uses the expert prior knowledge as input to extract spectral features through a second feature extraction network. Considering the small channel dimension and information sparseness of the exponential graph, the overall architecture of this second feature extraction network follows the "initial convolutional layer + multi-stage" framework of the first network. However, the number of basic units in each stage is appropriately reduced (the preferred number of basic units in the three stages of the second feature extraction network in this embodiment is 1, 3, and 1, respectively, which represents two fewer basic units per stage compared to the first feature extraction network), maintaining the same feature scale output as the spatial branch. Each stage still contains one downsampling unit and at least one basic unit. This ensures that the spectral features are strictly aligned with the spatial features in the spatial dimension, facilitating subsequent weight sharing and fusion. Furthermore, by reducing the convolution depth, the number of parameters is further reduced, achieving the goal of a lightweight spectral branch with minimal overall model overhead. The spectral features obtained through this processing accurately depict the regions of high response to the difference between visible and near-infrared reflectance in water. When combined with the spatial features, they effectively compensate for the shortcomings of morphological features under complex lighting and water quality conditions, thereby improving the overall robustness and generalization of the model.
[0037] Step 102: Generate a shared fusion weight according to the spatial features and the spectral features, and perform feature fusion on the spatial features and the spectral features according to the shared fusion weight to obtain a fusion feature.
[0038] After completing the parallel extraction of spatial and spectral features, the core purpose of step 102 is to organically integrate the complementary information contained in the two feature contexts at the pixel level through a lightweight and dynamic weighting mechanism. Relying solely on simple addition or splicing will lead to the accumulation of redundant information, increasing the computational complexity and weakening the ability to discriminate water areas. Therefore, this embodiment first calculates spatial attention for each of the two feature types, then uses a single convolution to generate a shared fusion weight. Finally, the two feature paths are simultaneously weighted according to a unified weight, significantly improving segmentation accuracy and robustness while maintaining real-time performance.
[0039] In one possible implementation, Figure 3 This is the third flow chart of the real-time water body segmentation method for multimodal remote sensing images adapted to UAVs provided by the present invention, such as Figure 3 As shown, in step 102, a specific implementation method for generating shared fusion weights according to spatial features and spectral features includes steps 301-303: Step 301: perform spatial attention calculation on the spatial features and spectral features respectively, and generate a spatial feature attention map and a spectral feature attention map respectively.
[0040] Step 302: The spatial feature attention map and the spectral feature attention map are spliced in the channel dimension to form a spliced feature map containing spatial information.
[0041] Step 303: Perform convolution operation on the concatenated feature map and activate it through Sigmoid to obtain shared fusion weights.
[0042] In steps 301-303, this embodiment aims to maintain the lightweight of the network while performing fine-grained, position-sensitive weight fusion of visible light spatial features and spectral features to overcome redundant interference caused by the heterogeneity of the two features.
[0043] To this end, first, in step 301, spatial attention calculation is performed on the spatial features and spectral features respectively to highlight the important areas of the spatial position in each feature map, and a spatial feature attention map and a spectral feature attention map are obtained.
[0044] Specifically, for spatial features or spectral features, the input feature map is aggregated along the channel dimension through global average pooling and global maximum pooling operations respectively to form two single-channel intermediate feature maps. The two single-channel feature maps are then spliced in the channel dimension and the spatial information is fused through convolution operation. The convolution result is then mapped to the range of [0,1] using the Sigmoid activation function to form a spatial feature attention map. and spectral feature attention map This “mean-extreme” complementary statistics can simultaneously retain both salient areas and detailed texture information, thereby highlighting the response of the target water area without additional channel overhead.
[0045] The computation of spatial attention is expressed as: in, and Represents the feature map after maximum pooling and average pooling, , Represents the feature map after the maximum pooling and spatial pooling results are spliced in the channel dimension, , represents the spatial attention weight map, represents the Sigmoid activation function, .
[0046] Next, in step 302, in order to further improve the fusion feature expression ability, the spatial feature attention map and spectral feature attention map Splicing in the channel dimension to form a splicing feature map containing spatial information , Unlike directly adding or multiplying feature maps, the concatenation operation fully preserves the differences between the two attention paths, and then unifies the scale through convolution, laying the foundation for the subsequent generation of shared weights.
[0047] Then, in step 303, the concatenated feature map is further convolved and processed through a Sigmoid activation function to obtain a shared fusion weight. Specifically, the concatenated feature map is convolved to fuse the two spatial attention information to produce a comprehensive expression; the Sigmoid activation function is then used to generate a shared fusion weight map in the range [0, 1], where the weight value represents the importance of the fused feature at the corresponding spatial position. The specific formula for calculating the shared fusion weight map is as follows: Where, is the shared fusion weight, is the splicing feature map, is a 1×1 convolution operation, is the Sigmoid activation function.
[0048] In a possible implementation, in step 102, after obtaining the fusion weight, the fusion feature is further obtained by the following formula: ; Where, To fusion features, is the spatial feature, is the spectral characteristic, is the shared fusion weight.
[0049] Specifically, through the above-mentioned fusion method, spatial features and spectral features can be dynamically fused according to the importance of spatial positions, significantly improving the recognition accuracy and generalization of water areas in remote sensing images, while effectively reducing the interference of background and other landforms in complex environments, and ensuring the reliability and stability of the model in actual complex scenes.
[0050] Step 103: Input the fused features into the water body segmentation prediction model to obtain the water body segmentation prediction result output by the water body segmentation prediction model; wherein the water body segmentation prediction model is trained based on the fused feature samples and the true pixel labels.
[0051] Step 103 of this embodiment aims to achieve accurate segmentation of water areas using the fused features obtained in the previous steps. Because water boundaries in remote sensing images are complex and changeable, and are easily confused with surrounding objects, it is necessary to train a water segmentation prediction model to accurately identify water areas from the fused features and achieve effective water segmentation.
[0052] Specifically, during model training, fused feature samples are fed into the main decoder and self-distillation decoder of the water segmentation prediction model. The main decoder decodes the fused features, generating a probability map for the target class through a 1×1 convolution. This probability map is then upsampled to the same resolution as the input remote sensing image using bilinear interpolation, resulting in the main decoder branch feature map and the first prediction result. The prediction result of the main decoder branch is used to determine whether a pixel belongs to the water target area, focusing on the overall distribution of the target area.
[0053] The self-distillation decoder, on the other hand, is used to enhance the model's ability to identify water area edges. Its structure is similar to the main decoder, but it further incorporates a multi-scale contextual feature extraction module to capture richer boundary region information, resulting in self-distillation branch feature maps and a second prediction result. This module includes convolution operations with various dilation rates (e.g., 6, 12, and 18) and a global average pooling layer. Smaller dilation rates capture local details near edges, while larger dilation rates capture broader spatial structure. The feature maps processed by this module are concatenated and then reduced in dimension using a 1×1 convolution, forming self-distillation branch feature maps that contain rich edge information.
[0054] To further improve the model's segmentation performance, the self-distillation branch is trained collaboratively with the main branch. The main loss function is calculated by comparing the difference between the first prediction of the main decoding branch and the true pixel labels to optimize the model's overall performance. Simultaneously, the first prediction of the main branch is binarized and its predicted edges are extracted to obtain dynamic edge labels. The dynamic edge labels are then used together with the second prediction of the self-distillation branch to calculate the edge loss function, dynamically guiding the model to gradually optimize its ability to recognize edge regions.
[0055] Furthermore, to ensure feature consistency between the main branch and the self-distillation branch, this embodiment constructs a feature consistency loss function based on the Euclidean distance between the main decoding branch feature map and the self-distillation branch feature map. By constraining this feature consistency loss, the self-distillation branch can effectively guide the main branch features to gradually focus on the boundary details of the target area during the training phase.
[0056] Finally, the main loss function, edge loss function, and feature consistency loss function are summed to form a total loss function, and the model network parameters are updated by optimizing the total loss function. After model training is completed, the self-distillation decoder is removed, and only the main decoder is retained to form a water segmentation prediction model for subsequent practical applications. In actual applications, this embodiment inputs the fused features obtained in step 102 into the main decoder of the water segmentation prediction model to obtain water segmentation prediction results.
[0057] The following examples will explain in detail the training method of the water body segmentation prediction model. Figure 4 This is the fourth flow chart of the method for real-time water body segmentation of multimodal remote sensing images adapted to UAVs provided by the present invention, such as Figure 4 As shown, the training method specifically includes steps 401-407: Step 401: Input the fused feature samples into the main decoder and self-distillation decoder of the initial model respectively, and obtain the main decoding branch feature map and the first prediction result output by the main decoder, and the self-distillation branch feature map and the second prediction result output by the self-distillation decoder.
[0058] Step 401 of this embodiment aims to achieve efficient feature extraction of the water target area and its edge details in the fusion feature through the coordinated optimization of the main decoder and the self-distillation decoder of the water body segmentation prediction model, so as to improve the model's ability to accurately segment water body areas in complex remote sensing images.
[0059] Since the water areas in remote sensing images are usually complexly intertwined with the surrounding environment, it is difficult to effectively capture the detailed information of the edges of the water areas by relying solely on the main decoder. Therefore, this embodiment sets up a self-distillation decoder to assist the main decoder to focus on edge features more accurately.
[0060] Specifically, this embodiment first inputs the fused feature samples into the main decoder and self-distillation decoder of the initial model. The main decoder can be a U-Net decoder or a SegNet decoder, while the self-distillation decoder can be a U-Net decoder, an ASPP decoder, or a shallow FCN-based decoder. The main decoder is used to generate a holistic water area prediction. In practice, the main decoder decodes the fused features through a 1×1 convolution to obtain a main decoder branch feature map. The main decoder branch feature map is then upsampled using bilinear interpolation to restore it to a resolution that matches the original input image size, generating the first prediction result, i.e., the preliminary prediction result for the entire water area.
[0061] Meanwhile, the self-distillation decoder focuses on extracting detailed information related to the edges of water bodies from the fused features. To this end, the self-distillation decoder includes a multi-scale contextual feature extraction module. This module uses convolutional layers with various dilation rates (e.g., 6, 12, and 18) and global average pooling layers to capture local texture features at the edge of the fused features and broader background structural information. This module then performs dimensionality reduction through feature concatenation and 1×1 convolution to produce self-distilled branch feature maps rich in edge details. The self-distilled branch feature maps are then similarly processed with 1×1 convolution and up-sampled using bilinear interpolation to generate the second prediction result, namely the prediction result for the edge position of the water body.
[0062] The main decoder and self-distillation decoder simultaneously process the same fused features, but each focuses on different levels of feature detail. The first prediction generated by the main decoder represents a global prediction of the entire region, while the second prediction generated by the self-distillation decoder represents a refined prediction of local edge details. This dual decoding structure effectively enables multi-angle analysis of the fused features, providing a more comprehensive and accurate feature representation.
[0063] Step 402: Binarize the first prediction result and extract the predicted edge to obtain a dynamic edge label.
[0064] Step 402 of this embodiment is intended to dynamically generate an edge supervision signal to guide the model to more effectively learn edge detail information of water areas in remote sensing images, thereby improving the boundary recognition capability and accuracy of the water body segmentation prediction model.
[0065] Specifically, the first prediction result generated by the main decoder usually reflects the spatial distribution information of the entire water body area, but the prediction precision of the target edge is insufficient. In order to further enhance the main decoder's perception ability of the water body edge area, this embodiment generates a dynamic edge label based on the first prediction result output by the main decoder as a supervision signal for the self-distillation decoder.
[0066] In the implementation, the first prediction output by the main decoder is first binarized to determine the clear boundaries of the predicted water area. Based on this binarization, an edge extraction algorithm is then used to extract the contour edges of the predicted area, generating a predicted edge image that reflects the current prediction level of the main decoder. This predicted edge image is then combined with the true pixel labels of the corresponding remote sensing image by performing a logical AND operation on the two. This operation retains only the boundary regions where the main branch's prediction confidence is high, resulting in the final dynamic edge labels.
[0067] The generation formula of dynamic edge labels is expressed as: in, is a dynamic edge label; Represents pixel points The true label of Pixel The binarization result of Indicates the extraction of predicted binary results The symbol & indicates that the above feature map is processed using a logical AND operation to achieve effective generation of dynamic edge labels.
[0068] Compared to traditional static edge labels, the dynamic edge labels generated in this way better reflect the model's actual learning level and current segmentation prediction capabilities during training. Specifically, dynamic edge labels not only effectively highlight areas where the main decoder has correctly predicted with high confidence, but also selectively focus on edge details where the model failed to accurately predict. This gradual expansion of supervisory signals from simple to complex effectively reduces the uncertainty of edge region predictions during model training.
[0069] Step 403: Construct a feature consistency loss function based on the Euclidean distance between the main decoding branch feature map and the self-distillation branch feature map.
[0070] Step 403 of this embodiment aims to construct a feature consistency loss function to achieve consistency constraints on the feature representation capabilities between the main decoding branch and the self-distillation branch, thereby effectively guiding the main decoder to gradually learn and focus on the important features of the edge area during the training phase, and improving the final water body segmentation prediction model's ability to express edge structures.
[0071] In the real-time water segmentation task of multimodal remote sensing images adapted to drones, the main decoder is mainly used to complete the overall prediction of the water area, while the self-distillation decoder is dedicated to mining and strengthening the local structural information of the edge area. Due to the differences in the training objectives of the two, if there is a lack of an effective feature alignment mechanism, the main branch may not learn enough edge features, affecting the final model's segmentation effect on complex boundary areas. Therefore, this embodiment guides the main branch and the self-distillation branch to maintain consistency in the feature space by constructing a feature consistency loss function based on the Euclidean distance metric, so as to encourage the main branch to learn more significant features related to the edge.
[0072] Specifically, this embodiment first extracts the main decoding branch feature map Branch feature map of the self-distilled decoder , where i represents the number of layers in the current upsampling stage and k is the total number of upsampling layers. To ensure the spatial dimension alignment of the feature map, the self-distillation branch feature map Need to use linear interpolation upsampling function Restore to Same resolution.
[0073] Afterwards, respectively and Perform binary activation operations , to enhance the response strength of the target area, and perform a logical AND operation (&) between the two to obtain the difference mask between the two in the edge area. Finally, the Euclidean distance calculation is performed on the difference mask and the upsampled self-distilled branch feature map to measure the degree of difference between it and the main branch feature map, thereby constructing the feature consistency loss function .
[0074] In one possible implementation, the feature consistency loss function is obtained by the following formula: ; Where, is the feature consistency loss function, i is the current number of upsampling layers, k is the total number of upsampling layers, is linear interpolation upsampling, is the self-distillation branch feature diagram, represents the binary activation process, F is the main decoding branch feature map, Represents the XOR process, Representation and process, Represents the Euclidean distance metric constraint.
[0075] This loss function measures the consistency of the responses of the two decoders at the same spatial location, specifically focusing on predicting differences in responses in edge regions. This effectively enhances the self-distillation branch's ability to guide the main branch during training. By continuously optimizing this loss function, the main branch can gradually focus on high-confidence features extracted by the self-distillation branch in edge regions without significantly increasing computational complexity, thereby improving the overall model's ability to segment regions with blurred boundaries.
[0076] Step 404: construct an edge loss function based on the difference between the second prediction result and the dynamic edge label.
[0077] Step 404 of this embodiment aims to construct an edge loss function for optimizing the edge prediction capability of water body areas. By measuring the difference between the edge prediction results of the self-distillation decoder and the dynamic edge labels, the model is guided to more accurately learn the edge details of the water body areas in the remote sensing image, thereby effectively improving the edge positioning accuracy of the water body segmentation model.
[0078] Because water objects in remote sensing images often have diverse shapes and fuzzy boundaries, traditional segmentation models are prone to problems such as breakage and misjudgment in edge regions. This is especially true in scenes with complex lighting or high background noise, where edge regions become a bottleneck for model performance. To address this, this embodiment introduces an edge loss supervision mechanism based on dynamic edge labels during the training phase, enabling the model to gradually optimize its boundary recognition capabilities and implement a training strategy that progresses from easy to difficult and gradually learns.
[0079] Specifically, the self-distillation decoder generates a second prediction result, namely an edge prediction map, in step 401, which is used to characterize the predicted probability of each pixel being the edge of a water body. At the same time, in step 402, the dynamic edge label is extracted based on the first prediction result of the main decoder and the true label. , used to characterize whether each pixel belongs to a high confidence edge region. Based on the above two, the edge loss function is constructed. .
[0080] In one possible implementation, the edge loss function is obtained by the following formula: ; Where, is the marginal loss function, Pixel is the probability of the edge of the target area, Pixel Dynamic edge labels.
[0081] In actual training, smaller values of this loss function indicate a closer match between the predicted edge and the true edge, thus encouraging the self-distillation decoder to make more precise judgments about edge regions. It is worth noting that compared to the conventional binary cross-entropy loss function, the edge Dice loss used in this embodiment is more robust to sample imbalance, preventing edge pixels from being neglected in the overall prediction and increasing the model's focus on a small number of key edge regions.
[0082] Step 405: Construct a main loss function based on the difference between the first prediction result and the true label of the pixel.
[0083] Step 405 of this embodiment aims to improve the ability of the subject decoder in the water segmentation prediction model to identify the overall distribution of water areas by constructing a subject loss function. The subject decoder serves as the primary output channel for fusion features, and its prediction results are directly used to generate the final water segmentation map of the remote sensing image. Therefore, its training accuracy plays a decisive role in the overall performance of the model.
[0084] In the real-time water segmentation task of multimodal remote sensing imagery adapted for drones, the target area often has features such as blurred boundaries and irregular shapes. At the same time, the background area may contain non-water areas with similar reflective properties to water bodies. Therefore, a loss function with strong class discrimination capabilities and stable training is required to accurately distinguish between water and non-water areas. To this end, this embodiment uses the cross-entropy loss function as the main loss function, measuring the difference between the first prediction result output by the main decoder and the true pixel-level label.
[0085] In the specific implementation process, firstly, the first prediction result output by the main decoder in step 401 is The true label of the pixel at the corresponding position Perform pixel-by-pixel comparison and calculate the difference in target category probability distribution between the two. Indicates the probability that the model predicts a water body at the pixel point (x, y). Indicates whether the pixel point (x, y) is a real water pixel (1 indicates water, 0 indicates non-water). The calculation formula of the main loss function is as follows: This formula is a standard two-class cross-entropy loss function, which effectively measures the degree of match between the predicted probability distribution and the true label. By continuously optimizing this loss function, the model is guided to continuously improve the accuracy of distinguishing water and non-water pixels during training, thereby improving the regional integrity and semantic consistency of the overall segmentation map.
[0086] In addition, the cross entropy loss function is also highly stable when dealing with the problem of unbalanced category proportions. It is suitable for situations where the distribution ratios of "water bodies" and "non-water bodies" in remote sensing images are quite different, and helps to reduce the model's tendency to overfit the background area.
[0087] Step 406: summing the feature consistency loss function, the edge loss function, and the main loss function to obtain a total loss function, and optimizing the network parameters of the initial model based on the total loss function.
[0088] Step 406 of this embodiment aims to comprehensively consider the model's ability to identify the overall distribution of water areas and edge details, and to achieve global performance improvement of the water segmentation prediction model in remote sensing images by constructing and optimizing a total loss function. This step combines the feature consistency loss function constructed in step 403, the edge loss function constructed in step 404, and the main loss function constructed in step 405. This step uses multi-dimensional objectives to constrain the model training process and collaboratively optimize the parameters of the main decoder and self-distillation decoder.
[0089] Because the spatial morphology and edge details of water bodies in remote sensing images are often complex and uncertain, a single loss function cannot fully measure model performance. This embodiment integrates semantic guidance at different levels into the training objective by fusing three types of loss signals. This allows the model to focus on the integrity of water bodies while also having stronger boundary recognition and branch consistency expression capabilities, thus achieving a balance between accuracy and robustness.
[0090] Specifically, the total loss function is expressed by the following formula ; During training, each iteration uses the aforementioned total loss function to backpropagate gradients and update network parameters. Specifically, all learnable parameters, including the fusion feature extraction network, the main decoder, and the self-distillation decoder, are optimized. By introducing edge supervision and branch consistency constraints, the model not only achieves stronger boundary discrimination but also improves edge response stability and generalization, avoiding blurred or broken boundaries caused by insufficient training of the main branches.
[0091] This multi-loss joint optimization mechanism is highly adaptable and scalable, and the weighting ratios of various losses can be further adjusted according to task requirements to adapt to application scenarios with different segmentation accuracy and efficiency. In the real-time water segmentation task proposed in this paper, a model trained using the aforementioned total loss function and retaining only the main decoder during the inference phase significantly reduces computational overhead while achieving high-precision segmentation of complex water bodies.
[0092] Step 407: After the training is completed, the self-distillation decoder is removed and the main decoder is retained to form a water body segmentation prediction model.
[0093] Step 407 of this embodiment aims to complete the final deployment optimization of the water body segmentation prediction model. After training is completed, the self-distillation decoder is removed, and only the main decoder is retained as the core output channel of the inference stage. This significantly compresses the model structure without reducing the segmentation accuracy, thereby improving its inference efficiency and deployment adaptability in practical applications.
[0094] Throughout the training phase, the self-distillation decoder serves as an auxiliary branch, sharing the fused feature input with the main decoder. Through dynamic edge labels, edge loss functions, and feature consistency loss functions, it significantly guides and supplements the main decoder's edge details. However, the core value of this distillation mechanism lies in the "soft supervision" of the main branch during training. Once training is complete, the main decoder is capable of accurately extracting the key edge structure and spatial distribution of the fused features. The self-distillation decoder no longer participates in model inference tasks, but instead introduces additional computational resource overhead.
[0095] Therefore, after model training converges—that is, when the total loss function reaches a preset convergence condition or the maximum number of iterations—the model structure is streamlined. Specifically, the self-distillation decoder and its corresponding convolution, upsampling, and multi-scale context modules are removed from the model, leaving only the backbone network consisting of the fusion feature extraction network and the main decoder. The first prediction output by this structure serves as the final water segmentation prediction, making it suitable for online inference or edge computing deployment of remote sensing imagery in real-world scenarios.
[0096] After processing in the above manner, the resulting water body segmentation prediction model has the following technical advantages: on the one hand, by introducing an edge-guiding mechanism during the training phase, the accuracy of water body boundary detection and the continuity of segmentation contours are significantly improved; on the other hand, during the inference phase, it does not rely on auxiliary branches, and the overall model parameter quantity and computational complexity are effectively controlled, making it more suitable for resource-constrained UAV systems.
[0097] Reference Figure 5 , Figure 5 This is a structural diagram of the multimodal remote sensing image real-time water segmentation system adapted to UAVs provided by the present invention. The system includes: A first processing module is used to determine spatial features and spectral features based on the remote sensing image to be processed; The second processing module is used to generate a shared fusion weight according to the spatial features and the spectral features, and perform feature fusion on the spatial features and the spectral features according to the shared fusion weight to obtain a fusion feature; The third processing module is used to input the fusion feature into the water body segmentation prediction model to obtain the water body segmentation prediction result output by the water body segmentation prediction model; wherein the water body segmentation prediction model is trained based on the fusion feature samples and the true pixel labels.
[0098] In a possible implementation, the second processing module is further configured to: Perform spatial attention calculations on spatial features and spectral features respectively, and generate spatial feature attention maps and spectral feature attention maps respectively; The spatial feature attention map and the spectral feature attention map are spliced in the channel dimension to form a spliced feature map containing spatial information; The concatenated feature map is convolved and activated by Sigmoid to obtain the shared fusion weights.
[0099] In a possible implementation, the second processing module is further configured to: The fusion feature is obtained by the following formula: ; Where, To fusion features, is the spatial feature, is the spectral characteristic, is the shared fusion weight.
[0100] In one possible implementation, the system further includes a model training module for: Input the fused feature samples into the main decoder and self-distillation decoder of the initial model respectively, and obtain the main decoding branch feature map and the first prediction result output by the main decoder, and the self-distillation branch feature map and the second prediction result output by the self-distillation decoder; Binarize the first prediction result and extract the predicted edge to obtain a dynamic edge label; Construct a feature consistency loss function based on the Euclidean distance between the main decoding branch feature map and the self-distillation branch feature map; Constructing an edge loss function based on the difference between the second prediction result and the dynamic edge label; Construct a main loss function based on the difference between the first prediction result and the true label of the pixel; The feature consistency loss function, edge loss function and main loss function are summed to obtain the total loss function, and the network parameters of the initial model are optimized according to the total loss function; After training is completed, the self-distillation decoder is removed and the main decoder is retained to form a water body segmentation prediction model.
[0101] In one possible implementation, the model training module is further configured to: The feature consistency loss function is obtained by the following formula: ; Where, is the feature consistency loss function, i is the current number of upsampling layers, k is the total number of upsampling layers, is linear interpolation upsampling, is the self-distillation branch feature diagram, represents the binary activation process, F is the main decoding branch feature map, Represents the XOR process, Representation and process, Represents the Euclidean distance metric constraint.
[0102] In one possible implementation, the model training module is further configured to: The marginal loss function is obtained by the following formula: ; Where, is the marginal loss function, Pixel is the probability of the edge of the target area, Pixel Dynamic edge labels.
[0103] In a possible implementation, the first processing module is further configured to: Perform band selection on the remote sensing image to be processed to obtain a visible light image, and calculate expert prior knowledge based on the preset bands; Inputting the visible light image into a first feature extraction network to obtain spatial features output by the first feature extraction network; wherein the first feature extraction network includes an initial convolutional layer and multiple first stages, each of which includes a downsampling unit and multiple basic units; The expert prior knowledge is input into the second feature extraction network to obtain the spectral features output by the second feature extraction network; wherein the second feature extraction network includes an initial convolution layer and multiple second stages, each second stage includes a downsampling unit and at least one basic unit.
[0104] It should be noted that the real-time water segmentation system for multimodal remote sensing images adapted to drones provided by the present invention can, during specific operation, execute the real-time water segmentation method for multimodal remote sensing images adapted to drones of any of the above-mentioned embodiments, which will not be elaborated in this embodiment.
[0105] Figure 6 Schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 6 As shown, the electronic device may include: a processor 610 (processor), a communication interface 620 (Communications Interface), a memory 630 (memory), and a communication bus 640, wherein the processor 610, the communication interface 620, and the memory 630 communicate with each other via the communication bus 640. The processor 610 may call the logic instructions in the memory 630 to execute a real-time water segmentation method for multimodal remote sensing images adapted for drones. The method includes: determining spatial features and spectral features based on the remote sensing image to be processed; generating shared fusion weights based on the spatial features and spectral features, and performing feature fusion on the spatial features and spectral features based on the shared fusion weights to obtain fused features; inputting the fused features into a water segmentation prediction model to obtain a water segmentation prediction result output by the water segmentation prediction model; wherein the water segmentation prediction model is trained based on fused feature samples and true pixel labels.
[0106] Furthermore, the logic instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0107] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the real-time water body segmentation method of multimodal remote sensing images adapted to drones provided in the above-mentioned embodiments. The method includes: determining spatial features and spectral features based on the remote sensing images to be processed; generating shared fusion weights based on the spatial features and spectral features, and performing feature fusion on the spatial features and spectral features based on the shared fusion weights to obtain fusion features; inputting the fusion features into a water body segmentation prediction model to obtain a water body segmentation prediction result output by the water body segmentation prediction model; wherein the water body segmentation prediction model is trained based on the fusion feature samples and the true pixel labels.
[0108] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the real-time water body segmentation method for multimodal remote sensing images adapted to drones provided in the above-mentioned embodiments, the method comprising: determining spatial features and spectral features based on the remote sensing image to be processed; generating shared fusion weights based on the spatial features and spectral features, and performing feature fusion on the spatial features and spectral features based on the shared fusion weights to obtain fusion features; inputting the fusion features into a water body segmentation prediction model to obtain a water body segmentation prediction result output by the water body segmentation prediction model; wherein the water body segmentation prediction model is trained based on fusion feature samples and true pixel labels.
[0109] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units. That is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0110] Through the description of the above embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods of each embodiment or certain portions of the embodiments.
[0111] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A real-time water segmentation method for multimodal remote sensing images adapted to UAVs, characterized by: include: Determine spatial and spectral characteristics based on the remote sensing image to be processed; generating a shared fusion weight according to the spatial feature and the spectral feature, and performing feature fusion on the spatial feature and the spectral feature according to the shared fusion weight to obtain a fusion feature; The fusion feature is input into a water body segmentation prediction model to obtain a water body segmentation prediction result output by the water body segmentation prediction model; wherein the water body segmentation prediction model is trained based on the fusion feature samples and the true pixel labels.
2. The method for real-time water segmentation of multimodal remote sensing images adapted to UAVs according to claim 1 is characterized in that: Generating a shared fusion weight according to the spatial feature and the spectral feature includes: Performing spatial attention calculations on the spatial features and the spectral features respectively to generate a spatial feature attention map and a spectral feature attention map respectively; Splicing the spatial feature attention map and the spectral feature attention map in the channel dimension to form a spliced feature map containing spatial information; A convolution operation is performed on the concatenated feature map and activated by Sigmoid to obtain the shared fusion weight.
3. The method for real-time water segmentation of multimodal remote sensing images adapted to UAVs according to claim 1 is characterized in that: The performing feature fusion on the spatial feature and the spectral feature according to the shared fusion weight to obtain a fusion feature includes: The fusion feature is obtained by the following formula: ; Where, is the fusion feature, is the spatial feature, is the spectral feature, is the shared fusion weight.
4. The method for real-time water segmentation of multimodal remote sensing images adapted to UAVs according to claim 1, characterized in that: Also included is a training method for the water body segmentation prediction model: Inputting the fused feature samples into the main decoder and the self-distillation decoder of the initial model respectively, obtaining the main decoding branch feature map and the first prediction result output by the main decoder, and the self-distillation branch feature map and the second prediction result output by the self-distillation decoder; Binarizing the first prediction result and extracting the predicted edge to obtain a dynamic edge label; Constructing a feature consistency loss function according to the Euclidean distance between the main decoding branch feature map and the self-distillation branch feature map; constructing an edge loss function according to a difference between the second prediction result and the dynamic edge label; Constructing a subject loss function based on the difference between the first prediction result and the true label of the pixel; Summing the feature consistency loss function, the edge loss function, and the main loss function to obtain a total loss function, and optimizing the network parameters of the initial model according to the total loss function; After the training is completed, the self-distillation decoder is removed and the main decoder is retained to form the water body segmentation prediction model.
5. The method for real-time water segmentation of multimodal remote sensing images adapted to UAVs according to claim 4 is characterized in that: The constructing a feature consistency loss function according to the Euclidean distance between the main decoding branch feature map and the self-distillation branch feature map includes: The feature consistency loss function is obtained by the following formula: ; Where, is the feature consistency loss function, i is the current upsampling layer number, k is the total upsampling layer number, is linear interpolation upsampling, is the self-distillation branch characteristic diagram, represents the binary activation process, F is the main decoding branch feature map, Represents the XOR process, Representation and process, Represents the Euclidean distance metric constraint.
6. The method for real-time water segmentation of multimodal remote sensing images adapted to UAVs according to claim 4, characterized in that: The constructing an edge loss function according to the difference between the second prediction result and the dynamic edge label includes: The marginal loss function is obtained by the following formula: ; Where, is the marginal loss function, Pixel is the probability of the edge of the target area, Pixel Dynamic edge labels.
7. The method for real-time water segmentation of multimodal remote sensing images adapted to UAVs according to claim 1, characterized in that: Determining the spatial characteristics and spectral characteristics based on the remote sensing image to be processed includes: Performing band selection on the remote sensing image to be processed to obtain a visible light image, and obtaining expert prior knowledge based on the preset band calculation; Inputting the visible light image into a first feature extraction network to obtain spatial features output by the first feature extraction network; wherein the first feature extraction network includes an initial convolutional layer and multiple first stages, each of the first stages includes a downsampling unit and multiple basic units; The expert prior knowledge is input into a second feature extraction network to obtain spectral features output by the second feature extraction network; wherein the second feature extraction network includes an initial convolutional layer and multiple second stages, each of the second stages includes a downsampling unit and at least one basic unit.
8. A real-time water segmentation system for multimodal remote sensing images adapted to UAVs, characterized by: include: A first processing module is used to determine spatial features and spectral features based on the remote sensing image to be processed; a second processing module, configured to generate a shared fusion weight according to the spatial feature and the spectral feature, and perform feature fusion on the spatial feature and the spectral feature according to the shared fusion weight to obtain a fusion feature; The third processing module is used to input the fusion feature into the water body segmentation prediction model to obtain the water body segmentation prediction result output by the water body segmentation prediction model; wherein, the water body segmentation prediction model is trained based on the fusion feature samples and the true pixel labels.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the real-time water segmentation method for multimodal remote sensing images adapted to drones as described in any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the real-time water segmentation method of multimodal remote sensing images adapted to drones as claimed in any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Unmanned aerial vehicle image boundary segmentation method and system, and storage medium
CN121482074A
Polymorphic water body segmentation method based on morphological prior and hybrid expert network
CN121616831A
A multi-morphology water body segmentation method based on morphology prior and hybrid expert network
CN121616831B