Mangrove forest remote sensing image processing method based on unmanned aerial vehicle
By using UAV remote sensing image processing methods, the reflectance of mangrove bottoms is reconstructed using water body attenuation coefficient and water depth characteristics. Combined with deep learning networks, the problem of spectral misjudgment in mangrove remote sensing images during tidal changes is solved, achieving high-precision identification and clear segmentation of mangroves.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG OCEAN UNIVERSITY
- Filing Date
- 2026-03-23
- Publication Date
- 2026-06-05
AI Technical Summary
The spectral characteristics of mangrove remote sensing images change significantly with tides, resulting in different spectral curves for the same tree species at high and low tides, which can easily lead to misjudgment when directly classified and identified.
Using UAV remote sensing image processing methods, the diffuse attenuation coefficient of water body is obtained by inverting the water body spectrum in deep water area, and bottom reflectance is reconstructed by combining pixel-by-pixel instantaneous water depth. A cross-attention mechanism is used to fuse water depth and spectral features, and a deep learning network is constructed for feature extraction and enhancement.
It improves the accuracy of mangrove identification, solves the identification difficulties caused by the large differences in the scale of mangrove canopies, enhances the edge accuracy of segmentation results, and makes the outline of mangroves clearer and more accurate.
Smart Images

Figure CN122156991A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of mangrove processing technology, specifically relating to a method for processing remote sensing images of mangroves based on unmanned aerial vehicles (UAVs). Background Technology
[0002] Mangroves are unique wetland ecosystems that grow in the intertidal zone of tropical and subtropical coasts, playing important ecological roles such as wind and wave protection, siltation and beach conservation, carbon sequestration and storage, and maintaining biodiversity. Accurately understanding the distribution range, area changes, and health status of mangroves is crucial for mangrove resource protection, ecological restoration, and coastal zone management.
[0003] Currently, mangrove monitoring mainly employs satellite remote sensing and UAV remote sensing technologies. Among these, UAV remote sensing, with its advantages of maneuverability, high spatial resolution, and ability to acquire multi-angle / multispectral data, has become an important means of precise mangrove monitoring. However, because mangroves grow in the intertidal zone and are affected by periodic tidal flooding, their spectral characteristics change significantly with the tides, posing the following technical challenges to remote sensing identification: During high tide, mangrove trunks and even part of the canopy are submerged by seawater. Sunlight penetrates the water layer to reach the canopy, and after reflection, it penetrates the water layer again and returns to the sensor. The water layer has a selective absorption effect on sunlight (strong absorption in the red band and strong penetration in the blue and green bands), causing distortion of the spectral signal received by the sensor. The same tree species exhibits different spectral curves at high and low tides, a phenomenon known as "different spectra for the same species," which can easily lead to misjudgments when directly classifying and identifying based on raw images. Summary of the Invention
[0004] To address the above problems, this invention proposes a method for processing remote sensing images of mangrove forests based on unmanned aerial vehicles (UAVs).
[0005] The technical solution of this invention is: a method for processing remote sensing images of mangroves based on unmanned aerial vehicles (UAVs) includes the following steps:
[0006] S1. Use drones to capture hyperspectral remote sensing images of the target mangrove area;
[0007] S2. Perform spectral processing and optimization inversion on the hyperspectral remote sensing image to obtain the corrected hyperspectral remote sensing image;
[0008] S3. Perform feature extraction and feature enhancement on the corrected hyperspectral remote sensing image to obtain the probability that each pixel belongs to the mangrove forest and determine the final mangrove forest area.
[0009] Furthermore, S2 includes the following sub-steps:
[0010] S21. Determine the tide level elevation during drone photography and identify the deep-water pixels in the hyperspectral remote sensing image.
[0011] S22. Extract the original spectral reflectance of all deep-water pixels in each band, calculate the average spectrum of each band, and obtain the water spectrum of the deep-water area.
[0012] S23. Subtract the water surface reflectance from the reflectance values of each band in the deep water spectrum to obtain the actual water reflectance away from the water.
[0013] S24. Based on the actual water body's reflectivity, perform optimization and inversion to obtain the water body's diffuse attenuation coefficient;
[0014] S25. Based on the diffuse attenuation coefficient of the water body, the bottom reflectance of each pixel is reconstructed, and the reconstructed bottom reflectance is assigned to the pixel position to obtain the corrected hyperspectral remote sensing image.
[0015] The beneficial effects of the above-mentioned further solutions are as follows: In this invention, mangroves grow in the intertidal zone, and during high tide, the trunks and even part of the crowns are submerged by seawater. Seawater has a strong selective absorption of sunlight (red light is absorbed, while blue and green light has strong penetrability), which will cause the same tree species to show completely different spectra received by the sensor at the moment of low tide and long after low tide, thus the same tree species will show different spectral curves. If the original images are used directly for classification, misjudgment is likely to occur.
[0016] The attenuation of light by water follows Beer-Lambert's law, meaning that light intensity decreases exponentially with water depth. Therefore, this invention utilizes the spectral information of pure water areas (deep water areas) and vegetation areas just above the water surface in an image to obtain two key parameters: the attenuation coefficient and the thickness of the water layer. With these two parameters, an equation can be established to remove the filtering effect of the water. The output low reflectance image is equivalent to simulating the true spectrum of mangroves under conditions of complete absence of tidal cover.
[0017] In S22, the original reflectance values of deep-water pixels in each band are extracted and set together. Then, the arithmetic mean of that band is calculated. After calculating for all bands separately, a spectral curve with the band as the independent variable and the average reflectance as the dependent variable is obtained, which is the water spectrum of the deep-water area.
[0018] Furthermore, in S21, based on the historical point cloud data of the target mangrove area, a DEM model is determined. The difference between the tide level at the time of drone photography and the DEM elevation of the pixel in the DEM model is taken as the instantaneous water depth of the pixel. Pixels with an instantaneous water depth greater than a set threshold are taken as deep water pixels.
[0019] Because the surface elevation of mudflats in mangrove areas is stable in the short term, historical point cloud data needs to be acquired at low tide.
[0020] In S23, , It is a set of values corresponding to the reflectivity of various wavebands. It is a constant; subtracting it gives the water's reflectance when it leaves the water. .
[0021] Furthermore, S24 includes the following sub-steps:
[0022] S241. Use the pure water absorption coefficient as the initial value of the total water absorption coefficient, and use the sum of the total water absorption coefficient and the backscattering coefficient as the initial inversion factor.
[0023] S242. The ratio of the backscattering coefficient to the initial inversion factor is taken as the inversion factor to be inverted.
[0024] S243. Multiply the illumination geometric constant with the inversion factor and perform inversion to minimize the difference between the theoretical water reflectance and the actual water reflectance, and obtain the inverted backscattering coefficient.
[0025] S244. Add the backscattering coefficient after inversion to the absorption coefficient of pure water to obtain the diffuse attenuation coefficient of the water body.
[0026] The beneficial effects of the above-mentioned further scheme are as follows: In this invention, during the inversion process, the inputs are the inherent optical properties (absorption coefficient and backscattering coefficient) of the water body; the output is the apparent optical quantity that can be observed by the sensor, i.e., the difference between the theoretical water reflectance and the actual water reflectance. The illumination geometry factor is determined based on the solar altitude angle and observation angle during UAV photography. For UAV images observed vertically, an empirical constant value can be approximated, i.e., the commonly used midday flight period. The calculated actual water reflectance is used as input to establish the relationship between the water reflectance and the inherent optical property parameters of the water body.
[0027] In the specific inversion process, the known absorption coefficient of pure water is first introduced as the initial value of the total absorption coefficient of the water body, and it is assumed that the contribution of other components in the water body (such as yellow substances and non-pigment suspended particles) to light absorption is negligible. Then, the present invention continuously adjusts the value of the backscattering coefficient to minimize the error between the theoretical water-leaving reflectance and the water-leaving reflectance measured in step S23. When the error between the two is minimized, the corresponding backscattering coefficient is the actual backscattering coefficient of the water body in that region.
[0028] Finally, the backscattering coefficient obtained from the inversion is added to the absorption coefficient of pure water to calculate the diffuse attenuation coefficient of the water body. This diffuse attenuation coefficient describes the degree of attenuation of light as it propagates in water.
[0029] Therefore, in this embodiment of the invention, the water reflectivity upon leaving the water is determined by the illumination geometric constant, the water backscattering coefficient, and the water total absorption coefficient. Specifically, the water reflectivity upon leaving the water is proportional to the illumination geometric constant and proportional to the ratio of the water backscattering coefficient to the water total absorption coefficient.
[0030] Furthermore, in S25, the expression for reconstructing the bottom reflectance is:
[0031] ;
[0032] in, Indicates the band as Low reflectivity, Indicates the band as The diffuse attenuation coefficient of water bodies, This indicates that the sensor received the band as apparent reflectivity, Indicates water surface reflectivity. Indicates the band as The actual water reflectance when it leaves the water. Indicates the instantaneous water depth of a pixel. Indicates an index.
[0033] The beneficial effect of the above-mentioned further solution is that, in this invention, bottom reflectivity reconstruction considers three physical processes: water surface reflection, water body scattering, and water layer attenuation. The apparent reflectivity received by the sensor consists of three parts: water surface reflection, water body scattering, and attenuated bottom reflection. Subtracting these components yields the attenuated bottom reflection signal, which is then multiplied by the exponentiation of the attenuation coefficient to recover the true bottom reflectivity before attenuation.
[0034] Furthermore, S3 includes the following sub-steps:
[0035] S31. Generate a pixel-by-pixel instantaneous water depth map with the same size as the hyperspectral remote sensing image based on the instantaneous water depth of the pixels.
[0036] S32. Normalize the pixel-by-pixel instantaneous water depth map and the corrected hyperspectral remote sensing image, and extract the corresponding spectral features and water depth features using three-dimensional convolution.
[0037] S33. Attention modulation is applied to the spectral features and water depth features to obtain the modulated spectral features;
[0038] S34. Perform 1×1 convolution, 3×3 dilated convolution with a dilation rate of 2, 3×3 dilated convolution with a dilation rate of 4, and global average pooling on the modulated spectral features respectively. Then, concatenate the outputs of the four branches along the channels to obtain the fused feature map.
[0039] S35. Extract the gradient map of the pixel-by-pixel instantaneous water depth map, downsample the gradient map to the same size as the fused feature map, and perform enhancement processing to obtain the enhanced feature map;
[0040] S36. Input the enhanced feature map into the decoder and use the Sigmoid function to output the mangrove probability map to obtain the probability that each pixel belongs to the mangrove forest and determine the final mangrove forest area.
[0041] The beneficial effects of the above-mentioned further scheme are as follows: In this invention, discrete water depth data is transformed into image form, allowing water depth information to spatially correspond one-to-one with hyperspectral images, facilitating the processing of images of two modalities. Attention modulation is used to complete the interaction between water depth and spectrum, relying more on water depth information in deep water areas and more on spectral information in shallow water areas, so that the modulated features retain both local spectral details and incorporate global water depth distribution information. Convolutions with different dilation rates enable the network to simultaneously perceive canopies at different scales. The outputs of the four branches are concatenated along the channel dimension, and can also be fused through 1×1 convolutions to obtain a fused feature map. The enhanced feature map is input into the decoder, which gradually restores the spatial resolution through multiple upsampling steps, and performs skip connections (concatenation) with the features of the corresponding layer of the encoder at each upsampling stage, fusing deep semantics and shallow details. The final output is a feature map of the same size as the original image, which is then activated by the Sigmoid activation function to obtain the mangrove probability value for each pixel. The Sigmoid output probability value can be subsequently thresholded (e.g., 0.5) to obtain a binarized mangrove distribution map.
[0042] In S33, the attention modulation module dynamically adjusts the weights of spectral features using water depth information, which conforms to the physical law that the greater the water depth, the lower the reliability of the spectrum. In S34, multi-scale dilated convolution captures the features of tree canopies of different sizes. In S35, gradient enhancement utilizes the characteristic that the boundary of the mangrove forest is where the water depth changes drastically to strengthen the edge features.
[0043] Based on the instantaneous water depth value of each pixel calculated in step S11, the pixels are arranged according to their spatial position in the hyperspectral image to generate a pixel-by-pixel instantaneous water depth map with the same size as the hyperspectral image. The value of each pixel in the water depth map is the water layer thickness at that location.
[0044] Furthermore, in S33, the water depth features are mapped to a query matrix through convolution, and the spectral features are mapped to a key matrix and a value matrix through convolution; the cross-attention weights between the query matrix and the key matrix are multiplied with the value matrix and reshaped to obtain the modulated spectral features.
[0045] The beneficial effects of the above-mentioned further solutions are as follows: In this invention, a cross-attention mechanism is used to achieve deep fusion of water depth and spectral features. Compared with simple splicing or weighted summation, this invention can complete cross-modal interaction. The water depth features at each location can include spectral features, and spectral information related to the current water depth can be aggregated, so that the modulated features retain spectral details and are integrated into the water depth context.
[0046] The similarity between the query matrix and the key matrix is calculated and normalized using Softmax to obtain the cross-attention weights. The cross-attention weights are then multiplied by the value matrix to obtain the weighted feature sequence, which is then reshaped into a feature map.
[0047] Furthermore, in S35, the feature map is enhanced. The expression is:
[0048] ;
[0049] in, Represents the fused feature map. This indicates element-wise multiplication. This represents the gradient map after downsampling. Represents a 1×1 convolution. This represents the activation function. This indicates that the activation result will be copied along the channel dimension. Second-rate, This represents the number of channels in the fused feature map.
[0050] The beneficial effect of the above-mentioned further solution is that, in this invention, areas of drastic water depth change often mark the boundary between mangroves and the water / mudflat. Therefore, this invention utilizes the gradient information of the water depth map to highlight the edge features of the mangroves. The gradient magnitude is calculated from the water depth map, and the gradient map... Downsampled to the same size as the fused feature map, resulting in Through a convolutional layer The mapping is applied to single-channel spatial attention weights and copied along the channel dimension to obtain an enhanced feature map. That is, the features of edge regions (with large gradients) are enhanced, while the features of non-edge regions remain unchanged.
[0051] The beneficial effects of this invention are:
[0052] (1) This invention utilizes the diffuse attenuation coefficient of water body in the deep water area in the hyperspectral remote sensing image to invert the water body spectrum, and combines it with the instantaneous water depth per pixel to reconstruct the bottom reflectance of vegetation pixels submerged by tide.
[0053] (2) This invention constructs a water depth-guided deep learning network, extracts spectral features and water depth features from the corrected hyperspectral image, and uses a cross-attention mechanism to achieve interaction between the two modes, so that each water depth position can be connected to all spectral positions, thereby improving the recognition accuracy of complex tidal environments; it uses dilated convolution with different dilation rates to extract features in parallel, and combines global average pooling to obtain full-image context information, which can perceive mangrove patches of different scales, solve the problem of recognition difficulties caused by large differences in the scale of mangrove canopies, improve the edge accuracy of the segmentation results, and make the outline of mangroves clearer and more accurate. Attached Figure Description
[0054] Figure 1 This is a flowchart of a method for processing remote sensing images of mangroves based on unmanned aerial vehicles (UAVs). Detailed Implementation
[0055] The embodiments of the present invention will be further described below with reference to the accompanying drawings.
[0056] like Figure 1 As shown, this invention provides a method for processing remote sensing images of mangroves based on unmanned aerial vehicles (UAVs), comprising the following steps:
[0057] S1. Use drones to capture hyperspectral remote sensing images of the target mangrove area;
[0058] S2. Perform spectral processing and optimization inversion on the hyperspectral remote sensing image to obtain the corrected hyperspectral remote sensing image;
[0059] S3. Perform feature extraction and feature enhancement on the corrected hyperspectral remote sensing image to obtain the probability that each pixel belongs to the mangrove forest and determine the final mangrove forest area.
[0060] In this embodiment of the invention, S2 includes the following sub-steps:
[0061] S21. Determine the tide level elevation during drone photography and identify the deep-water pixels in the hyperspectral remote sensing image.
[0062] S22. Extract the original spectral reflectance of all deep-water pixels in each band, calculate the average spectrum of each band, and obtain the water spectrum of the deep-water area.
[0063] S23. Subtract the water surface reflectance from the reflectance values of each band in the deep water spectrum to obtain the actual water reflectance away from the water.
[0064] S24. Based on the actual water body's reflectivity, perform optimization and inversion to obtain the water body's diffuse attenuation coefficient;
[0065] S25. Based on the diffuse attenuation coefficient of the water body, the bottom reflectance of each pixel is reconstructed, and the reconstructed bottom reflectance is assigned to the pixel position to obtain the corrected hyperspectral remote sensing image.
[0066] In this invention, mangroves grow in the intertidal zone, and at high tide, their trunks and even parts of their crowns are submerged by seawater. Seawater has a strong selective absorption of sunlight (red light is absorbed, while blue and green light are highly penetrating), which causes the same tree species to exhibit completely different spectra received by the sensor at low tide and long after low tide, resulting in the same tree species displaying different spectral curves. If the raw images are used directly for classification, misclassification is likely to occur.
[0067] The attenuation of light by water follows Beer-Lambert's law, meaning that light intensity decreases exponentially with water depth. Therefore, this invention utilizes the spectral information of pure water areas (deep water areas) and vegetation areas just above the water surface in an image to obtain two key parameters: the attenuation coefficient and the thickness of the water layer. With these two parameters, an equation can be established to remove the filtering effect of the water. The output low reflectance image is equivalent to simulating the true spectrum of mangroves under conditions of complete absence of tidal cover.
[0068] In S22, the original reflectance values of deep-water pixels in each band are extracted and set together. Then, the arithmetic mean of that band is calculated. After calculating for all bands separately, a spectral curve with the band as the independent variable and the average reflectance as the dependent variable is obtained, which is the water spectrum of the deep-water area.
[0069] In this embodiment of the invention, in S21, a DEM model is determined based on the historical point cloud data of the target mangrove area. The difference between the tide level at the time of drone shooting and the DEM elevation of the pixel in the DEM model is taken as the instantaneous water depth of the pixel. Pixels with an instantaneous water depth greater than a set threshold are taken as deep water pixels.
[0070] Because the surface elevation of mudflats in mangrove areas is stable in the short term, historical point cloud data needs to be acquired at low tide.
[0071] In S23, , It is a set of values corresponding to the reflectivity of various wavebands. It is a constant; subtracting it gives the water's reflectance when it leaves the water. .
[0072] In this embodiment of the invention, S24 includes the following sub-steps:
[0073] S241. Use the pure water absorption coefficient as the initial value of the total water absorption coefficient, and use the sum of the total water absorption coefficient and the backscattering coefficient as the initial inversion factor.
[0074] S242. The ratio of the backscattering coefficient to the initial inversion factor is taken as the inversion factor to be inverted.
[0075] S243. Multiply the illumination geometric constant with the inversion factor and perform inversion to minimize the difference between the theoretical water reflectance and the actual water reflectance, and obtain the inverted backscattering coefficient.
[0076] S244. Add the backscattering coefficient after inversion to the absorption coefficient of pure water to obtain the diffuse attenuation coefficient of the water body.
[0077] In this invention, during the inversion process, the inputs are the inherent optical properties of the water body (absorption coefficient and backscattering coefficient); the output is the apparent optical quantity observable by sensors, i.e., the difference between the theoretical water reflectance and the actual water reflectance. The illumination geometry factor is determined based on the solar altitude angle and observation angle during UAV photography. For UAV images observed vertically, an empirical constant value can be approximated, i.e., the commonly used midday flight period. The calculated actual water reflectance is used as input to establish the relationship between the water reflectance and the inherent optical property parameters of the water body.
[0078] In the specific inversion process, the known absorption coefficient of pure water is first introduced as the initial value of the total absorption coefficient of the water body, and it is assumed that the contribution of other components in the water body (such as yellow substances and non-pigment suspended particles) to light absorption is negligible. Then, the present invention continuously adjusts the value of the backscattering coefficient to minimize the error between the theoretical water-leaving reflectance and the water-leaving reflectance measured in step S23. When the error between the two is minimized, the corresponding backscattering coefficient is the actual backscattering coefficient of the water body in that region.
[0079] Finally, the backscattering coefficient obtained from the inversion is added to the absorption coefficient of pure water to calculate the diffuse attenuation coefficient of the water body. This diffuse attenuation coefficient describes the degree of attenuation of light as it propagates in water.
[0080] Therefore, in this embodiment of the invention, the water reflectivity upon leaving the water is determined by the illumination geometric constant, the water backscattering coefficient, and the water total absorption coefficient. Specifically, the water reflectivity upon leaving the water is proportional to the illumination geometric constant and proportional to the ratio of the water backscattering coefficient to the water total absorption coefficient.
[0081] In this embodiment of the invention, in S25, the expression for reconstructing the bottom reflectivity is:
[0082] ;
[0083] in, Indicates the band as Low reflectivity, Indicates the band as The diffuse attenuation coefficient of water bodies, This indicates that the sensor received the band as apparent reflectivity, Indicates water surface reflectivity. Indicates the band as The actual water reflectance when it leaves the water. Indicates the instantaneous water depth of a pixel. Indicates an index.
[0084] In this invention, bottom reflectivity reconstruction considers three physical processes: water surface reflection, water scattering, and water layer attenuation. The apparent reflectivity received by the sensor consists of three parts: water surface reflection, water scattering, and attenuated bottom reflection. Subtracting these components yields the attenuated bottom reflection signal, which is then multiplied by the exponentiation of the attenuation coefficient to recover the true bottom reflectivity before attenuation.
[0085] In this embodiment of the invention, S3 includes the following sub-steps:
[0086] S31. Generate a pixel-by-pixel instantaneous water depth map with the same size as the hyperspectral remote sensing image based on the instantaneous water depth of the pixels.
[0087] S32. Normalize the pixel-by-pixel instantaneous water depth map and the corrected hyperspectral remote sensing image, and extract the corresponding spectral features and water depth features using three-dimensional convolution.
[0088] S33. Attention modulation is applied to the spectral features and water depth features to obtain the modulated spectral features;
[0089] S34. Perform 1×1 convolution, 3×3 dilated convolution with a dilation rate of 2, 3×3 dilated convolution with a dilation rate of 4, and global average pooling on the modulated spectral features respectively. Then, concatenate the outputs of the four branches along the channels to obtain the fused feature map.
[0090] S35. Extract the gradient map of the pixel-by-pixel instantaneous water depth map, downsample the gradient map to the same size as the fused feature map, and perform enhancement processing to obtain the enhanced feature map;
[0091] S36. Input the enhanced feature map into the decoder and use the Sigmoid function to output the mangrove probability map to obtain the probability that each pixel belongs to the mangrove forest and determine the final mangrove forest area.
[0092] In this invention, discrete water depth data is transformed into image form, allowing water depth information to spatially correspond one-to-one with hyperspectral images, facilitating the processing of images from both modalities. Attention modulation is used to achieve interaction between water depth and spectrum, relying more on water depth information in deep water and more on spectral information in shallow water, ensuring that the modulated features retain both local spectral details and global water depth distribution information. Convolutions with different dilation rates enable the network to simultaneously perceive canopies at different scales. The outputs of the four branches are concatenated along the channel dimension, and can also be fused using 1×1 convolutions to obtain a fused feature map. The enhanced feature map is input into the decoder, which gradually restores spatial resolution through multiple upsampling steps, and performs skip connections (concatenation) with the features of the corresponding layer of the encoder at each upsampling stage, fusing deep semantics and shallow details. The final output is a feature map of the same size as the original image, which is then processed by a Sigmoid activation function to obtain the mangrove probability value for each pixel. The Sigmoid output probability value can be subsequently thresholded (e.g., 0.5) to obtain a binarized mangrove distribution map.
[0093] In S33, the attention modulation module dynamically adjusts the weights of spectral features using water depth information, which conforms to the physical law that the greater the water depth, the lower the reliability of the spectrum. In S34, multi-scale dilated convolution captures the features of tree canopies of different sizes. In S35, gradient enhancement utilizes the characteristic that the boundary of the mangrove forest is where the water depth changes drastically to strengthen the edge features.
[0094] Based on the instantaneous water depth value of each pixel calculated in step S11, the pixels are arranged according to their spatial position in the hyperspectral image to generate a pixel-by-pixel instantaneous water depth map with the same size as the hyperspectral image. The value of each pixel in the water depth map is the water layer thickness at that location.
[0095] In this embodiment of the invention, in S33, the water depth features are mapped to a query matrix through convolution, and the spectral features are mapped to a key matrix and a value matrix through convolution; the cross-attention weights between the query matrix and the key matrix are multiplied with the value matrix and reshaped to obtain the modulated spectral features.
[0096] In this invention, a cross-attention mechanism is employed to achieve deep fusion of water depth and spectral features. Compared to simple splicing or weighted summation, this invention can accomplish cross-modal interaction. The water depth features at each location can include spectral features, and spectral information related to the current water depth can be aggregated, so that the modulated features retain spectral details while incorporating the water depth context.
[0097] The similarity between the query matrix and the key matrix is calculated and normalized using Softmax to obtain the cross-attention weights. The cross-attention weights are then multiplied by the value matrix to obtain the weighted feature sequence, which is then reshaped into a feature map.
[0098] In this embodiment of the invention, in S35, the feature map is enhanced. The expression is:
[0099] ;
[0100] in, Represents the fused feature map. This indicates element-wise multiplication. This represents the gradient map after downsampling. Represents a 1×1 convolution. This represents the activation function. This indicates that the activation result will be copied along the channel dimension. Second-rate, This represents the number of channels in the fused feature map.
[0101] In this invention, areas of drastic water depth change often mark the boundary between mangroves and the water / mudflat. Therefore, this invention utilizes gradient information from water depth maps to highlight mangrove edge features. The gradient magnitude is calculated from the water depth map, and the gradient map... Downsampled to the same size as the fused feature map, resulting in Through a convolutional layer The mapping is applied to single-channel spatial attention weights and copied along the channel dimension to obtain an enhanced feature map. That is, the features of edge regions (with large gradients) are enhanced, while the features of non-edge regions remain unchanged.
[0102] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of this invention.
Claims
1. A method for processing remote sensing images of mangrove forests based on unmanned aerial vehicles (UAVs), characterized in that, Includes the following steps: S1. Use drones to capture hyperspectral remote sensing images of the target mangrove area; S2. Perform spectral processing and optimization inversion on the hyperspectral remote sensing image to obtain the corrected hyperspectral remote sensing image; S3. Perform feature extraction and feature enhancement on the corrected hyperspectral remote sensing image to obtain the probability that each pixel belongs to the mangrove forest and determine the final mangrove forest area.
2. The method for processing remote sensing images of mangroves based on unmanned aerial vehicles (UAVs) according to claim 1, characterized in that, S2 includes the following sub-steps: S21. Determine the tide level elevation during drone photography and identify the deep-water pixels in the hyperspectral remote sensing image. S22. Extract the original spectral reflectance of all deep-water pixels in each band, calculate the average spectrum of each band, and obtain the water spectrum of the deep-water area. S23. Subtract the water surface reflectance from the reflectance values of each band in the deep water spectrum to obtain the actual water reflectance away from the water. S24. Based on the actual water body's reflectivity, perform optimization and inversion to obtain the water body's diffuse attenuation coefficient; S25. Based on the diffuse attenuation coefficient of the water body, the bottom reflectance of each pixel is reconstructed, and the reconstructed bottom reflectance is assigned to the pixel position to obtain the corrected hyperspectral remote sensing image.
3. The method for processing remote sensing images of mangroves based on unmanned aerial vehicles (UAVs) according to claim 2, characterized in that, In step S21, a DEM model is determined based on historical point cloud data of the target mangrove area. The difference between the tide level at the time of drone photography and the DEM elevation of the pixel in the DEM model is taken as the instantaneous water depth of the pixel. Pixels with an instantaneous water depth greater than a set threshold are taken as deep water pixels.
4. The method for processing remote sensing images of mangroves based on unmanned aerial vehicles (UAVs) according to claim 2, characterized in that, S24 includes the following sub-steps: S241. Use the pure water absorption coefficient as the initial value of the total water absorption coefficient, and use the sum of the total water absorption coefficient and the backscattering coefficient as the initial inversion factor. S242. The ratio of the backscattering coefficient to the initial inversion factor is taken as the inversion factor to be inverted. S243. Multiply the illumination geometric constant with the inversion factor and perform inversion to minimize the difference between the theoretical water reflectance and the actual water reflectance, and obtain the inverted backscattering coefficient. S244. Add the backscattering coefficient after inversion to the absorption coefficient of pure water to obtain the diffuse attenuation coefficient of the water body.
5. The method for processing remote sensing images of mangroves based on unmanned aerial vehicles (UAVs) according to claim 2, characterized in that, In step S25, the expression for reconstructing the bottom reflectivity is: ; in, Indicates the band as Low reflectivity, Indicates the band as The diffuse attenuation coefficient of water bodies, This indicates that the sensor received the band as apparent reflectivity, Indicates water surface reflectivity. Indicates the band as The actual water reflectance when it leaves the water. Indicates the instantaneous water depth of a pixel. Indicates an index.
6. The method for processing remote sensing images of mangroves based on unmanned aerial vehicles (UAVs) according to claim 1, characterized in that, S3 includes the following sub-steps: S31. Generate a pixel-by-pixel instantaneous water depth map with the same size as the hyperspectral remote sensing image based on the instantaneous water depth of the pixels. S32. Normalize the pixel-by-pixel instantaneous water depth map and the corrected hyperspectral remote sensing image, and extract the corresponding spectral features and water depth features using three-dimensional convolution. S33. Attention modulation is applied to the spectral features and water depth features to obtain the modulated spectral features; S34. Perform 1×1 convolution, 3×3 dilated convolution with a dilation rate of 2, 3×3 dilated convolution with a dilation rate of 4, and global average pooling on the modulated spectral features respectively. Then, concatenate the outputs of the four branches along the channels to obtain the fused feature map. S35. Extract the gradient map of the pixel-by-pixel instantaneous water depth map, downsample the gradient map to the same size as the fused feature map, and perform enhancement processing to obtain the enhanced feature map; S36. Input the enhanced feature map into the decoder and use the Sigmoid function to output the mangrove probability map to obtain the probability that each pixel belongs to the mangrove forest and determine the final mangrove forest area.
7. The method for processing remote sensing images of mangrove forests based on unmanned aerial vehicles according to claim 6, characterized in that, In step S33, the water depth features are mapped to a query matrix through convolution, and the spectral features are mapped to a key matrix and a value matrix through convolution. The cross-attention weights between the query matrix and the key matrix are multiplied by the value matrix and reshaped to obtain the modulated spectral features.
8. The method for processing remote sensing images of mangroves based on unmanned aerial vehicles (UAVs) according to claim 6, characterized in that, In S35, the feature map is enhanced. The expression is: ; in, Represents the fused feature map. This indicates element-wise multiplication. This represents the gradient map after downsampling. Represents a 1×1 convolution. This represents the activation function. This indicates that the activation result will be copied along the channel dimension. Second-rate, This represents the number of channels in the fused feature map.