Estimation method and system for total flux of methane emission of coastal benthic animal cave based on unmanned aerial vehicle remote sensing and deep learning
By combining UAV remote sensing with deep learning, the contradiction between measurement accuracy and spatial coverage in estimating the total methane emission flux from benthic animal burrows in coastal mudflats was resolved. This approach enabled clear differentiation and efficient counting of burrows down to the centimeter level, providing an efficient and reliable regional assessment method.
Patent Information
- Application Number
- CN202610657464.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-13
- Publication Date
- 2026-08-04
AI Technical Summary
Existing technologies for estimating the total methane emission flux from benthic animal burrows in coastal mudflats face the dilemma of not being able to simultaneously achieve measurement accuracy and spatial coverage. Static box methods have poor spatial representativeness, manual visual counting methods are inefficient and highly subjective, and satellite remote sensing inversion methods have insufficient resolution, making it difficult to meet the requirements for high accuracy over a large area.
This study employs a method combining UAV remote sensing and deep learning. Ultra-high resolution images are acquired through aerial surveys during low tide. Cave detection is performed using a deep learning model with a sliding window strategy and attention mechanism. Geographic reference information is then used for fusion and deduplication to estimate the total methane emission flux.
It enables clear differentiation and accurate counting of individual cave specimens at the centimeter level, improving detection efficiency and accuracy, reducing monitoring costs and time consumption, and providing an efficient and reliable regional-scale assessment method.
Smart Images

Figure CN122510754A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of environmental monitoring and remote sensing technology, and in particular relates to a method and system for estimating the total methane emission flux from benthic animal burrows in coastal mudflats based on UAV remote sensing and deep learning. Background Technology
[0002] Coastal wetlands, as important global carbon sink ecosystems, play a crucial role in regulating the greenhouse gas balance and are also significant sources of greenhouse gases such as methane. Recent ecological and biogeochemical studies have shown that the burrowing systems constructed by widely distributed burrowing benthic animals (such as fiddler crabs and mudskippers) in coastal mudflats significantly alter gas exchange processes at the sediment-atmosphere interface, becoming "hotspots" for methane emissions. It is estimated that the total methane emission flux in areas affected by bioturbation can be 2 to 20 times that of surrounding undisturbed areas. Therefore, accurately estimating the total methane emission flux from benthic animal burrows in coastal mudflats is of great significance for assessing the greenhouse gas balance of coastal wetlands.
[0003] Currently, the main method for measuring methane flux in individual benthic animal burrows relies on the static box method. This method involves placing a sealed box on the Earth's surface and monitoring the change in gas concentration within the box over time to calculate the flux. It is a classic method for measuring gas exchange at the soil-atmosphere interface, with advantages such as mature principles and high measurement accuracy, and is widely used in point-scale flux observations. Methods for estimating the number of benthic animal burrows mainly include: First, the manual visual counting method, which involves manually identifying and counting burrows through field surveys or based on low-resolution aerial imagery. This is a fundamental means of obtaining burrow density distribution, is simple to operate, and requires no complex equipment. Second, the satellite remote sensing inversion method, which uses medium- to high-resolution satellite imagery (such as Sentinel-2 and Landsat series) to invert tidal flat ecological parameters. This method can cover a large geographical area and provides data support for large-scale ecological assessments.
[0004] However, the aforementioned existing technologies all have significant limitations in practical applications. While the static box method can achieve high-precision measurements, its spatial representativeness is poor; a single box typically covers only 0.01 to 0.25 square meters, making it difficult to capture flux variations caused by the heterogeneity of cave spatial distribution. Furthermore, large-scale deployment of static boxes requires substantial manpower and time, is difficult in muddy tidal flat environments, and the placement of the boxes may physically interfere with the cave's micro-topography, affecting the accuracy of the measurement results. Manual visual counting is inefficient, requiring days or even weeks on a regional scale of several to tens of hectares. Moreover, the identification results are highly subjective, with different observers using varying standards, resulting in a high rate of missed detections for small caves or those covered by algae, making it difficult to meet the needs of large-scale, high-precision statistics. Satellite remote sensing inversion methods are limited by the spatial resolution of images. The ground resolution of mainstream satellite images is mostly 10 to 30 meters, which is far from being able to identify benthic animal burrows with a diameter of only 1 to 8 centimeters. In addition, due to the influence of satellite revisit cycles and cloud cover, it is difficult to obtain effective image data during critical low tide periods, which further limits its application in refined monitoring.
[0005] In summary, existing technologies for estimating the total methane emission flux from benthic animal burrows in coastal mudflats at a regional scale have always faced the core contradiction of "the incompatibility between measurement accuracy and spatial coverage," and there is an urgent need for a technical solution that can achieve rapid estimation over a large area while ensuring accuracy. Summary of the Invention
[0006] To address the aforementioned technical problems, this invention proposes a method and system for estimating the total methane emission flux from benthic animal burrows in coastal mudflats based on UAV remote sensing and deep learning, thereby resolving the issues present in the existing technologies.
[0007] To achieve the above objectives, in a first aspect, the present invention provides a method for estimating the total methane emission flux from benthic animal burrows in coastal mudflats based on UAV remote sensing and deep learning, comprising: During low tide, drones are used to conduct aerial surveys of the target tidal flat area to obtain raw remote sensing data for constructing ultra-high resolution images that can distinguish the morphology of benthic animal burrows. The raw remote sensing data is processed to generate a digital orthophoto covering the target tidal flat area; The digital orthophoto image is segmented using a sliding window strategy to generate multiple image subsets that conform to the input size of the deep learning model. Each of the aforementioned image subsets is input into a trained deep learning detection model, which integrates an attention mechanism to enhance cave feature extraction capabilities, and outputs cave detection results for each image subset. Based on georeferenced information, the cave detection results of each image subset are fused in a unified coordinate system, and duplicate detections in the fused results are deduplicated to obtain the total number of benthic animal caves in the target tidal flat area. Based on the total number of caves and the average methane emission flux per cave obtained through field sampling, the total methane emission flux from benthic animal caves in the target tidal flat area is estimated.
[0008] Preferably, the flight altitude, heading overlap, and lateral overlap of the aerial survey are configured such that the ground sampling distance of the original remote sensing data reaches the sub-centimeter level.
[0009] Preferably, the sliding window strategy uses a window with a predetermined overlap rate to traverse and cut the digital orthophoto. The size of the window is set according to the input requirements of the deep learning detection model, and the predetermined overlap rate is configured to ensure that cave targets located at the edge of the slice are completely detected.
[0010] Preferably, the backbone network of the deep learning detection model integrates a spatial attention module and a channel attention module; the spatial attention module is used to enhance the extraction of cave edge contour features, and the channel attention module is used to enhance the extraction of features in the cave center region.
[0011] Preferably, the spatial attention module generates spatial attention weights by fusing average pooling and max pooling features and performing convolution processing; the channel attention module processes average pooling and max pooling features using a shared multilayer perceptron and fuses them to generate channel attention weights.
[0012] Preferably, the cave detection results of each image subset are fused in a unified coordinate system based on geographic reference information. Specifically, this includes converting the pixel coordinates of the detection results in each image subset into global geographic coordinates based on the ground sampling distance and reference point geographic coordinates of the digital orthophoto.
[0013] Preferably, the duplicate detections in the fusion result are deduplicated, specifically including: calculating the overlap index between multiple detection results located in the overlapping areas of different image subsets, and removing redundant detection results with an overlap index higher than a preset threshold.
[0014] Preferably, the total methane emission flux from benthic animal burrows in the target tidal flat area is estimated using the following formula: F_total = N × f_avg; Wherein, F_total is the total methane emission flux from benthic animal burrows in the target tidal flat area, N is the total number of burrows, and f_avg is the measured average methane emission flux per burrow.
[0015] Secondly, a system for estimating the total methane emission flux from benthic animal burrows in coastal mudflats based on UAV remote sensing and deep learning includes: The aerial survey module is used to control UAVs to conduct aerial surveys of target tidal flat areas during low tide and acquire raw remote sensing data for constructing ultra-high resolution images that distinguish the individual morphology of benthic animal burrows. The image processing module is used to process the raw remote sensing data to generate a digital orthophoto covering the target tidal flat area; The segmentation module is used to segment the digital orthophoto using a sliding window strategy to generate multiple image subsets that conform to the input size of the deep learning model. The detection module has a built-in trained deep learning detection model, which integrates an attention mechanism to enhance the ability to extract cave features. It is used to detect caves on a subset of the input image and output the detection results. The fusion counting module is used to fuse the cave detection results of each image subset in a unified coordinate system based on geographic reference information, remove duplicate detections in the fusion results, and count the total number of benthic animal caves in the target tidal flat area. The estimation module is used to estimate the total methane emission flux from benthic animal burrows in the target tidal flat area based on the total number of burrows and the preset average methane emission flux per burrow.
[0016] Preferably, the deep learning detection model in the detection module includes: The backbone network is used to extract deep features from a subset of the input image. The spatial attention submodule is used to perform spatial dimension weighting on the depth features to enhance the response of the cave edge contour features; The channel attention submodule is used to perform channel-dimensional weighting on the spatially weighted features to enhance the response of features in the central region of the cave. The detection head is used to output the detection bounding box and confidence score of the cave target based on the features enhanced by both spatial and channel attention.
[0017] Compared with the prior art, the present invention has the following advantages and technical effects: This invention utilizes unmanned aerial vehicles (UAVs) to conduct aerial surveys during low tide and generate digital orthophotos with sub-centimeter ground sampling distances. For the first time, it achieves clear differentiation of the morphology of benthic animal burrows with diameters as small as centimeters at a regional scale. This technological feature overcomes the technical bottleneck of satellite remote sensing's inability to observe microscopic biological structures due to insufficient spatial resolution, providing a reliable data foundation for subsequent accurate counting and fundamentally resolving the core contradiction in existing technologies where "measurement accuracy and spatial coverage are mutually exclusive."
[0018] This invention employs a sliding window strategy to segment ultra-large orthophotos with billions of pixels into standardized image subsets, and combines this with a deep learning detection model integrating an attention mechanism to achieve automatic identification and accurate counting of tiny caves in massive images. This technology improves detection efficiency to the minute level, eliminates subjective differences in manual identification, and keeps the average relative error between detection accuracy and manual counting within a very small range. It significantly improves the objectivity and accuracy of regional-scale cave density statistics, overcoming the shortcomings of low efficiency, strong subjectivity, and high false negative rate of manual visual counting.
[0019] This invention, based on obtaining the precise total number of caves, combines the average methane emission flux from a single cave obtained through field sampling, and directly uses a linear model to calculate the total regional methane emission flux. This technique abandons the indirect extrapolation path of traditional methods that rely on multiple environmental sensor data or complex gas diffusion models, directly transforming the "spatial distribution quantity of biological structures" into the "total greenhouse gas emission flux from benthic animal caves." This represents a technological leap from point-based measurement to regional estimation, significantly reducing monitoring costs and time consumption, and providing a completely new technical approach for estimating methane emissions from benthic animal caves.
[0020] This invention forms a complete technical closed loop from data acquisition to total throughput output by precisely selecting tidal windows in the acquisition step, implementing edge protection strategies for overlapping windows in the segmentation step, enhancing attention mechanisms for cave morphology in the detection step, and stitching cross-slice results based on geographic coordinates in the fusion step. This integrated solution is systematically optimized for special conditions such as muddy coastal mudflat environments, low-tide operating windows, and the small scale of caves. Compared with general environmental monitoring technologies, it has significant advantages in scene adaptability, processing robustness, and result reliability, providing an efficient and reliable technical means for the accurate assessment of greenhouse gas emissions from benthic animal caves in coastal wetlands. Attached Figure Description
[0021] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a general technical flowchart of an embodiment of the present invention; Figure 2 This is a schematic diagram of UAV aerial survey parameters and route planning according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the ultra-large DOM sliding window slicing strategy according to an embodiment of the present invention; Figure 4 This is a diagram of the attention-enhanced cave detection network architecture according to an embodiment of the present invention; Figure 5This is a flowchart of the global stitching and NMS deduplication process for slice detection results in an embodiment of the present invention. Figure 6 This is a schematic diagram of the cave aperture classification and flux correlation model according to an embodiment of the present invention; Figure 7 This is a heat map showing the spatial distribution of cave density in the study area according to an embodiment of the present invention. Figure 8 This is a system hardware and software architecture diagram according to an embodiment of the present invention. Detailed Implementation
[0022] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0023] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0024] Example 1 like Figure 1 As shown, this embodiment provides a method for estimating the total methane emission flux from benthic animal burrows in coastal mudflats based on UAV remote sensing and deep learning, including: Step 1: During low tide, use drones to conduct aerial surveys of the target tidal flat area to obtain raw remote sensing data for constructing ultra-high resolution images that can distinguish the individual morphology of benthic animal burrows. Furthermore, the flight altitude, heading overlap, and lateral overlap of the aerial survey are configured such that the ground sampling distance of the original remote sensing data reaches the sub-centimeter level.
[0025] Specifically, during low tide, drones equipped with RTK modules are used to conduct low-altitude, high-overlap aerial surveys of the tidal flats to acquire RGB image sequences.
[0026] This embodiment innovatively employs an extremely low-altitude flight altitude of 12 m combined with a 70% high-overlap aerial survey strategy to achieve ultra-high resolution imaging with a GSD of 0.3~0.5 cm, targeting the minute-scale features (1~8 cm in diameter) of benthic animal burrows.
[0027] Step 2: Process the raw remote sensing data to generate a digital orthophoto covering the target tidal flat area. The ground sampling distance of the digital orthophoto is configured to clearly identify benthic animal burrows. Specifically, sub-centimeter-level GSD digital orthophotos are generated. The principle of GSD (Ground Sampling Distance) digital orthophotos is to eliminate topographic projection differences by performing pixel-by-pixel geometric correction and radiometric optimization on high-resolution remote sensing images, combined with a digital elevation model (DEM), to generate an orthophoto dataset with an accurate ground scale.
[0028] Step 3: The digital orthophoto is segmented using a sliding window strategy to generate multiple image subsets that conform to the input size of the deep learning model; Furthermore, the sliding window strategy uses a window with a predetermined overlap rate to traverse and cut the digital orthophoto. The size of the window is set according to the input requirements of the deep learning detection model, and the predetermined overlap rate is configured to ensure that cave targets located at the edge of the slice are completely detected.
[0029] Specifically, a sliding window strategy is used to divide the super-large DOM into standard-sized slices suitable for the input of deep learning models, and then normalization is performed.
[0030] This embodiment addresses the contradiction between the massive amount of DOM image data (100-500 billion pixels) and the extremely small detection targets (occupying 10-100 pixels). It proposes a sliding window slice detection strategy, which achieves seamless stitching and deduplication of global results through geographic coordinate mapping and the NMS algorithm.
[0031] Step 4: Input each of the image subsets into the trained deep learning detection model, which integrates an attention mechanism to enhance the ability to extract cave features, and outputs the cave detection results for each image subset; Furthermore, the backbone network of the deep learning detection model integrates a spatial attention module and a channel attention module; the spatial attention module is used to enhance the extraction of cave edge contour features, and the channel attention module is used to enhance the extraction of features in the cave center region.
[0032] Furthermore, the spatial attention module generates spatial attention weights by fusing average pooling and max pooling features and performing convolution processing; the channel attention module processes average pooling and max pooling features through a shared multilayer perceptron and fuses them to generate channel attention weights.
[0033] Specifically, a trained attention-enhanced deep learning detection model is used to detect cave targets in each slice, outputting detection boxes and confidence scores.
[0034] This embodiment introduces a spatial attention module (SAM) and a channel attention module (CAM) into the YOLO or Faster R-CNN detection network to enhance the model's ability to extract discriminative features such as the circular edges of caves and the central dark area, effectively reducing the false detection rate.
[0035] Step 5: Based on georeferenced information, the cave detection results of each image subset are fused in a unified coordinate system, and duplicate detections in the fused results are deduplicated to obtain the total number of benthic animal caves in the target tidal flat area. Furthermore, based on georeferenced information, the cave detection results of each image subset are fused in a unified coordinate system. Specifically, this includes converting the pixel coordinates of the detection results in each image subset into global geographic coordinates according to the ground sampling distance and reference point geographic coordinates of the digital orthophoto.
[0036] Furthermore, duplicate detections in the fusion results are deduplicated, specifically by calculating the overlap index between multiple detection results located in overlapping regions of different image subsets, and removing redundant detection results with overlap indices higher than a preset threshold.
[0037] Specifically, the detection results in each slice are mapped back to the global geographic coordinate system using georeferenced information, and non-maximum suppression (NMS) is used to eliminate duplicate detections in overlapping areas. Step Six: Based on the total number of caves and the average methane emission flux per cave obtained through field sampling, estimate the total methane emission flux from benthic animal caves in the target tidal flat area.
[0038] Furthermore, the total regional methane emission flux is estimated based on the total number of caves and the average methane emission flux per cave, using the following formula: F_total = N × f_avg Wherein, F_total is the total methane emission flux from benthic animal burrows in the target tidal flat area, N is the total number of burrows, and f_avg is the measured average methane emission flux per burrow.
[0039] Specifically, by combining measured methane emission flux data from typical caves, the total methane emission flux from benthic caves in the region was inverted based on cave counting results.
[0040] This embodiment establishes a correlation model between the number of burrows and the total methane emission flux from benthic animal burrows. The total number of burrows is obtained through a detection network, and combined with the measured average flux per burrow, a rapid regional-scale inversion of the total methane emission flux is achieved.
[0041] Example 2 Based on the technical solution of Embodiment 1, this embodiment selects a typical mudflat in the Futian mangrove forest along the coast of Guangdong Province as the study area. The intertidal zone of this area is 1100 m wide and is mainly inhabited by large mudskippers (…). Boleophthalmus pectinirostris ) and the pleasing big-eyed crab ( Macrophthalmus eratoTwo species of burrowing benthic animals were observed. Preliminary surveys indicated that the burrow density in the area ranged from 20 to 150 burrows per m², with burrow diameters ranging from 1 to 6 cm.
[0042] Step 1: Data Collection; A schematic diagram of UAV aerial survey parameters and flight path planning, such as... Figure 2 As shown.
[0043] (1) Aerial survey equipment: This embodiment uses the DJI Mavic 3 Enterprise drone as the aerial survey platform, and its key parameters are as follows: Table 1 (2) Selection of aerial survey time window: The aerial survey will be conducted on December 25, 2025, specifically from 8:00 AM to 11:30 AM. This time slot was chosen based on the following considerations: Tidal conditions: According to local tide forecasts, the low tide time on this day is 9:10, and the low tide level is 0.38 m. The aerial survey period will cover 1.5 hours after low tide to ensure that a large area of the mudflats is exposed and free of standing water. Meteorological conditions: The weather will be relatively sunny, with northerly winds of 5.5–7.9 m / s and visibility ≥10 km. Sufficient sunlight will be beneficial for acquiring high-quality imagery, while low wind speeds will ensure flight stability. Solar altitude angle: During this period, the solar altitude angle is 30°~50°, which avoids the high light reflection caused by strong direct sunlight at noon and ensures sufficient illumination of ground objects.
[0044] (3) Flight parameter settings: To achieve sub-centimeter-level GSD, this embodiment employs an extremely low-altitude flight strategy, with the following specific parameters: Table 2 The theoretical formula for calculating GSD (Ground Sampling Distance) is: GSD = (H × Sw) / (f × Iw); Where: H is the relative flight altitude (12 m); Sw is the sensor width (17.3 mm); f is the lens focal length (12.3 mm); Iw is the image width in pixels (5280 pixels).
[0045] Substitute into the calculation: GSD = (12 × 17.3) / (12.3 × 5280) ≈ 0.0032 m = 0.32 cm; This GSD value means that each pixel corresponds to a 0.32 cm × 0.32 cm area on the ground, and can clearly distinguish cave targets with a diameter >1 cm.
[0046] (4) Data volume collected: A total of 5,158 valid RGB images were acquired during this aerial survey, with a total data volume of 42.6 GB. The image format is JPEG, and the size of each image is 5280×3956 pixels.
[0047] Step 2: High-precision DOM construction; (1) Processing software and environment; This embodiment uses DJI Terra version 5.1.1 for image processing.
[0048] (2) DOM generation and quality assessment; Table 3 Step 3: Cave Identification and Counting; (1) Analysis of technical challenges; The generated DOM image has a total pixel count of 226,562 × 162,843 = 36.9 billion pixels. The input size of the deep learning object detection network is 640×640 to 1024×1024 pixels. Therefore, this invention employs a small object detection strategy based on sliding window slicing.
[0049] (2) Sliding window slicing strategy; A diagram illustrating the strategy for slicing large DOM sliding windows, as shown below. Figure 3 As shown.
[0050] Slicing Parameter Design – Assume the DOM image size is W×H pixels, the slice window size is w×h pixels, and the window step size is sx×sy pixels. Then the number of rows and columns in the slice are: nc = (W - w) / sx + 1; nr = (H - h) / sy + 1; Table 4 The number of slices generated in this embodiment is: N_total = nc × nr = 295 × 212 = 62,540 slices (3) Attention-enhanced deep learning detection model; Attention-enhanced cave detection network architecture diagram, such as Figure 4 As shown.
[0051] This invention preferably uses YOLOv8 as the basic detection framework and introduces a dual attention mechanism in the backbone network: Spatial Attention Module (SAM): Ms(F) = σ(f 7×7 ([AvgPool(F); MaxPool(F)])); Channel Attention Module (CAM): Mc(F) = σ(MLP(AvgPool(F)) + MLP(MaxPool(F))); The model performance evaluation results are shown in Table 5.
[0052] Table 5 (4) Coordinate mapping and deduplication; Formula for converting pixel coordinates to geographic coordinates: E = E0 + X_global × GSD; N = N0 - Y_global × GSD; Intersection over Union (IoU) definition: IoU(B i B j ) = |B i ∩ B j | / |B i ∪ B j |; Among them, B i For the i-th bounding box, B j For the j-th bounding box; The NMS threshold was set to τ = 0.5. The final number of valid caves detected was 3,847,562, with an average cave density of 74 caves / m².
[0053] The flowchart for global stitching and NMS deduplication of slice detection results is as follows: Figure 5 As shown.
[0054] A schematic diagram of the correlation model between cave aperture classification and flux, as shown below. Figure 6 As shown.
[0055] The study area features a thermal map of the spatial distribution of cave density, such as... Figure 7 As shown.
[0056] Step 4: Inversion model of total methane flux in regional benthic animal burrows; (1) Basic inversion model: F_total = N × f_avg; Where: F_total is the total methane emission flux from benthic burrows in the region; N is the total number of burrows detected; and f_avg is the measured average methane emission flux per burrow.
[0057] (2) The measured flux data of benthic animal burrows are shown in Table 6.
[0058] Table 6 (3) Total flux inversion results: Total flux inversion results based on cavern counts: F_total = 3,847,562 × 5.45 = 20,969,213 μmol / h; F_area = 20,969,213 / 57,000 ≈ 367.9 μmol / (m²·h); Step 5: Verify the results; Table 7 Step Six: Calculate efficiency; Table 8 Example 3 Based on the technical solution described in Embodiment 1, this embodiment provides a system for estimating the total methane emission flux from benthic animal burrows in coastal mudflats using UAV remote sensing and deep learning. Figure 8 As shown, it includes: The aerial survey module is used to control UAVs to conduct aerial surveys of target tidal flat areas during low tide and acquire raw remote sensing data for constructing ultra-high resolution images that distinguish the individual morphology of benthic animal burrows. The image processing module is used to process the raw remote sensing data to generate a digital orthophoto covering the target tidal flat area; The segmentation module is used to segment the digital orthophoto using a sliding window strategy to generate multiple image subsets that conform to the input size of the deep learning model. The detection module has a built-in trained deep learning detection model, which integrates an attention mechanism to enhance the ability to extract cave features. It is used to detect caves on a subset of the input image and output the detection results. The fusion counting module is used to fuse the cave detection results of each image subset in a unified coordinate system based on geographic reference information, remove duplicate detections in the fusion results, and count the total number of benthic animal caves in the target tidal flat area. The estimation module is used to estimate the total methane emission flux from benthic animal burrows in the target tidal flat area based on the total number of burrows and the preset average methane emission flux per burrow.
[0059] Furthermore, the deep learning detection model in the detection module includes: The backbone network is used to extract deep features from a subset of the input image. The spatial attention submodule is used to perform spatial dimension weighting on the depth features to enhance the response of the cave edge contour features; The channel attention submodule is used to perform channel-dimensional weighting on the spatially weighted features to enhance the response of features in the central region of the cave. The detection head is used to output the detection bounding box and confidence score of the cave target based on the features enhanced by both spatial and channel attention.
[0060] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for estimating the total methane emission flux from benthic animal burrows in coastal mudflats based on UAV remote sensing and deep learning, characterized in that... Includes the following steps: During low tide, drones are used to conduct aerial surveys of the target tidal flat area to obtain raw remote sensing data for constructing ultra-high resolution images that can distinguish the morphology of benthic animal burrows. The raw remote sensing data is processed to generate a digital orthophoto covering the target tidal flat area; The digital orthophoto image is segmented using a sliding window strategy to generate multiple image subsets that conform to the input size of the deep learning model. Each of the aforementioned image subsets is input into a trained deep learning detection model, which integrates an attention mechanism to enhance cave feature extraction capabilities, and outputs cave detection results for each image subset. Based on georeferenced information, the cave detection results of each image subset are fused in a unified coordinate system, and duplicate detections in the fused results are deduplicated to obtain the total number of benthic animal caves in the target tidal flat area. Based on the total number of caves and the average methane emission flux per cave obtained through field sampling, the total methane emission flux from benthic animal caves in the target tidal flat area is estimated.
2. The method according to claim 1, characterized in that, The flight altitude, heading overlap, and lateral overlap of the aerial survey are configured to ensure that the ground sampling distance of the original remote sensing data reaches the sub-centimeter level.
3. The method according to claim 1, characterized in that, The sliding window strategy uses a window with a predetermined overlap rate to traverse and cut the digital orthophoto. The size of the window is set according to the input requirements of the deep learning detection model, and the predetermined overlap rate is configured to ensure that cave targets located at the edge of the slice are completely detected.
4. The method according to claim 1, characterized in that, The backbone network of the deep learning detection model integrates a spatial attention module and a channel attention module; the spatial attention module is used to enhance the extraction of cave edge contour features, and the channel attention module is used to enhance the extraction of features in the cave center region.
5. The method according to claim 4, characterized in that, The spatial attention module generates spatial attention weights by fusing average pooling and max pooling features and performing convolution processing; the channel attention module processes average pooling and max pooling features using a shared multilayer perceptron and fusing them to generate channel attention weights.
6. The method according to claim 1, characterized in that, Based on georeferenced information, the cave detection results of each image subset are fused in a unified coordinate system. Specifically, this includes converting the pixel coordinates of the detection results in each image subset into global geographic coordinates based on the ground sampling distance and reference point geographic coordinates of the digital orthophoto.
7. The method according to claim 1, characterized in that, The duplicate detections in the fusion results are deduplicated, specifically by calculating the overlap index between multiple detection results located in the overlapping areas of different image subsets, and removing redundant detection results with an overlap index higher than a preset threshold.
8. The method according to claim 1, characterized in that, The total methane emission flux from benthic animal burrows in the target tidal flat area is estimated using the following formula: F_total = N × f_avg; Wherein, F_total is the total methane emission flux from benthic animal burrows in the target tidal flat area, N is the total number of burrows, and f_avg is the measured average methane emission flux per burrow.
9. A system for estimating total methane emission flux from benthic animal burrows in coastal mudflats based on UAV remote sensing and deep learning, characterized in that, include: The aerial survey module is used to control UAVs to conduct aerial surveys of target tidal flat areas during low tide and acquire raw remote sensing data for constructing ultra-high resolution images that distinguish the individual morphology of benthic animal burrows. The image processing module is used to process the raw remote sensing data to generate a digital orthophoto covering the target tidal flat area; The segmentation module is used to segment the digital orthophoto using a sliding window strategy to generate multiple image subsets that conform to the input size of the deep learning model. The detection module has a built-in trained deep learning detection model, which integrates an attention mechanism to enhance the ability to extract cave features. It is used to detect caves on a subset of the input image and output the detection results. The fusion counting module is used to fuse the cave detection results of each image subset in a unified coordinate system based on geographic reference information, remove duplicate detections in the fusion results, and count the total number of benthic animal caves in the target tidal flat area. The estimation module is used to estimate the total methane emission flux from benthic animal burrows in the target tidal flat area based on the total number of burrows and the preset average methane emission flux per burrow.
10. The system according to claim 9, characterized in that, The deep learning detection model in the detection module includes: The backbone network is used to extract deep features from a subset of the input image. The spatial attention submodule is used to perform spatial dimension weighting on the depth features to enhance the response of the cave edge contour features; The channel attention submodule is used to perform channel-dimensional weighting on the spatially weighted features to enhance the response of features in the central region of the cave. The detection head is used to output the detection bounding box and confidence score of the cave target based on the features enhanced by both spatial and channel attention.