Field-ditch-pond extraction method based on unmanned aerial vehicle photogrammetry and deep learning

Through the method of combining drone photogrammetry and deep learning, a U-net algorithm model with built-in attention mechanism is built, and the image data of the small watershed field-gul-tang system is fused with multiple features, which solves the problems of low extraction efficiency and insufficient accuracy in the existing technology, and achieves high-precision field-gul-tang system extraction.

CN120388307AActive Publication Date: 2025-07-29CHUZHOU UNIV

Patent Information

Application Number
CN202510468377.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-29
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

The prior art has problems such as strong artificial dependence, low extraction efficiency and insufficient accuracy in the extraction of small watershed field-ditch-pool system, especially in complex terrain and multiple types of farmland.

Method used

Using a combination of drone photogrammetry and deep learning, image data is obtained through tilt photogrammetry, a U-net algorithm model with built-in attention mechanism is built, spectral, terrain and texture features are fused, and multi-scale segmentation is combined with ISODATA clustering and mountain shadowing tools to generate high-precision field-groove-tang vector data.

Benefits of technology

The extraction accuracy and efficiency of the small watershed field-ditch-pool system is improved, and it can adapt to complex terrain and multiple types of farmland, and provides high-precision data to support the simulation of agricultural non-point source pollution migration trajectory and fine farmland management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388307A_ABST
    Figure CN120388307A_ABST
Patent Text Reader

Abstract

The invention discloses a field-ditch-pond extraction method based on unmanned aerial vehicle photogrammetry and deep learning, and belongs to the technical field of remote sensing technology and agricultural informationization, and the method comprises the steps: employing an unmanned aerial vehicle to carry a multispectral camera and an optical lens, and obtaining original image data through an oblique photogrammetry technology; preprocessing the obtained original image data to generate a multispectral image map (DOM) and a digital surface model (DSM); starting from the spectrum, terrain and texture features of a small watershed field-ditch-pond system, a U-net algorithm with a built-in attention mechanism, an image segmentation model supporting zero sample learning, a multi-scale segmentation technology and a multiple extraction framework of a shape feature extraction algorithm are constructed, image segmentation, classification and interpretation are carried out, and a multi-scale image segmentation model is constructed. Fine extraction of the small watershed field-ditch-pond system is realized; the design can improve the extraction precision and efficiency under complex terrains, and is suitable for current small watershed agricultural non-point source pollution migration trajectory simulation and farmland fine management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of remote sensing technology and agricultural informatization technology, and particularly relates to a method for extracting fields, ditches and ponds by using unmanned aerial vehicle (UAV) photogrammetry and deep learning. Background Art

[0002] The small watershed field-ditch-pond system (fields, ditches and ponds) is an important part of the agricultural landscape and a key channel for the migration of non-point source pollution in the agricultural area, and is of great significance for agricultural production, irrigation and water environment monitoring. Traditional methods for extracting the small watershed field-ditch-pond system mainly rely on manual delineation and ground surveys, which are time-consuming and laborious, and it is difficult to achieve large-scale and high-precision automatic extraction. With the development of remote sensing technology, methods for extracting large-scale watershed fields, rivers and waters based on satellite remote sensing have gradually been applied and popularized. However, limited by the acquisition of high-precision remote sensing data and the multi-target recognition algorithm technology under complex ground object scenes, there are still problems in the small watershed such as incomplete and inaccurate extraction of the boundaries of small ground objects, poor extraction accuracy of small ditches, and spectral mixing.

[0003] Due to the characteristics of high resolution, strong flexibility, and timely data acquisition, UAV photogrammetry technology has gradually become an important means for small-scale remote sensing monitoring; UAVs are equipped with multiple types of lenses, and through oblique photogrammetry, they can provide high-precision and rich ground object spectral information and terrain information, and have unique advantages in ground object classification and target extraction; however, existing extraction methods mostly separate the spectral, texture or terrain information of ground objects, resulting in limited extraction accuracy and efficiency of ground objects. At the same time, existing algorithms have not provided a systematic solution for the fine extraction of the field-ditch-pond system. Especially in the face of complex terrain and multiple types of farmland, the extraction effect is still not ideal.

[0004] Aiming at the deficiencies of the existing technology, the present invention provides a method for extracting fields, ditches and ponds by using UAV photogrammetry and deep learning, aiming to solve the above problems. Summary of the Invention

[0005] The purpose of the present invention is to overcome the deficiencies in the existing technology and provide a method for extracting fields, ditches and ponds by using UAV photogrammetry and deep learning, which can solve the problems of strong manual dependence, low extraction efficiency and insufficient accuracy under complex terrain in the existing technology.

[0006] To achieve the above object, in a first aspect, the present invention provides a method for extracting fields, ditches and ponds by using UAV photogrammetry and deep learning, the method comprising:

[0007] S1: Using a UAV equipped with a multispectral camera and an optical lens, and obtaining original image data through oblique photogrammetry technology;

[0008] S2: Preprocess the acquired original image data to generate a digital orthophoto map (DOM) and a digital surface model (DSM).

[0009] S3: Construct a U-net algorithm model with an embedded attention mechanism, and combine multi-feature fusion technology to train the preprocessed image data. The fused features include spectral features, terrain features, and texture features.

[0010] S4: Use the trained U-net algorithm model to extract water ponds from the image data of the target area and generate vector data of the water ponds.

[0011] Classify the preprocessed multi-spectral data using the ISODATA clustering algorithm to extract data of the areas to be extracted for fields; use an image segmentation model supporting zero-shot learning to perform image segmentation on the data of the areas to be extracted for fields and generate vector data of the fields.

[0012] Use the hillshade tool combined with DSM data to extract the ridge areas, and perform multi-scale segmentation and shape feature extraction to generate vector data of the ditches.

[0013] S5: Perform spatial overlay analysis on the vector data of the water ponds, fields, and ditches, check the topological consistency, and generate the final vector data of fields-ditches-ponds.

[0014] Combined with the first aspect, the U-net algorithm model with an embedded attention mechanism described in S3 includes:

[0015] An encoder module, adopting a five-level downsampling structure. Each level consists of two groups of depthwise separable convolutions with a kernel size of 3×3, a stride of 1, in cooperation with a ReLU activation function, and an SE (Squeeze-and-Excitation) channel attention module is embedded to strengthen the water body spectral features (such as the reflection difference in the near-infrared band) through adaptive channel weight adjustment.

[0016] A decoder module, adopting a five-level upsampling structure. Each level realizes feature recovery through bilinear interpolation and transposed convolution, and performs cross-layer skip connections with the feature maps of the corresponding levels of the encoder module. A dynamic gating mechanism is introduced at the connection, and the Sigmoid function is used to control the feature fusion weight; and the decoder module innovatively integrates a multi-scale dilated convolution pyramid module (MDCP) in the fourth-level decoding layer, which includes four groups of parallel dilated convolution layers, and the output features are input into a 1×1 convolution for multi-scale context feature fusion after channel splicing.

[0017] In combination with the first aspect, a 2×2 max-pooling layer is set at the end of each stage in the encoder module, the size of the feature map is gradually reduced, and the number of channels is doubled to 512 at the same time.

[0018] In combination with the first aspect, the multi-feature fusion technology in S3 includes:

[0019] Performing spectral feature extraction on multi-spectral data, including calculating the reflectance of the blue band, green band, red band, red-edge band, and near-infrared band;

[0020] Performing terrain feature extraction on DSM data, including calculating slope, curvature, and aspect;

[0021] Performing texture feature extraction on image data, including calculating texture intensity and texture direction.

[0022] In combination with the first aspect, the U-net algorithm model is used to label water bodies through the LabelMe tool, and the bicubic interpolation algorithm is used to perform size normalization processing on the image data.

[0023] In combination with the first aspect, the multi-scale segmentation technology uses the hillshade tool in ArcGIS Pro, sets the solar altitude angle to 50°, and extracts hillshades at multiple solar azimuth angles to enhance the shadow features of the image.

[0024] In the second aspect, the present invention provides a field-ditch-pond extraction system for unmanned aerial vehicle photogrammetry and deep learning; the system includes:

[0025] A data acquisition module for controlling an unmanned aerial vehicle equipped with a multi-spectral camera and an optical lens to obtain the original image data of the target area through oblique photogrammetry technology;

[0026] A data preprocessing module for preprocessing the obtained original image data to generate a multi-spectral image map (DOM) and a digital surface model (DSM);

[0027] A model training module for constructing a U-net algorithm model with an internal attention mechanism and training the preprocessed image data in combination with multi-feature fusion technology, where the fused features include spectral features, terrain features, and texture features;

[0028] A data extraction module for:

[0029] Using the trained U-net model to extract ponds from the image data of the target area to generate vector data of the ponds;

[0030] The ISODATA clustering algorithm is used to classify the preprocessed multi-spectral data, and the data of the area to be extracted in the field block is extracted; an image segmentation model supporting zero-shot learning is used to perform image segmentation on the data of the area to be extracted in the field block to generate vector data of the field block;

[0031] The hill shadow tool is used in combination with DSM data to extract the ridge area, and multi-scale segmentation and shape feature extraction are performed to generate vector data of the field ditch;

[0032] A data integration module is used to perform spatial overlay analysis on the vector data of the pond, field block, and field ditch, check topological consistency, and generate the final field-ditch-pond vector data.

[0033] Combined with the second aspect, the data extraction module further includes:

[0034] A pond extraction unit is used to handle the problem of tree occlusion, and the pond boundary is repaired through the convex hull algorithm and curvature calculation;

[0035] A field block extraction unit is used to handle the boundary effect of sliding cropping, and the influence of the boundary effect is reduced through weight assignment;

[0036] A field ditch extraction unit is used to handle the misclassification problem in the ridge area, and the misclassified area is removed through maximum bounding rectangle analysis and threshold method.

[0037] In the third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.

[0038] In the fourth aspect, the present invention provides a device, including:

[0039] A memory for storing instructions;

[0040] A processor for executing the instructions, so that the device executes the steps of the method described in the first aspect.

[0041] Compared with the prior art, the beneficial effects achieved by the present invention:

[0042] 1. In the data acquisition module of the present invention, a multi-spectral camera and an optical lens are carried by a drone to obtain high-resolution original image data, providing a high-quality data basis for subsequent fine extraction. The data is preprocessed through geometric correction, radiometric correction, and aerial triangulation measurement. In the model training module, by combining spectral features, terrain features, and texture features, training is performed through a U-net algorithm model with an built-in attention mechanism, significantly improving the model's perception ability of complex ground objects. Through multi-feature fusion technology, the model can more accurately identify and extract ponds, fields, and field ditches, improving the extraction accuracy. In the data extraction module, vector data of ponds is extracted through the U-net model, vector data of fields is extracted through the ISODATA clustering algorithm and an image segmentation model, and vector data of field ditches is extracted through the hill shade tool and multi-scale segmentation technology. Finally, all the data is subjected to spatial overlay analysis and topological consistency check to generate the final field-ditch-pond vector data, providing high-precision data support for subsequent simulation of agricultural non-point source pollution migration trajectories and fine management of farmland. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is a flowchart of the extraction method of the present invention.

[0044] Figure 2 is the multi-spectral image map (DOM) of the experimental area in the embodiment of the present invention.

[0045] Figure 3 is the digital surface model (DSM) of the experimental area in the embodiment of the present invention.

[0046] Figure 4 is the extraction result map of the experimental area in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and should not be used to limit the protection scope of the present invention.

[0048] Embodiment 1

[0049] Refer to Figures 1 - 4 , the present invention provides a method for extracting fields, ditches, and ponds by using drone photogrammetry and deep learning. The method includes:

[0050] S1: Use a drone to carry a multi-spectral camera and an optical lens, and obtain original image data through oblique photogrammetry technology;

[0051] S2: Preprocess the obtained original image data to generate a multi-spectral image map (DOM) and a digital surface model (DSM); The multi-spectral image map (DOM) specifically refers to Figure 2, the digital surface model (DSM) can be specifically referred to Figure 3 ;

[0052] S3: Construct a U-net algorithm model with an built-in attention mechanism, and combine multi-feature fusion technology to train the preprocessed image data, where the fused features include spectral features, terrain features, and texture features;

[0053] S4: Use the trained U-net algorithm model to extract ponds from the image data of the target area and generate vector data of the ponds;

[0054] Adopt the ISODATA clustering algorithm to classify the preprocessed multi-spectral data and extract the data of the area to be extracted for the fields; use an image segmentation model supporting zero-shot learning to perform image segmentation on the data of the area to be extracted for the fields and generate vector data of the fields;

[0055] Use the hill shade tool combined with DSM data to extract the ridge area, and perform multi-scale segmentation and shape feature extraction to generate vector data of the ditches;

[0056] S5: Perform spatial overlay analysis on the vector data of the ponds, fields, and ditches, check the topological consistency, and generate the final field-ditch-pond vector data.

[0057] Furthermore, the preprocessing described in S1 includes: geometric correction, radiometric correction, and aerial triangulation.

[0058] Specifically, in this embodiment, the U-net algorithm model with an built-in attention mechanism described in S3 includes:

[0059] An encoder module, adopting a five-level downsampling structure, each level consists of two groups of depthwise separable convolutions, with a kernel size of 3×3, a stride of 1, combined with a ReLU activation function, and embedded with an SE (Squeeze-and-Excitation) channel attention module to strengthen the water body spectral features (such as the reflection difference in the near-infrared band) through adaptive channel weight adjustment;

[0060] The decoder module adopts a five-level upsampling structure. At each level, feature recovery is achieved through bilinear interpolation and transposed convolution, and cross-layer skip connections are made with the feature maps of the corresponding levels of the encoder module. A dynamic gating mechanism is introduced at the connection, and the Sigmoid function is used to control the feature fusion weight. Moreover, the decoder module innovatively integrates a multi-scale dilated convolution pyramid module (MDCP) in the fourth-level decoding layer, which includes four groups of parallel dilated convolution layers (with dilation rates of 1, 3, 5, and 7 respectively). The output features are input into a 1×1 convolution for multi-scale context feature fusion after channel concatenation.

[0061] It should be noted that in this embodiment, a 2×2 max pooling layer is set at the end of each level in the encoder module, the size of the feature map is gradually reduced, and the number of channels is doubled to 512.

[0062] Furthermore, the multi-feature fusion technology in S3 includes:

[0063] Spectral feature extraction is performed on the multi-spectral data, including the calculation of reflectance in the blue band, green band, red band, red edge band, and near-infrared band.

[0064] Terrain feature extraction is performed on the DSM data, including the calculation of slope, curvature, and aspect.

[0065] Texture feature extraction is performed on the image data, including the calculation of texture intensity and texture direction.

[0066] Furthermore, the U-net algorithm model is used to label water bodies through the LabelMe tool, and the bicubic interpolation algorithm is used to perform size normalization processing on the image data.

[0067] Furthermore, the multi-scale segmentation technology uses the hillshade tool in ArcGIS Pro, sets the solar elevation angle to 50°, and extracts hillshades at multiple solar azimuth angles to enhance the shadow features of the image.

[0068] It should be noted that in this embodiment, the data acquisition / obtaining and data preprocessing in S1 and S2 specifically include the following steps:

[0069] For oblique photogrammetry, a DJI M350 multi-rotor unmanned aerial vehicle (UAV) is used, equipped with a Red Edge-MX multispectral camera and a five-lens gimbal for ground data collection. The multispectral image data covers the blue band, green band, red band, red edge band, and near-infrared band. The oblique aerial photography flight is set with a forward overlap of 80%, a side overlap of 70%, and a flight altitude of 200 meters. During the shooting process, a photo is taken every 2.5 seconds to ensure the consistency of the image data and flight efficiency, and to ensure efficient coverage of the target area. Through these parameters, high-quality and high-coverage original image data can be obtained.

[0070] Since the ground radiation energy received by the spectral camera may be affected by factors such as the photoelectric noise of the sensor itself, the vignetting effect of the camera, the change of the solar altitude angle, and the aerosol optical thickness, there is a certain difference between the collected radiation value and the actual radiation value of the ground object. Therefore, photos of the radiation calibration board are taken before and after the UAV flight, and the reflectivities of the calibration board are set to 25%, 50%, and 75% respectively. Through these calibration board images, radiometric calibration is carried out to eliminate the radiation errors caused by various factors and ensure the accuracy of the subsequent image data.

[0071] After the image data collection is completed, the DJI Zhitu software is used to perform geometric correction, radiometric correction, and aerial triangulation on the image data collected by the UAV, and centimeter-level multispectral and RGB orthophoto maps (DOM) and digital surface models (DSM) are generated respectively.

[0072] It should be noted that in this embodiment, the training of the U-net model in S3 specifically includes the following steps:

[0073] Select 500 RGB original aerial photos of the study area and its surrounding areas, ensure that the image resolution is not lower than 0.5 meters / pixel, and the storage format is GeoTIFF to retain the geographic coordinate information. Use LabelMe9 (an open-source image annotation tool mainly used for data preparation in the fields of computer vision and machine learning) to annotate the water bodies, and save the annotation results as a JSON file in the form of polygon vectors, and then convert them into binary mask data (the water body value is 255 and the background value is 0). When annotating, all types of water bodies (such as rivers, lakes, ponds, etc.) need to be covered. First, determine the equal-proportion scaling benchmark according to the maximum side length, construct a square blank canvas and fill it with zeros, and map the original image data to the corresponding area of the canvas in the center; use the bicubic interpolation algorithm to perform high-precision size normalization processing on the filled image to achieve distortion-free scaling of the image content to 512×512 pixels. Use the GDAL function to read the EXIF metadata of the JPEG file to obtain the position information of the picture and record it, which is convenient for stitching the predicted images.

[0074] Among them, in this embodiment, the dataset is also divided and enhanced, that is: 500 images are proportionally divided into a training set (400 images), a validation set (80 images), and a test set (20 images). Ensure uniform distribution of different terrains (such as mountains and plains) during division. Geometric transformation, radiometric transformation, and noise injection are performed on the training data to improve the robustness of the model.

[0075] During data training, the Adam optimizer is used, with an initial learning rate of 1e-4, decaying by 50% every 20 epochs, and a batch size of 8; combining weighted cross-entropy loss and Dice Loss to handle class imbalance and improve boundary detection accuracy; with a batch size of 8 and 100 training epochs, the test set needs to achieve IoU≥0.87 and Recall≥0.91, and the false detection rate for shadow areas needs to be less than 5%;

[0076] After training, the data is predicted: both the original training data and the annotated data are scaled to 512×512 pixel size, and the position information is retained, then prediction is performed. The generated binary file is then converted back to the original image size and the original position information read is written out to facilitate image stitching. Based on the GDAL library, the optimized binary mask is converted into a polygon vector, and the output format is Shapefile.

[0077] It should be noted that in this embodiment, the extraction of vector data of ponds, fields, and ditches in S4 specifically includes the following steps:

[0078] S41: Pond data extraction

[0079] The original aerial photos of the study area are scaled and the position information is retained, then input into the trained model for prediction. After predicting the water body data, it is scaled back to the original size, merged based on the original geographical location, and then the binary water body data is processed by raster to polygon conversion to obtain the vector data of the water body. Thus, according to the spectral characteristics, the pond extraction is completed.

[0080] S42: Field data extraction

[0081] S421: The ISODATA clustering algorithm is used to classify the multispectral data. It is classified 12 times in three iterations, and then forest land and construction land are extracted through class merging, and then post-classification processing is performed to eliminate the salt-and-pepper effect. The multispectral data of the study area is excluded from forest land, construction land, and the ponds extracted in the previous step to obtain the data of the area to be extracted for fields, reducing the influence of the subsequent image segmentation background;

[0082] S422: Based on the data of the area to be extracted in the field block, perform sliding cropping at a size of 1600×1600 pixels with a step size of 800 pixels (repetition rate 50%), save the location information of the cropped TIF file and convert it to JPG format; use an image segmentation model that supports zero-shot learning (Segment Anything Model) for image segmentation, and automatically extract each raster instance of the image according to the spectral and texture features of the data (each field block can be clearly extracted), then convert it into a vector boundary, and then calculate the spectral values within each vector block to remove non-field-block areas, thus completing the extraction of field blocks.

[0083] S43: Extraction of field ditch data

[0084] S431: To further optimize the demarcation between the field block and the surrounding environment, the present invention uses the hillshade tool in ArcGIS Pro (a desktop GIS software widely used in spatial data management, advanced map analysis, and creation of visualization effects), sets the solar altitude angle to 50°, and extracts hillshades at multiple solar azimuth angles (45°, 135°, 225°, 305°). The four extracted hillshade bands are fused with the DSM data to enhance the shadow features of the image and improve the recognition accuracy of the field block and the ridge area.

[0085] S432: Based on the data of the area to be extracted in the field block, use eCognition software (a professional remote sensing image processing software widely used in fields such as geoscience, environmental monitoring, and natural resource management) to perform multi-scale segmentation and object-oriented classification on the image data. During this process, set the weight of the four hillshade bands to 1, the weight of the DSM band to 0.6, the scale parameter to 80, the shape parameter to 0.5, and the compactness parameter to 0.6. These parameter settings can effectively perform image segmentation, ensuring that both the field block and the ridge can be accurately extracted during the segmentation process while avoiding over-segmentation.

[0086] S433: After the segmentation is completed, select about ten ridge samples and twenty background samples, configure the nearest neighbor features, select the mean value and standard deviation as feature parameters, and then perform classification. In this step, through sample training, the ridge area is recognized and distinguished from other categories. The classification result is exported in Shapefile (shp) format, and the classification result is written into the attribute table, and the ridge area is extracted based on the classification result.

[0087] S434: Since some areas may be misclassified as ridge categories, a threshold method is used to eliminate these misclassified areas. To ensure the accuracy of the results, the WGS1984 geographic coordinate system is used for coordinate transformation, and the data is projected into the WGS1984 UTM Zone 50N projection coordinate system. Subsequently, the aspect ratio of the maximum bounding rectangle of each feature and the area ratio of the feature to the maximum bounding rectangle are calculated. According to the set threshold, the eligible ridge areas are screened out, and the misclassified areas are removed. In this example, the threshold selected is that the feature with a ratio greater than 0.3 of its own maximum bounding rectangle area and an aspect ratio less than 1.5 is regarded as a mis-extracted area and deleted.

[0088] S435: In ArcGIS Pro, the merge tool is used to merge adjacent ridge surface areas to form continuous ridge areas. By smoothing the boundaries of these areas, noise is further removed. Through manual investigation, it is known that the width of the field ditch is generally about 0.2 to 0.3 m. A 0.4 m buffer zone is generated for the extracted fields, and the erase tool is used to erase the field data and ridge data, thereby obtaining the field ditch data.

[0089] In this embodiment, the present invention can obtain high-resolution original image data by using a multi-spectral camera and an optical lens carried by a drone in the data acquisition module, providing a high-quality data basis for subsequent fine extraction, and preprocessing the data through geometric correction, radiometric correction, and aerial triangulation measurement; in the model training module, combined with spectral features, terrain features, and texture features, training is carried out through a U-net algorithm model with an internal attention mechanism, significantly improving the model's perception ability of complex ground objects. Through the multi-feature fusion technology, the model can more accurately identify and extract ponds, fields, and field ditches, improving the extraction accuracy; in the data extraction module, the pond extracts vector data through the U-net model, the field extracts appropriate data through the ISODATA clustering algorithm and the image segmentation model, and the field ditch extracts vector data through the hill shade tool and the multi-scale segmentation technology; finally, all the data is subjected to spatial overlay analysis and topological consistency check to generate the final field-ditch-pond vector data, providing high-precision data support for subsequent simulation of agricultural non-point source pollution migration trajectories and fine management of farmland.

[0090] Embodiment Two

[0091] Based on Embodiment One, the present invention also provides a field-ditch-pond extraction system for drone photogrammetry and deep learning; the system includes:

[0092] A data acquisition module for controlling a drone to carry a multi-spectral camera and an optical lens to obtain original image data of a target area through oblique photogrammetry technology;

[0093] A data preprocessing module for preprocessing the acquired raw image data to generate a multispectral image map (DOM) and a digital surface model (DSM);

[0094] A model training module for constructing a U-net algorithm model with an built-in attention mechanism and training the preprocessed image data by combining multi-feature fusion technology, where the fused features include spectral features, terrain features, and texture features;

[0095] A data extraction module for:

[0096] Using the trained U-net model to extract ponds from the image data of the target area and generate vector data of the ponds;

[0097] Adopting the ISODATA clustering algorithm to classify the preprocessed multispectral data and extract data of the areas to be extracted for fields; using an image segmentation model supporting zero-shot learning to perform image segmentation on the data of the areas to be extracted for fields and generate vector data of the fields;

[0098] Using the mountain shadow tool in combination with DSM data to extract ridge areas, and performing multi-scale segmentation and shape feature extraction to generate vector data of the ditches;

[0099] A data integration module for performing spatial overlay analysis on the vector data of the ponds, fields, and ditches, checking topological consistency, and generating final field-ditch-pond vector data.

[0100] Furthermore, the data extraction module further includes:

[0101] A pond extraction unit for dealing with the problem of tree occlusion and repairing the pond boundary through the convex hull algorithm and curvature calculation;

[0102] A field extraction unit for dealing with the boundary effect of sliding cropping and reducing the influence of the boundary effect through weight assignment;

[0103] A ditch extraction unit for dealing with the misclassification problem in the ridge area and removing the misclassified area through maximum bounding rectangle analysis and threshold method.

[0104] It should be noted that, in this embodiment, the specific operation of the pond extraction unit is: for the part of the pond that is more obscured by trees, it is selected separately and the Convex The Hull convex hull algorithm calculates its convex polygon and then determines the curvature by calculating the vector cross product of three adjacent points. A certain curvature threshold is set to extract areas with high curvature. The vertical distance from the point with high curvature value to the convex hull is calculated to exclude the naturally curved parts of the pond. Then, it is determined whether the outside of these concave areas are trees. If so, they are filled. Specifically, for each vertex, the previous point, current point and next point are taken, and the cross product of the forward vector and the backward vector is calculated. Then, the curvature value is normalized to obtain the curvature value. The point with a curvature value greater than 0.15 and a vertical distance from the point to the convex hull less than 7m is taken as the tree-occluded vertex. Then, from this point to the left and right, points with a curvature value greater than -0.1 are searched to form an arc segment. The convex hull of this arc segment is constructed and the spectral value of this convex hull is calculated to determine whether it is a tree. If so, it is merged into the pond surface. If not, it is not merged. This method can compensate for the accuracy loss caused by the loss of information due to tree occlusion, thereby improving the accuracy of pond extraction.

[0105] It should be noted that, in this embodiment, the specific operation of the field extraction unit is as follows: due to the boundary effect generated during sliding cropping, the integrity of the boundary features will be lost, so the intermediate image of the sliding cropping (for example, the first, second, and third images in the sliding cropping, the second image being the middle area between the first and third images) is given the highest weight, and then when vector merging is performed, the vector elements selected from 50% of the area in the intermediate image are the highest weighted. If they overlap with other images during merging, they are merged into the highest weighted area, thus reducing the impact of the boundary effect. For smaller areas, they are merged to the outside to reduce the impact of holes, and then for each vectorized field area, the internal spectral value is calculated, and non-field areas are eliminated by setting thresholds for each band to ensure that the areas finally retained are all actual field parts.

[0106] It should be noted that in this embodiment, the ridge extraction unit specifically operates as follows: to remove misclassified areas, the maximum bounding rectangle of each segmented area is analyzed, its aspect ratio and area percentage are calculated, and misclassified areas are removed based on a set threshold (maximum bounding rectangle aspect ratio ≤ 1.5 and element proportion of its maximum bounding rectangle area ≥ 0.3). This step effectively removes misclassified or non-compliant ridge areas, ensuring the accuracy of the extraction results. The ridge boundary is then smoothed to extract the ridge boundary.

[0107] Embodiment 3

[0108] Based on the first embodiment, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in the first embodiment are implemented.

[0109] Embodiment Four

[0110] Based on the first embodiment, the present invention further provides a device, including:

[0111] A memory for storing instructions;

[0112] A processor for executing the instructions, so that the device executes the steps of the method described in the first embodiment.

[0113] In summary, through the UAV photogrammetry technology and deep learning algorithm, combined with the multi-feature fusion technology, the present invention provides an efficient and high-precision method for extracting fields, ditches and ponds. This method not only significantly improves the extraction accuracy, but also improves the extraction efficiency, and can adapt to the extraction requirements of complex terrains and multiple types of farmlands, providing strong technical support for the treatment of agricultural non-point source pollution and the fine management of farmlands in small watersheds.

[0114] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0115] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0116] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions in the flowFigure 1 one process or multiple processes and / or blocks Figure 1 the functions specified in one block or multiple blocks.

[0117] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple processes and / or the steps of the functions specified in one block or multiple blocks.

[0118] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A method for extracting fields, ditches, and ponds by using UAV photogrammetry and deep learning, characterized in that, Including: S1: Use a drone to carry a multispectral camera and an optical lens, and obtain the original image data through oblique photogrammetry technology; S2: Preprocess the obtained original image data to generate a multispectral orthophoto map (DOM) and a digital surface model (DSM); S3: Construct a U-net algorithm model with an embedded attention mechanism, and combine multi-feature fusion technology to train the preprocessed image data, where the fused features include spectral features, terrain features, and texture features; S4: Use the trained U-net algorithm model to extract ponds from the image data of the target area and generate vector data of the ponds; Use the ISODATA clustering algorithm to classify the preprocessed multispectral data and extract the data of the area to be extracted in the field plot; use an image segmentation model supporting zero-shot learning to perform image segmentation on the data of the area to be extracted in the field plot and generate vector data of the field plot; Use the hillshade tool combined with DSM data to extract the ridge area, and perform multi-scale segmentation and shape feature extraction to generate vector data of the field ditch; S5: Perform spatial overlay analysis on the vector data of the ponds, field plots, and field ditches, check the topological consistency, and generate the final field-ditch-pond vector data.

2. The method for extracting fields, ditches and ponds by using UAV photogrammetry and deep learning according to claim 1, wherein The U-net algorithm model with an embedded attention mechanism described in S3 includes: An encoder module, adopting a five-level downsampling structure, each level consists of two groups of depthwise separable convolutions with a kernel size of 3×3, a stride of 1, combined with a ReLU activation function, and an SE channel attention module is embedded to strengthen the water body spectral features through adaptive channel weight adjustment; A decoder module, adopting a five-level upsampling structure, each level realizes feature restoration through bilinear interpolation and transposed convolution, and performs cross-layer skip connection with the feature map of the corresponding level of the encoder module. A dynamic gating mechanism is introduced at the connection, and the Sigmoid function is used to control the feature fusion weight; and the decoder module innovatively integrates a multi-scale dilated convolution pyramid module in the fourth-level decoding layer, which includes four groups of parallel dilated convolution layers, and the output features are input into a 1×1 convolution for multi-scale context feature fusion after channel concatenation.

3. The method for extracting fields, ditches, and ponds by drone photogrammetry and deep learning according to claim 2, characterized in that, A 2×2 max-pooling layer is set at the end of each level in the encoder module, the feature map size is reduced level by level, and the number of channels is doubled to 512 at the same time.

4. The method for extracting fields, ditches, and ponds by drone photogrammetry and deep learning according to claim 1, wherein The multi-feature fusion technology described in S3 includes: Extract spectral features from the multispectral data, including the calculation of the reflectance of the blue band, green band, red band, red edge band, and near-infrared band; Extract terrain features from the DSM data, including the calculation of slope, curvature, and aspect; Extract texture features from the image data, including the calculation of texture intensity and texture direction.

5. The method for extracting fields, ditches and ponds by drone photogrammetry and deep learning according to claim 1, characterized in that The U-net algorithm model is used to label water bodies through the LabelMe tool, and the bilinear interpolation algorithm is used to perform size normalization processing on the image data.

6. The method for extracting fields, ditches and ponds by drone photogrammetry and deep learning according to claim 1, characterized in that, The multi-scale segmentation technology uses the hillshade tool in ArcGIS Pro, sets the solar elevation angle to 50°, and extracts hillshades at multiple solar azimuth angles to enhance the shadow features of the image.

7. A field-ditch-pond extraction system for UAV photogrammetry and deep learning, characterized in that, Including: A data acquisition module, which is used to control the drone to carry a multispectral camera and an optical lens, and obtain the original image data of the target area through the oblique photogrammetry technology; A data preprocessing module, which is used to preprocess the obtained original image data to generate a multispectral orthophoto map (DOM) and a digital surface model (DSM); A model training module, which is used to construct a U-net algorithm model with an attention mechanism built in, and train the preprocessed image data in combination with the multi-feature fusion technology, where the fused features include spectral features, terrain features and texture features; A data extraction module, which is used for: Using the trained U-net model to extract ponds from the image data of the target area to generate vector data of the ponds; Adopting the ISODATA clustering algorithm to classify the preprocessed multispectral data to extract the data of the area to be extracted in the field plot; using an image segmentation model supporting zero-shot learning to perform image segmentation on the data of the area to be extracted in the field plot to generate vector data of the field plot; Using the mountain shadow tool in combination with the DSM data to extract the ridge area, and performing multi-scale segmentation and shape feature extraction to generate vector data of the field ditch; A data integration module, which is used to perform spatial overlay analysis on the vector data of the ponds, field plots and field ditches, check the topological consistency, and generate the final vector data of the field-ditch-pond; 8. The field-ditch-pond extraction system for drone photogrammetry and deep learning according to claim 7, characterized in that, The data extraction module further includes: A pond extraction unit, which is used to handle the problem of tree occlusion, and repair the pond boundary through the convex hull algorithm and curvature calculation; A field plot extraction unit, which is used to handle the boundary effect of sliding clipping, and reduce the influence of the boundary effect through weight assignment; A field ditch extraction unit, which is used to handle the misclassification problem in the ridge area, and eliminate the misclassified area through the maximum circumscribed rectangle analysis and threshold method; 9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method for extracting field-ditch-pond by drone photogrammetry and deep learning as described in any one of claims 1-6.

10. A device, characterized in that, It includes: A memory, which is used to store instructions; A processor, which is used to execute the instructions, so that the device executes the steps of the method for extracting field-ditch-pond by drone photogrammetry and deep learning as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Unmanned aerial vehicle image crown canopy segmentation method based on combination of morphology and mark controlling

    CN106875407A

  • Field ditch pond parameter acquisition method and device

    CN110207676A

  • Method and device for extracting cultivated land parcels based on SE-U-Net + + model

    CN114419430A

Cited By

  • Method and system for counting progress of rice transplanting operation of rice machine

    CN120892479A

  • A rice machine transplanting operation progress statistics method and system

    CN120892479B

  • Farmland drainage capacity monitoring method and device, electronic equipment and storage medium

    CN121275056A