Farmland area determination method based on boundary determination and semantic segmentation of G-Lap algorithm
By using the G-Lap algorithm and a modified Deeplab model with Swin Transformer and BiFPN networks, the method addresses precision and adaptability issues in drone-based agricultural field area calculation, achieving high precision and real-time performance.
Patent Information
- Application Number
- CN202510779943.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-12
AI Technical Summary
The existing UAV farmland area calculation method relies on fixed flight altitude and adopts regular square pixel geometric approximation models, resulting in limited calculation accuracy and insufficient universality; the semantic segmentation model has limitations in the trade-off between real-time and precision; the farmland boundaries are not obvious in complex situations, making it difficult to achieve high-precision boundary judgment, which affects the labeling accuracy and model training effect.
The G-Lap algorithm is used to enhance the farmland boundaries, combine the improved Deeplab model for semantic segmentation, enhance the farmland edges through the G-Lap algorithm and build a multi-scene training set. The backbone network of the Deeplab is improved to be the Swin Transformer and BiFPN networks, and calculate the farmland area with aerial photography altitude information.
The accuracy of farmland boundary judgment and the segmentation effect of the model are improved, and high-precision area calculations are achieved in different farmland scenarios, with higher versatility and calculation efficiency, the segmentation accuracy reaches 95.7% mIoU, and the inference speed remains 17.3 FPS.
Smart Images

Figure CN120318303A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of farmland information science, and specifically to a method for measuring farmland area based on improved feature-enhanced semantic segmentation. Background Art
[0002] With the rapid development of global technology, the production mode of agricultural workers has changed greatly, such as the emergence of smart agriculture. Smart agriculture accurately monitors and analyzes the crop growth situation, soil fertility, pest and disease conditions, and irrigation situation in farmland information through technologies such as remote sensing and the Internet of Things, thereby improving efficiency. For the calculation of farmland area, which is most needed for irrigation, traditional satellite remote sensing and manual mapping have problems such as low efficiency, high cost, and lag. Compared with satellite remote sensing and manual mapping, drones show their unique advantages in calculating the area of small and medium-sized farmland. Their centimeter-level image accuracy and operation convenience can not only reduce the acquisition cost but also achieve high-frequency dynamic monitoring. Therefore, the present invention proposes a new method for automatically calculating the farmland area based on drone aerial images.
[0003] Traditional drone area calculation methods mainly rely on image processing technology and geometric formulas, and the extraction of the calculation area in the image mainly relies on manual extraction. For example, using drones to collect images and combining with the Gaussian area formula to calculate the farmland area. The principle of such methods is relatively simple, and for complex situations, the calculation accuracy may be affected to a certain extent. For the improvement of traditional geometric methods, the size information of the target in multiple aspects in a single drone optical image can also be measured and the accuracy improved by obtaining and calculating the focal length, detector size, pitch and roll angles of the drone, and the relative elevation angle between the aerial camera and the target point. In addition to traditional image processing and geometric formulas, existing computer software such as Adobe Photoshop and ARCGIS are used to process, analyze, and calculate the area of drone images with a fixed flight height. The above existing area calculation methods rely on a fixed flight height and use a regular square pixel geometric approximation model, resulting in problems such as limited calculation accuracy and insufficient universality in practical applications, and further improvement of the area calculation method is required.
[0004] In recent years, deep learning methods in computer vision have surpassed traditional image segmentation methods in extracting high-dimensional features and have been widely used in the segmentation and extraction of corresponding regions in remote sensing images, effectively improving efficiency. The convolutional networks of deep learning usually contain many convolutional layers and can learn deep features from training data. Therefore, deep convolutional neural networks have been introduced into the segmentation task and have achieved good results. The methods for extracting target regions in remote sensing images through deep learning can be divided into two categories: semantic segmentation and instance segmentation. Semantic segmentation refers to classifying each pixel in an image, and instance segmentation can be regarded as an extension of semantic segmentation. However, compared with semantic segmentation, instance segmentation requires a large number of training images, and there are not many adjacent farmland areas in the complete farmland parcels in drone images. Therefore, semantic segmentation can be used to extract the pixels of the complete farmland parcel area in the image and calculate the farmland area.
[0005] In terms of model selection, common semantic segmentation models include U-Net, DeepLab, SegFormer and their improved models, etc. Research shows that an adaptive reconnaissance strategy combined with edge computing can achieve accurate identification of rice lodging areas. For typical scenarios such as tobacco planting in mountainous areas of the plateau and plastic film farmland, the U-Net series of models show significant advantages in small sample learning and complex ground object recognition. Combining drone multi-temporal images with an optimized SegFormer model, a dynamic monitoring method is proposed to effectively achieve mountainous cultivated land change monitoring and farmland distribution extraction. The above work content shows that semantic segmentation models have made remarkable progress in drone farmland image recognition and segmentation, but there are still some limitations, such as the trade-off between real-time performance and accuracy. Lightweight models meet the real-time requirements, but the segmentation accuracy is lower than that of complex models;
[0006] In the process of making the dataset of the semantic segmentation model, it is necessary to manually determine and label the feature regions. For farmland images taken by drones, when the drone takes pictures of farmland images, due to complex farmland environments, light changes, different shooting angles, etc., the farmland boundaries in the images are not obvious and it is difficult to achieve high-precision boundary determination. This problem will affect the subsequent annotation accuracy, the training effect of the model and the area calculation accuracy. However, this problem and its solution have not been considered in existing research. Summary of the Invention
[0007] The technical problem to be solved by the present invention is that the existing area calculation method relies on a fixed flight altitude and adopts a regular square pixel geometric approximation model, resulting in limited calculation accuracy and insufficient universality in practical applications; the limitations of the semantic segmentation model in the trade-off between real-time performance and accuracy; in complex situations, the farmland boundaries in aerial images are not obvious, making it difficult to achieve high-precision boundary determination, thus leading to a decrease in annotation accuracy. The present invention provides a method for measuring the area of farmland based on the G-Lap algorithm for boundary determination and semantic segmentation.
[0008] To achieve the above object, the specific technical solution adopted by the present invention is as follows:
[0009] A method for measuring the area of farmland based on the G-Lap algorithm for boundary determination and semantic segmentation, comprising the following steps:
[0010] Step 1: Collect images of agricultural regions in multiple scenarios through an unmanned aerial vehicle, and then construct a multi-scenario farmland image set;
[0011] Step 2: Further enhance the edges of the farmland in the multi-scenario farmland image set in Step 1 through the G-Lap algorithm, and determine the farmland boundaries, so as to achieve high-precision boundary determination and annotation of the farmland, and further construct a multi-scenario farmland training image set;
[0012] Step 3: Construct and train a semantic segmentation model. The semantic segmentation model is based on the Deeplab model as the basic network. In the backbone network of the Deeplab model, the feature extraction network Swin Transformer is used to replace the original ResNet, the FPN network of Deeplab is replaced by the BiFPN network, and the training image set is used to train the frozen backbone network to obtain an initial model;
[0013] Step 4: Use an unmanned aerial vehicle to collect unmanned aerial vehicle farmland images of the farmland area to be measured, establish a test image set, and process the test image set in the manner of multi-scale edge enhancement in Step 2. The test image is embedded with aerial photography altitude information;
[0014] Step 5: Calculate the actual area corresponding to each pixel in each test image at different aerial photography altitudes through an improved unmanned aerial vehicle area calculation method;
[0015] Step 6: Use the trained segmentation model to segment and extract the pixels of the farmland area in the test image set, and statistically calculate the number of pixels segmented by the segmentation model through the image binarization method. According to the number of pixels, combined with the actual area corresponding to each pixel at different aerial photography altitudes, the farmland area is obtained.
[0016] Preferably, in the step 1, the constructed multi-scenario farmland image set is farmland under rice, wheat, fallow, and other farm crop scenarios, and the image has the edge features of a complete farmland parcel.
[0017] Preferably, in the step 1, the collected images are three-channel RGB images.
[0018] Preferably, in the step 2, the method for boundary determination of the farmland edge by the G-Lap algorithm is as follows:
[0019] Apply the Laplacian operator to the multi-scenario farmland image set in step 1 to enhance the boundary of the complete farmland parcel. Subsequently, using the image enhanced by the Laplacian operator as the guiding image, adopt the guided filtering method to further enhance the boundary of the original farmland RGB image, and at the same time suppress the crop texture noise. Based on the enhanced boundary, construct connected regions and fill the closed regions as the annotation regions, so as to achieve high-precision boundary determination and annotation of the farmland.
[0020] Preferably, in the step 2, the method for using the Laplacian operator to enhance the boundary of the complete farmland parcel is as follows:
[0021] First, define the Laplacian operator of the two-dimensional image as a discrete convolution kernel:
[0022] ;
[0023] where represents the Laplacian operator of the function, and respectively represent the second-order partial derivatives of the function in the and directions;
[0024] Use the 3×3 extended diagonal convolution kernel of to perform convolution operation on the original image ;
[0025] Subsequently, superimpose the edge response map and the original image proportionally to enhance the high-frequency edge information;
[0026] Finally, perform normalization output. Since the convolution result may be negative, the output image needs to be normalized to the preset range.
[0027] Preferably, the guided filtering method is as follows:
[0028] ;
[0029] ;
[0030] ;
[0031] Among them, i represents the pixel subscript, and are the linear coefficients within the window, guiding the image to be the image enhanced by the Laplacian operator, is the mean value of the guiding image within the window ; is the variance of the guiding image within the window ; is the mean value of the input original image within the window ; is the regularization parameter to prevent from being too large.
[0032] Preferably, in step 2, the method for boundary determination of the original farmland RGB image through guided filtering is as follows:
[0033] In the boundary region of the guiding image, when , increases, the high-frequency details of the input image are retained. In the boundary region of the guiding image, when , decreases, , and the noise of the input image is suppressed and smoothed; the farmland boundary is determined according to the distribution of , that is, , is the threshold, which is the high-value region of the farmland boundary, that is, the region of the farmland boundary, where the threshold is determined by the following formula:
[0034]
[0035] Among them, μ edge and are respectively the distribution mean and standard deviation of .
[0036] Preferably, in step 2, the boundary of the farmland is labeled: based on the high-value region of the farmland boundary, through a combination of morphological dilation and erosion operations, effective connection of the boundary is achieved to form a closed region, and the labeled region can be obtained by filling the closed region. Filling the closed region serves as the labeled region, thereby realizing the boundary labeling of the farmland and further constructing a multi-scene farmland training image set.
[0037] Preferably, in step 3, an atrous spatial pyramid pooling is introduced after the BiFPN network, and the dilation rate of the atrous convolution in the atrous spatial pyramid pooling is optimized.
[0038] Preferably, the method for optimizing the dilation rate of the atrous convolution is as follows: a dilation rate sequence is dynamically designed according to the size of the feature map output by the atrous spatial pyramid pooling, and by the formula:
[0039] ,
[0040] For a 384×512 feature map, the dilation rate is adjusted to the combination of [1, 6, 12, 24].
[0041] Preferably, two parallel network branches are introduced into the atrous spatial pyramid pooling, and the two parallel network branches are a Transformer branch and a pooling layer branch respectively.
[0042] Preferably, in step 5, the method for calculating the area of the drone is as follows:
[0043] ;
[0044] ;
[0045] Among them, is the total area, is the actual area corresponding to each pixel, , are the numbers of pixels in the , directions of the image respectively, is the focal length, is the height of the camera from the ground, is half of the width of the camera negative.
[0046] The present invention also provides a farmland area measurement system based on the G-Lap algorithm for boundary determination and semantic segmentation, including: a processor and a storage device; the processor loads and executes the instructions and data in the storage device to implement a method for farmland area measurement based on the G-Lap algorithm for boundary determination and semantic segmentation.
[0047] The present invention has the following characteristics and beneficial effects:
[0048] In view of the fact that the existing area calculation method relies on a fixed flight altitude and adopts a regular square pixel geometric approximation model, an improved area calculation method based on the camera imaging principle is proposed. An area calculation algorithm adapted to altitude changes is established through a theoretical pixel dynamic adjustment model. To further improve the accuracy of farmland boundary determination, the G-Lap algorithm is used to further enhance the farmland boundary, and the farmland boundary is determined according to the distribution of ak. On this basis, a connected domain is constructed, and the closed area is filled as the annotation area. And the improved Deeplab is used as the segmentation model, and the improved model is trained by the method of frozen training. The improved Deeplab has the advantages of simple structure, small demand for the training set, and remarkable segmentation effect. By testing the drone altitude data embedded in the images, the actual area corresponding to the pixels of each test image is obtained in combination with the improved area calculation method. After the improved Deeplab segmentation and extraction, the farmland pixels belonging to the calculated area in each drone farmland image can be obtained. Finally, the number of each pixel is statistically obtained through binarization, and the area of the farmland in each test image is obtained. The results obtained by this method are not less accurate than the traditional calculation method, and the calculation efficiency is improved, with higher versatility. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 is the flowchart of the farmland area measurement method based on the boundary determination and semantic segmentation of the G-Lap algorithm of the present invention;
[0050] Figure 2 is the schematic diagram of the improved area algorithm combined with the drone altitude;
[0051] Figure 3 is the flowchart of the traditional image processing extraction method;
[0052] Figure 4 is the schematic diagram of the improved Deeplab;
[0053] Figure 5 is the flowchart of the example of calculating the farmland area of the drone image based on the improved Deeplab;
[0054] Figure 6 is the qualitative experimental comparison diagram of various semantic segmentation algorithms;
[0055] Figure 7 is the quantitative index comparison diagram of each semantic segmentation algorithm;
[0056] Figure 8 is the comparison diagram of the area calculated by the present invention, the traditional method and the true value; DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] The present invention will be described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0058] Embodiment 1
[0059] This embodiment provides a method for measuring farmland area based on G-Lap algorithm for boundary determination and semantic segmentation, as Figure 1 shown, including the following steps:
[0060] Step 1: Use a drone to collect images of a specified agricultural area, and the collected images are three-channel RGB images.
[0061] Step 2: First, select the collected images constructed in Step 2, delete the images without complete farmland parcels, and then, through the G-Lap algorithm, further enhance the farmland edges in the multi-scene farmland image set in Step 1, and determine the farmland boundary according to the distribution of a k . On the basis of enhancing the boundary, construct a connected domain, fill the closed area as the annotation area. Thus, high-precision boundary determination and annotation of the farmland are realized, and a multi-scene farmland training image set is further constructed.
[0062] Specifically, the G-Lap algorithm: Use the Laplacian operator to enhance the boundary of the complete farmland parcel in the multi-scene farmland image set in Step 1, and then, with the image enhanced by the Laplacian operator as the guiding image, adopt the method of guided filtering to further enhance the boundary of the original farmland RGB image while suppressing the crop texture noise.
[0063] Edge enhancement of the original image farmland parcel by the Laplacian operator: The Laplacian operator of a two-dimensional image can be approximated as the following discrete convolution kernel:
[0064]
[0065] where represents the Laplacian operator of the function, and respectively represent the second-order partial derivatives of the function f in the x and y directions. Use the 3×3 extended diagonal convolution kernel to perform convolution operation on the original image O(x,y) to obtain the edge response map L(x,y):
[0066]
[0067] where * represents the convolution operation, Indicates the Laplacian extended diagonal convolution kernel. Subsequently, the edge response map is superimposed with the original image in proportion to enhance the high-frequency edge information:
[0068]
[0069] where the parameter controls the edge enhancement intensity and is usually taken as . Finally, normalization output is performed. Since the convolution result may be negative, the output image needs to be normalized to a reasonable range (such as 0~255):
[0070]
[0071] The process of guided filtering is as follows. For an input image , through the guidance image , the output image is obtained after filtering, where and are both inputs of the algorithm. Guided filtering defines a linear filtering process as shown below. Assume that the output image and the guidance image are linearly related within the local window :
[0072]
[0073] where i represents the pixel subscript, and are the linear coefficients within the window , and the guidance image is the image enhanced by the Laplacian operator.
[0074] For and , they are solved by the least squares method:
[0075]
[0076]
[0077] where is the mean of the guidance image within the window , is the variance of the guidance image within the window , is the mean of the input original image within the window , is the regularization parameter to prevent from being too large (to suppress noise amplification).
[0078] In the boundary region of the guiding image, when , increases, the high-frequency details of the input image are retained. In the boundary region of the guiding image, when , decreases, , the noise of the input image is suppressed and smoothed; according to 's distribution, the farmland boundary is determined, that is , is the threshold, which is the high-value area of the farmland boundary, that is, the area of the farmland boundary, where the threshold is determined by the following formula:
[0079]
[0080] where μ edge and are respectively 's distribution mean and standard deviation.
[0081] Furthermore, on the basis of enhancing the boundary and determining the farmland boundary, through the combination of morphological dilation and erosion operations, the effective connection of the boundary is realized to form a closed area. Subsequently, the labeled area can be obtained by filling the closed area. Filling the closed area serves as the labeled area. Thus, the determination and labeling of the farmland boundary are realized. In this embodiment, thus, a multi-scene farmland training image set is further constructed.
[0082] Specifically, let the predefined shape set (such as rectangle, circle), there is a high-value area edge point set , the spacing , the dilation operation is to expand the boundary and connect the breaks; the erosion operation is to retain the connection and remove the noise. Subsequently, the labeled area can be obtained by filling the closed area. Filling the closed area serves as the labeled area, thus realizing the determination and labeling of the farmland boundary, and further constructing a multi-scene farmland training image set.
[0083] It can be understood that the edge of the high-value area is actually not an ideal line, and it is composed of pixel points.
[0084] It should be noted that in this embodiment, different from the publicly available datasets used in previous studies, the datasets used in this study are all high-resolution drone images taken manually by experimenters in real farmlands. There are obvious and relatively complete farmland contours in these dataset images. The constructed multi-scene farmland images are of farmlands under rice, wheat, fallow, and other farm crop scenarios, and there are edge features wrapped by complete farmlands in the images.
[0085] Step 3: Construct a semantic segmentation model. Specifically, improve it using the Deeplab algorithm, and use the multi-scenario farmland training image set in Step 2 to train the frozen backbone to obtain an initial model. For the improvement of the Deeplab model, in this embodiment, improve the backbone network of Deeplab, use the more powerful feature extraction network SwinTransformer to replace the original ResNet to enhance the model's representation ability. Replace the FPN network of Deeplab with the BiFPN network to achieve the fusion of features at different levels, enhance the model's receptive field, and improve the recognition of image details. Deepen the Deeplab network structure, introduce the Atrous Spatial Pyramid Pooling (ASPP) structure after the encoder, and optimize the dilation rates of the atrous convolution in the ASPP module to improve the capture of multi-scale information and the model's ability to mine data features. Finally, fuse the Transformer module in the ASPP module, add a Transformer branch in the parallel branches of the ASPP, and fuse it with the convolutional features, as shown in the appendix Figure 4 shown. Perform frozen training by freezing the backbone network of the model to obtain an initial model. Subsequently, perform full training on the unfrozen backbone of the initial model, and train for multiple rounds until convergence to obtain the final segmentation model;
[0086] Step 4: Use a drone to collect drone farmland images of the farmland area to be measured, establish a test image set, and process the test image set using the multi-scale edge enhancement method in Step 2. The test images are embedded with aerial photography height information;
[0087] Step 5: Calculate the actual area corresponding to each pixel in each test image at different aerial photography heights by improving the drone area calculation method.
[0088] Step 6: Use the trained segmentation model to segment and extract the pixels of the farmland area in the test image set, and use the image binarization method to count the number of pixels segmented by the segmentation model. According to the number of pixels, combined with the actual area corresponding to each pixel at different aerial photography heights, the farmland area can be obtained.
[0089] Specifically, the calculation process of the improved Deeplab algorithm in Step 3 is as follows:
[0090] Step 1: Image input and preprocessing: Input a 3000×4000 pixel RGB farmland image, and scale it to 1536×2048 while optimizing the size and maintaining the aspect ratio.
[0091] Step 2: Swin Transformer backbone network:
[0092] 1. Patch Division: The image is cut into 4×4 windows, and the output tensor is [1×48×384×512].
[0093] 2. Hierarchical Feature Extraction: Stage 1, 4 Swin Transformer Blocks, Window Multi-Head Self-Attention (W-MSA): 7×7 local window, output resolution: 384×512 → 192×256. Stage 2: 6 Swin Transformer Blocks, the window is extended to 14×14, output resolution: 192×256 → 96×128. Stage 3: 18 Swin Transformer Blocks, Shifted Window Multi-Head Self-Attention (SW-MSA) is introduced, output resolution: 96×128 → 48×64. Stage 4: 6 Swin Transformer Blocks, global attention mechanism, output resolution remains 48×64.
[0094] 3. Feature Pyramid Construction: Output 4 levels of feature maps: C1: [1×96×384×512] (shallow edge features), C2: [1×192×256] (intermediate semantic features), C3: [1×384×96×128] (high-level semantic features), C4: [1×768×48×64] (global context features).
[0095] Step 3, BiFPN Feature Fusion: Bidirectional Cross-Scale Connections:
[0096] 1. Top-Down Path: Upsample C4 (48×64) by 2 times and fuse it with C3 (96×128), using depthwise separable convolution to reduce the computational cost.
[0097] 2. Bottom-Up Path: Downsample the fused features and fuse them with C2 (192×256) again, introducing learnable feature weights (implemented by 1×1 convolution).
[0098] 3. Multiple Rounds of Iteration: Repeat the bidirectional fusion 5 times (the depth design of BiFPN), and finally the output feature map resolution is 384×512 (the same size as the original image);
[0099] Step 4, Improved Atrous Spatial Pyramid Pooling (ASPP Module):
[0100] 1. Parallel multi-branch processing: Stage 1, Dilated Convolution Branch: Branch 1, 1×1 convolution (dilation rate 1) → Capture local details. Branch 2, 3×3 dilated convolution (dilation rate 6) → Medium receptive field. Branch 3, 3×3 dilated convolution (dilation rate 12) → Large-scale context. Branch 4, 3×3 dilated convolution (dilation rate 18) → Global correlation. Stage 2, Transformer Branch: Divide the feature map into 16×16 token sequences, and introduce 8-head self-attention calculation (each head dimension 96). Finally, the positional encoding adopts learnable relative positional encoding. Stage 3, Global Pooling Branch: Global average pooling → 1×1×768 → Bilinear interpolation to restore the size;
[0101] 2. Feature fusion: Use channel attention mechanism (SE Block) to dynamically allocate weights: CNN branch weight: 0.42 (dominated by local details), Transformer branch weight: 0.38 (supplemented by global relationships), pooling branch weight: 0.20 (assisted by background information). Finally, the output feature map size: [1×256×256×512];
[0102] Step 5: Decoder reconstruction:
[0103] 1. Feature upsampling: Use transposed convolution to gradually upsample (×2 → ×4 → ×8), and add element-wise to the shallow features output by BiFPN.
[0104] 2. Detail restoration: Introduce skip connection at the 256×512 resolution layer, and eliminate the artifacts brought by splicing through 3×3 convolution.
[0105] 3. Final prediction: 1×1 convolution compresses the number of channels to the number of classes, and bilinear interpolation upsamples to the image resolution of 1536×2048.
[0106] Step 6: Post-processing and output: Scale the image back to the original 3000×4000 pixel RGB farmland image, and generate a semantic segmentation mask map for visual output.
[0107] Specifically, the altitude information embedded in the UAV aerial image in Step 4 is obtained by on-board GPS.
[0108] Specifically, the improved UAV area calculation method in Step 5 is implemented by the following formula:
[0109] , , ,
[0110] As shown in the appendix Figure 2 as follows, where is the focal length, is the height of the camera from the ground, is half of the width of the camera film, is 1 / 2 of the width of the image plane and the object plane, is the camera half of the field of view angle in the direction. , are the number of pixels in the image , direction respectively, is the pixel corresponding size in the direction, is the pixel corresponding size in the direction, is the actual area under the corresponding height of each pixel.
[0111] The above formula is based on the idealized assumption that the pixels are in a standard square structure. However, according to the formula:
[0112]
[0113] the rectangular pixel characteristics commonly existing in the actual imaging system( ).
[0114] This will produce deformation errors and will also change dynamically according to the manufacturing errors of the sensor, where PAR is the pixel aspect ratio, is the length of the camera film, is the pixel physical size in the direction, is the pixel physical size in the direction. By analyzing the influence mechanism of the pixel aspect ratio on the deformation error, a dynamic adjustment model is established for the ideal pixels, so as to convert it into the dynamic adjustment and parameter compensation of , as shown in formula , where is the after dynamic adjustment, is the dynamic adjustment factor, and is set to 1 unit length in the model. Considering that different resolutions will affect the imaging, its aspect ratio ( ) is introduced to set the dynamic adjustment interval of k [1, . On this basis, through the endpoint average formula where represents the calculated value of the th endpoint average method, , to perform dynamic adjustment on to achieve parameter compensation for the deformation error of the ideal pixels. Finally is expressed as Therefore, the improved formula is as follows:
[0115] ,
[0116] where is the actual area corresponding to each dynamically adjusted pixel.
[0117] Finally, the area can be obtained by counting the pixels belonging to the measured area in the image. The area calculation formula is as follows: , where is the area, is the number of pixels in the measured area of the image.
[0118] Traditional UAV area calculation methods mainly rely on image processing technology and geometric formulas, and the extraction of the calculation area in the image mainly depends on manual extraction. Improvements to geometric methods, such as the introduction of UAV parameters and the application of other image software, can further improve the area calculation accuracy. At the same time, to overcome the limitations of UAV image technology in farmland area calculation, in recent years, scholars have proposed a method of integrating UAV images with deep learning to improve the extraction efficiency and segmentation accuracy of farmland plots. Currently, the mainstream semantic segmentation models include FCN, U-Net, PSPNet, and Deeplab, etc. Among them, Deeplab has a good effect in image segmentation tasks. Currently, researchers have also proposed a model of Deeplabv3+ based on Deeplab. Deeplabv3+ is a deep learning model for semantic segmentation proposed by Google AI, which mainly improves the performance of the Deeplab series. Its design goal is to more efficiently extract multi-scale features and accurately restore spatial information. Its core innovation lies in the combination of atrous convolution and the encoder-decoder framework. The encoder part includes a backbone network and an atrous spatial pyramid pooling (ASPP). The backbone network for feature extraction is mainly ResNet or Xception, etc. In the process of making the dataset of the semantic segmentation model, it is necessary to manually determine and label the feature area. For UAV aerial images of farmland, when the UAV takes pictures of farmland images, due to complex farmland environments, lighting changes, different shooting angles, etc., the farmland boundaries in the images are not obvious, making it difficult to achieve high-precision boundary determination. However, this problem and its solution have not been considered in existing research.
[0119] In this task, based on the camera imaging principle, an area calculation algorithm adaptable to height changes was established through a theoretical pixel dynamic adjustment model. In terms of farmland boundary determination, annotation, and dataset construction, the G-Lap algorithm was used to further enhance the farmland boundary, and the farmland boundary was determined according to the distribution of ak. On this basis, a connected domain was constructed, and the closed area was filled as the annotation area. For improving the Deeplab model and training method, the freezing training method was used to train the FCN, U-Net, PSPNet, and improved Deeplab models. For the improved Deeplab, the backbone network of the improved Deeplab was modified. The more powerful feature extraction network Swin Transformer was used to replace the original ResNet to enhance the model's representation ability. The FPN network of Deeplab was replaced with the BiFPN network to achieve the fusion of features at different levels, expand the model's receptive field, and improve the recognition of image details. The Deeplab network structure was deepened, and the atrous spatial pyramid pooling (ASPP) structure was introduced after the encoder. The dilation rates of the atrous convolution were optimized in the ASPP module to improve the capture of multi-scale information and the model's ability to mine data features. Finally, the Transformer module was fused in the ASPP module, and a Transformer branch was added to the parallel branches of the ASPP to fuse with the convolutional features. In our task, the improved Deeplab was used to extract the boundary area of the complete farmland.
[0120] The process of combining examples is as Figure 5 shown. The qualitative experiments of various semantic segmentation algorithms are as Figure 6 shown. The first column is the original image, the second column is the label image, the third column is the segmentation result of the method of the present invention, the fourth column is the segmentation result of PSP-Net, the fifth column is the segmentation result of U-Net, and the sixth column is the segmentation result of FCN; the quantitative index comparison of the above algorithms is as Figure 7 shown.
[0121] As Figure 8 shown is the comparison between the method of the present invention, the traditional image method, and the true area value. We took 23 farmland pictures with drone aerial photography height information by drone, and compared the areas calculated by the method of the present invention with the traditional image method and the true value.
[0122] In summary, in view of the fact that the existing area calculation method relies on a fixed flight altitude and adopts a regular square pixel geometric approximation model, the present invention proposes an improved area calculation method based on the camera imaging principle, and establishes an area calculation algorithm adaptable to altitude changes through a theoretical pixel dynamic adjustment model. To further improve the accuracy of farmland boundary determination, the G-Lap algorithm is used to further enhance the farmland boundary, and the farmland boundary is determined according to the distribution of ak. On this basis, a connected domain is constructed, and the closed area is filled as the annotation area. The improved Deeplab is used as the segmentation model, and the method of freezing training is adopted for training. After the improved Deeplab segmentation and extraction, the farmland pixels belonging to the area to be calculated in each UAV farmland image can be obtained. Finally, the number of each pixel is statistically obtained through binarization, and the area of the farmland in each test image is obtained. On the test set of this method, the improved Deeplab model reaches 95.7% mIoU, and the inference speed remains at 17.3 FPS (1080Ti), meeting the lightweight requirement under the premise of high accuracy, and having a good segmentation effect on different farmland scenes. The results obtained by this method are not inferior to the traditional calculation methods in terms of accuracy, and the calculation efficiency is improved, with higher versatility.
[0123] Embodiment 2
[0124] This embodiment also provides a farmland area measurement system based on the boundary determination and semantic segmentation of the G-Lap algorithm, including: a processor and a storage device; the processor loads and executes the instructions and data in the storage device to implement a farmland area measurement method based on the boundary determination and semantic segmentation of the G-Lap algorithm disclosed in Embodiment 1.
[0125] The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for measuring the area of farmland based on the G-Lap algorithm for boundary determination and semantic segmentation, characterized in that, It includes the following steps: Step 1: Collect images of agricultural areas in multiple scenarios through drones, and then construct a multi-scenario farmland image set; Step 2: Through the G-Lap algorithm, further enhance the farmland edges in the multi-scenario farmland image set in Step 1, and determine the farmland boundaries, so as to achieve high-precision boundary determination and annotation of the farmland, and further construct a multi-scenario farmland training image set; Step 3: Construct and train a semantic segmentation model. The semantic segmentation model is based on the Deeplab model as the basic network. In the backbone network of the Deeplab model, the feature extraction network Swin Transformer is used to replace the original ResNet, the FPN network of Deeplab is replaced by the BiFPN network, and the training image set is used to train the frozen backbone network to obtain an initial model. Subsequently, the initial model undergoes full training with the unfrozen backbone, and after multiple rounds of training convergence, the final segmentation model is obtained; Step 4: Use drones to collect drone farmland images of the farmland area to be measured, establish a test image set, and process the test image set in the multi-scale edge enhancement manner of Step 2. The test images are embedded with aerial photography height information; Step 5: Calculate the actual area corresponding to each pixel in each test image at different aerial photography heights by improving the drone area calculation method; Step 6: Use the trained segmentation model to segment and extract the farmland area pixels in the test image set, and use the image binarization method to count the number of pixels segmented by the segmentation model. According to the number of pixels, combined with the actual area corresponding to each pixel at different aerial photography heights, the farmland area is obtained.
2. The method according to claim 1, wherein In Step 1, the constructed multi-scenario farmland image set is for farmland under rice, wheat, fallow, and other farm crop scenarios, and the image has edge features wrapped by a complete farmland.
3. The method according to claim 1, characterized in that, In Step 1, the collected images are three-channel RGB images.
4. The method according to claim 1, characterized in that In Step 2, the method for boundary determination of the farmland edge through the G-Lap algorithm is as follows: Use the Laplacian operator to enhance the boundary of the complete farmland wrap in the multi-scenario farmland image set in Step 1. Subsequently, using the image enhanced by the Laplacian operator as a guiding image, adopt the guided filtering method to further enhance the boundary of the original farmland RGB image, and at the same time suppress crop texture noise. Based on the enhanced boundary, construct a connected domain and fill the closed area as the annotation area, so as to achieve high-precision boundary determination and annotation of the farmland.
5. The method according to claim 4, characterized in that In Step 2, the method for enhancing the boundary of the complete farmland wrap using the Laplacian operator is as follows: First, define the Laplacian operator of the two-dimensional image as a discrete convolution kernel: ; Among them, denotes the Laplacian operator of the function, and respectively denote the second-order partial derivatives of the function in the and directions; The original image is convolved using a 3×3 extended diagonal convolution kernel to obtain an edge response map ; Subsequently, the edge response map and the original image are superimposed proportionally to enhance the high-frequency edge information; Finally, perform normalized output. Since the convolution result may be negative, the output image needs to be normalized to a preset range.
6. The method according to claim 5, characterized in that The method of guided filtering is as follows: ; ; ; where \(i\) represents the pixel subscript, and are the linear coefficients within the window that guide the image to be the image enhanced by the Laplacian operator, is the mean of the guiding image within the window , is the variance of the guiding image within the window , is the mean of the input original image within the window , is the regularization parameter to prevent from being too large.
7. The method according to claim 6, wherein In Step 2, the method for boundary determination of the original farmland RGB image through guided filtering is as follows: In the boundary region of the guiding image, when , increases, the high-frequency details of the input image are retained. In the boundary region of the guiding image, when , decreases, , the noise of the input image is suppressed and smoothed; the farmland boundary is determined according to the distribution of , that is, , is the threshold, which is the high-value region of the farmland boundary, that is, the region of the farmland boundary, where the threshold is determined by the following formula: ; where μ edge and are respectively the distribution mean and standard deviation of 8. The method according to claim 7, characterized in that, In step 2, the boundaries of the farmland need to be marked: based on the high-value areas of the farmland boundaries, through a combination of morphological dilation and erosion operations, effective connection of the boundaries is achieved to form a closed area. By filling the closed area, the marked area can be obtained. The filled closed area is used as the marked area, thereby realizing the boundary marking of the farmland and further constructing a multi-scenario farmland training image set.
9. The method according to claim 1, characterized in that, In step 3, an atrous spatial pyramid pooling is introduced after the BiFPN network, and the dilation rate of the atrous convolution in the atrous spatial pyramid pooling is optimized.
10. The method according to claim 9, wherein The optimization method for the dilation rate of the atrous convolution is as follows: Dynamically design the dilation rate sequence according to the size of the feature map output by the atrous spatial pyramid pooling. The formula is: , For a 384×512 feature map, the dilation rate is adjusted to the combination of [1, 6, 12, 24].
11. The method according to claim 10, wherein, Two parallel network branches are introduced into the atrous spatial pyramid pooling. The two parallel network branches are the Transformer branch and the global pooling branch respectively.
12. The method according to claim 1, wherein In step 5, the method for calculating the area of the drone is as follows: ; ; Wherein, is the total area, is the actual area corresponding to each pixel, , are respectively the numbers of pixels in the , directions of the image, is the focal length, is the height of the camera from the ground, is half of the width of the camera film.
13. A farmland area measurement system for boundary determination and semantic segmentation based on the G-Lap algorithm, characterized in that, Including: A processor and a storage device; the processor loads and executes the instructions and data in the storage device to implement any one of the farmland area measurement methods for boundary determination and semantic segmentation based on the G-Lap algorithm described in claims 1 to 12.
Citation Information
Patent Citations
Measuring and calculating method, device and equipment for farmland area of aerial image, and medium
CN111080526A
Tunnel foreign matter detection method based on Retinex and YOLOv3 model
CN112115767A
Pavement crack pixel level detection method based on Seg-CapsNet algorithm
CN113643300A
Lightweight remote sensing image semantic segmentation method based on improved Deeplabv3 +
CN115984850A
Farmland boundary positioning information extraction method based on unmanned aerial vehicle low-altitude remote sensing image
CN117788822A
Cited By
Farmland intelligent identification method of interactive double-branch network based on deep learning
CN121353780A
Intelligent farmland recognition method based on deep learning interactive dual-branch network
CN121353780B
Flue-cured tobacco yield prediction method fusing airborne laser radar point cloud and orthoimage
CN121459225A