Artificial intelligence automatic extraction system and method for farmland planting plots from drone images

By introducing DCT transformation and multi-task learning frameworks into the farmland planting plot extraction system of drone images, the problems of blurred, adhesion and pseudo-boundary plot boundaries in drone images are solved, and a higher precision farmland planting plot extraction is achieved.

CN119360235BActive Publication Date: 2025-05-13BEIJING NORMAL UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411267914.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-11
Publication Date
2025-05-13
Estimated Expiration
2044-09-11

AI Technical Summary

Technical Problem

When the prior art uses ultra-high resolution drone images to extract farmland planting plots, problems such as blurred plot boundaries, plot adhesions, and pseudo-boundaries are prone to occur, resulting in inaccurate extraction results.

Method used

A drone image artificial intelligence automatic extraction system for farmland planting plots is proposed, including input module, feature extraction module, feature decomposition module and multi-task learning integrated decoding module. High-frequency and low-frequency components are extracted through DCT transformation and feature separation in the feature decomposition module, and feature decoding is performed through a multi-task learning framework to obtain more accurate plot extraction results.

Benefits of technology

By introducing the DCT transformation and multi-task learning framework, the recognition of the boundaries of farmland planting plots and the accuracy of the extraction results are significantly improved, and the problems of blurred, adhesion and pseudo-boundary plot boundaries are alleviated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119360235B_ABST
    Figure CN119360235B_ABST
Patent Text Reader

Abstract

The present application provides an artificial intelligence automatic extraction system and method for farmland planting plots from drone images, which belongs to the field of image data processing technology. It includes: an input module, a feature extraction module, a feature decomposition module, and a multi-task learning integrated decoding module. After the drone image data is feature extracted by the feature extraction module, the feature decomposition module performs DCT transformation and feature separation on the high-dimensional feature map in the frequency domain to form high-frequency components and low-frequency components, and then obtains a first regional feature map reflecting the plot area and a first boundary feature map reflecting the plot boundary through inverse transformation. Finally, the multi-task learning integrated decoding module is used to perform feature decoding on the first regional feature map and the first boundary feature map to obtain the results of farmland planting plot extraction. Due to the introduction of DCT transformation for frequency domain feature extraction and feature separation, the plot area and boundary feature representation capabilities are increased, and the extraction accuracy of farmland planting plots is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image data processing technology, and in particular to an artificial intelligence automatic extraction system and method for farmland planting plots from drone images. Background Art

[0002] Farmland plots refer to specific land areas used to grow crops in agricultural production. Using GIS, GPS and remote sensing technology to accurately divide and analyze farmland plots can better reflect the real situation of the land, help optimize the allocation of agricultural resources, support refined agricultural management, and enhance decision support.

[0003] With the development of technology, the plot extraction algorithm based on deep learning method has been applied to farmland plots. According to the characteristics of the model, deep learning methods can be divided into: single-task semantic segmentation and multi-task semantic segmentation methods. In the multi-task semantic segmentation method, some scholars have used multi-task learning models such as Psi-Net, BSINet, MLGNet, SEANet and GF-AFD to carry out research and practice on farmland plot extraction in meter-level resolution remote sensing images. However, when using ultra-high-resolution drone images to extract farmland planting plots, the extraction results of the above models are prone to problems such as blurred plot boundaries, plot adhesion, and pseudo-boundaries.

[0004] Therefore, it is necessary to provide an improved technical solution that can significantly improve the recognition of farmland planting plots. Summary of the invention

[0005] The purpose of this application is to provide an artificial intelligence automatic extraction system and method for farmland planting plots from drone images to solve or alleviate the problems existing in the above-mentioned prior art.

[0006] In order to achieve the above objectives, this application provides the following technical solutions:

[0007] This application provides an artificial intelligence automatic extraction system for farmland planting plots from drone images, including:

[0008] Input module, feature extraction module, feature decomposition module, multi-task learning integrated decoding module;

[0009] The input module is used to obtain drone image data of the farmland planting plots and input the image data into the feature extraction module;

[0010] The feature extraction module is used to receive the image data and perform feature extraction on the image data to obtain a high-dimensional feature map;

[0011] The feature decomposition module includes a first DCT transformation unit, a feature separation unit and an inverse transformation unit;

[0012] The first transformation unit is used to receive the high-dimensional feature map, and perform a first DCT transformation on the high-dimensional feature map to obtain a feature map after the first DCT transformation;

[0013] The feature separation unit is used to perform feature separation on the high-dimensional feature map to obtain high-frequency components and low-frequency components;

[0014] The inverse transformation unit is used to multiply the high-frequency component with the feature map after the first DCT transformation, and then perform an inverse transformation, and the result is used as the first boundary feature map of the farmland planting plot; multiply the low-frequency component with the feature map after the first DCT transformation, and then perform an inverse transformation, and the result is used as the first regional feature map of the farmland planting, and output the first boundary feature map and the first regional feature map to the multi-task learning integrated decoding module;

[0015] The multi-task learning integrated decoding module, based on the framework of multi-task learning, performs feature decoding on the first regional feature map and the first boundary feature map respectively to obtain the result of farmland planting plot extraction.

[0016] In a possible implementation, the feature extraction module includes a plurality of sequentially connected encoding layers;

[0017] The feature extraction module is used to receive the image data and input it into a plurality of sequentially connected coding layers for feature extraction, thereby obtaining a feature map of each coding layer;

[0018] The feature extraction module is further used to: use the feature map of the last coding layer as the high-dimensional feature map, and output it to the first DCT transformation unit.

[0019] In a possible implementation, the feature separation unit includes: a second DCT transformation subunit, a high-frequency information extraction subunit and a low-frequency information extraction subunit;

[0020] The second DCT transformation subunit uses a 1×1 convolution kernel to extract global spatial information from the high-dimensional feature map, and performs a second DCT transformation on the extracted spatial information to obtain a feature map after the second DCT transformation;

[0021] The low-frequency information extraction unit uses a Sigmoid function to perform a nonlinear transformation on the feature map after the second DCT transformation to obtain a nonlinear transformation result; determines whether the nonlinear transformation result is greater than or equal to 0.5, and if so, multiplies the nonlinear transformation result by a first constant matrix to obtain the low-frequency component;

[0022] The high-frequency information extraction subunit is used to obtain the difference between the first constant matrix and the low-frequency component to obtain the high-frequency component;

[0023] The first constant matrix is ​​a two-dimensional matrix having the same height and width as the high-dimensional feature map and a value of 1.

[0024] In a possible implementation, the multi-task learning integrated decoding module comprises a plurality of dual-branch channel attention DBAM blocks connected in sequence; the number of DBAM blocks is equal to the number of the encoding layers;

[0025] The DBAM block adopts a dual-branch structure, including: a regional branch and a boundary branch, wherein the regional branch and the boundary branch are arranged in parallel;

[0026] The region branch is used to sequentially perform 2-fold upsampling, 1×1 convolution, Sigmoid activation, and jump connection with the feature map of the corresponding encoding layer on the first region feature map to obtain a second region feature map;

[0027] The boundary branch is used to sequentially perform 2-fold upsampling, 1×1 convolution, Sigmoid activation, and jump connection with the feature map of the corresponding encoding layer on the first boundary feature map to obtain a second boundary feature map;

[0028] The second regional feature map and the second boundary feature map are provided to the next DBAM block, and the next DBAM block uses the regional branch and the boundary branch of the DBAM block to process the second regional feature map and the second boundary feature map, and then provides them to the next DBAM block again until all DBAM blocks are processed.

[0029] In a possible implementation, the system further includes: an information embedding block, wherein the information embedding block is used to embed the boundary information into the region information;

[0030] The information embedding block uses the ASPP method to perform feature enhancement on the second boundary feature map, and sums the enhanced second boundary feature map with the second region feature map to obtain a third region feature map;

[0031] The third region feature map replaces the second region feature map and is provided to the next DBAM block for processing.

[0032] In a possible implementation, the information embedding block adopts a dual-branch structure, including a channel attention branch and a spatial attention branch; the channel attention branch is arranged in parallel with the spatial attention branch;

[0033] The spatial attention branch is used to perform average pooling on the second boundary feature map and then input it into four parallel two-dimensional hole convolution kernels, each two-dimensional hole convolution applies a different expansion rate to obtain four two-dimensional feature maps of different scales, combine the four two-dimensional feature maps, perform a 1×1 convolution operation, and use a Sigmoid function for nonlinear transformation to obtain an enhanced boundary spatial attention map;

[0034] The channel attention branch is used to perform global pooling on the second boundary feature map to obtain a one-dimensional feature map, use four one-dimensional hole convolution kernels with different expansion rates to process the one-dimensional feature map to obtain four one-dimensional feature maps, combine the four one-dimensional feature maps, perform a 1×1 convolution operation, and perform nonlinear transformation with a Sigmoid function to obtain an enhanced boundary channel attention map;

[0035] The information embedding block is also used to sum the enhanced boundary space attention map, the enhanced boundary channel attention map and the second region feature map, and then perform a 3×3 convolution operation to obtain the third region feature map.

[0036] In a possible implementation, the multi-task learning integrated decoding module further includes: a plurality of integrated sub-decoding SDM (Ensemble Sub-Decoding Module) modules; except for the last DBAM block, each of the plurality of DBAM blocks corresponds to one SDM module;

[0037] An SDM module is used to obtain the third regional feature map and the second boundary feature map of the DBAM block corresponding thereto, and generate a regional depth supervision loss, a boundary depth supervision loss and a distance depth supervision loss of the DBAM block by using the third regional feature map and the second boundary feature map;

[0038] The multi-task learning integrated decoding module is also used to weight the regional depth supervision loss, the boundary depth supervision loss and the distance depth supervision loss of each DBAM block to obtain an overall loss function.

[0039] In a possible implementation manner, after the SDM module, further comprising: an uncertainty perception module;

[0040] The uncertainty perception module is used to process the third regional feature map to obtain an uncertainty map FU for measuring the foreground land parcel. M and uncertainty map BU of background plot M At the same time, the second boundary feature map is processed to obtain an uncertainty map FU for measuring the boundary of the foreground block Band background uncertainty map BU B ;

[0041] The uncertainty perception module is further used to enhance uncertainty perception according to the following rules:

[0042]

[0043] Where T is the preset uncertainty threshold and α is the scaling factor.

[0044] In a possible implementation, the multi-task learning integrated decoding module further includes: a spatial group enhancement SGE module, a second uncertainty perception module and a plot boundary connectivity module;

[0045] The SGE module is used to receive the feature map output by the last DBAM block, and obtain the plot area map, the plot boundary map and the distance map by using the multi-task output head;

[0046] The second uncertainty perception module is used to enhance the uncertainty of the land parcel area map output by the SGE module to obtain a final land parcel area extraction result;

[0047] The plot boundary connectivity module includes four strip convolution kernels in different directions, which are used to enhance the connectivity of the plot boundary map output by the SGE module to obtain the final plot boundary extraction result;

[0048] The final plot area extraction result, the plot boundary extraction result and the distance map are used as the result of farmland planting plot extraction.

[0049] The embodiment of the present application provides a method for automatic artificial intelligence extraction of farmland planting plots from drone images, the method being executed by any of the above-described systems for automatic artificial intelligence extraction of farmland planting plots from drone images, the method comprising:

[0050] Step S101: Acquire ultra-high resolution image data of farmland planting plots, and input the image data into the feature extraction module;

[0051] Step S102: extracting features from the image data to obtain a high-dimensional feature map;

[0052] Step S103: performing a first DCT transformation on the high-dimensional feature map to obtain a feature map after the first DCT transformation; performing feature separation on the high-dimensional feature map to obtain high-frequency components and low-frequency components;

[0053] Step S104: multiplying the high-frequency component with the feature map after the first DCT transformation, and then performing an inverse transformation, and the result obtained is used as the first boundary feature map of the farmland planting plot; multiplying the low-frequency component with the feature map after the first DCT transformation, and then performing an inverse transformation, and the result obtained is used as the first regional feature map of the farmland planting;

[0054] Step S105: Based on the framework of multi-task learning, feature decoding is performed on the first regional feature map and the first boundary feature map respectively to obtain the result of farmland planting plot extraction.

[0055] The technical solution of the embodiment of the present application has the following beneficial effects:

[0056] The artificial intelligence automatic extraction system of farmland planting plots from drone images provided in this embodiment includes: an input module, a feature extraction module, a feature decomposition module, and a multi-task learning integrated decoding module. After the drone image data of the farmland planting plot is subjected to feature extraction by the feature extraction module, the feature decomposition module performs DCT transformation and feature separation in the frequency domain on the high-dimensional feature map to form high-frequency components and low-frequency components, and then obtains the first regional feature map reflecting the plot area and the first boundary feature map reflecting the plot boundary through inverse transformation. Finally, based on the framework of multi-task learning, the first regional feature map and the first boundary feature map are feature decoded to obtain the results of farmland planting plot extraction. Due to the introduction of DCT transformation for frequency domain feature extraction and feature separation, the plot boundary feature representation capability is increased, which effectively alleviates the problems of blurred plot boundaries, plot adhesion, and pseudo-boundaries in the extraction of farmland planting plots from drone images, and improves the recognition of farmland planting plot boundaries and the accuracy of the final extraction results. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 This is a schematic diagram of the structure of an artificial intelligence automatic extraction system for farmland planting plots from drone images provided according to some embodiments of the present application.

[0058] Figure 2 A schematic diagram of the overall structure of a farmland planting plot extraction model provided according to some embodiments of the present application.

[0059] Figure 3 A schematic diagram of the structure of an information embedding block provided according to some embodiments of the present application.

[0060] Figure 4 A schematic diagram of analysis results of a transition region provided according to some embodiments of the present application.

[0061] Figure 5 A scenario simulation diagram of a transition area provided according to some embodiments of the present application.

[0062] Figure 6 A schematic diagram of the structure of an uncertainty perception module provided according to some embodiments of the present application.

[0063] Figure 7 A schematic diagram of the structure of a plot boundary connectivity module provided according to some embodiments of the present application.

[0064] Figure 8 A schematic flow chart of an artificial intelligence automatic extraction method of farmland planting plots from drone images provided according to some embodiments of the present application. DETAILED DESCRIPTION

[0065] The terms "first", "second", "third" and "fourth" etc. in the specification and claims of the present application and the drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally includes steps or units that are not listed, or optionally includes other steps or units inherent to these processes, methods, products or devices.

[0066] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0067] The embodiments of the present application are described below in conjunction with the accompanying drawings.

[0068] System Example

[0069] This embodiment provides an artificial intelligence automatic extraction system for farmland planting plots from drone images, such as Figure 1 As shown, the system includes: an input module, a feature extraction module, a feature decomposition module, and a multi-task learning integrated decoding module.

[0070] The input module is used to obtain drone image data of farmland planting plots and input the image data into the feature extraction module.

[0071] The feature extraction module is used to receive image data and perform feature extraction on the image data to obtain a high-dimensional feature map.

[0072] The feature decomposition module includes a first DCT transform unit, a feature separation unit and an inverse transform unit;

[0073] The first DCT transformation unit is used to receive the high-dimensional feature map and perform a first DCT transformation on the high-dimensional feature map to obtain a feature map after the first DCT transformation.

[0074] The feature separation unit is used to separate the features of the high-dimensional feature map to obtain high-frequency components and low-frequency components.

[0075] The inverse transformation unit is used to multiply the high-frequency component with the feature map after the first DCT transformation, and then perform an inverse transformation, and the result is used as the first boundary feature map of the farmland planting plot; multiply the low-frequency component with the feature map after the first DCT transformation, and then perform an inverse transformation, and the result is used as the first regional feature map of the farmland planting, and the first boundary feature map and the first regional feature map are output to the multi-task learning integrated decoding module.

[0076] The multi-task learning integrated decoding module, based on the multi-task learning framework, performs feature decoding on the first region feature map and the first boundary feature map respectively to obtain the results of farmland planting plot extraction.

[0077] Among them, drone image data refers to high-resolution images or data obtained by drone-mounted camera equipment (such as optical cameras, multispectral cameras, etc.). Since drones usually fly at lower altitudes, they can obtain ultra-high-resolution image data with a resolution of several centimeters, which is much higher than the meter-level resolution of traditional satellite images.

[0078] In this embodiment, the input module is used to obtain drone image data of farmland planting plots and input the image data into the feature extraction module.

[0079] In order to clarify the meaning of agricultural cultivation plots, the concepts of agricultural plots, agricultural cultivation plots (Agricultural Cultivation Field Parcels, CFP), and agricultural natural field plots (Agricultural Natural Field Parcels, NFP) are briefly explained below.

[0080] In this embodiment, the farmland plots include two plot subtypes: farmland planting plots and farmland natural plots.

[0081] Among them, farmland planting plots refer to the land within which only a single type of crop is planted in each agricultural production cycle. It is the smallest unit for farmers to carry out agricultural activities such as sowing, management, and harvesting. It is the basic unit for implementing precision agriculture and should theoretically be the smallest unit for crop area statistics and yield estimation. In drone images of farmland planting plots, the crop planting area is divided by the "conceptual" boundaries composed of spectral and texture differences between ridges, different crops, and the same crops due to planting time and management methods.

[0082] Corresponding to the cultivated plots of farmland is the natural plot of farmland, which refers to the crop planting area divided by "physical entity" boundaries of a certain width such as roads, ditches, ridges and non-agricultural land. A variety of crops can be planted within this area.

[0083] Farmland planting plots and farmland natural plots together constitute farmland plots. The core difference between the two lies in the different planting attributes and boundary characteristics of the two regions: there may be multiple crops planted inside natural plots of farmland, while the internal attributes of farmland planting plots are only one crop; the boundaries of natural plots of farmland are "physical entities" such as roads and ridges, which are wider and have clearer features, while the boundaries of farmland planting plots include not only physical entities such as ridges, but also "conceptual" boundaries composed of spectral and texture differences between different crops and the same crops due to planting time and management methods. The boundaries are more complex and diverse.

[0084] At present, domestic and foreign research mainly focuses on the extraction of natural plots of farmland, and there is little research on the extraction of planted plots of farmland. The research on natural plots of farmland is mainly based on deep learning methods, relying on time-series high-resolution satellite data to enhance the characteristics of planted plots. However, due to the limitations of spatial resolution and model feature expression capabilities, it is directly applied to ultra-high-resolution drone images to extract more detailed farmland planted plots, and the extraction effect needs to be improved. In other words, the current identification of farmland plots mainly uses meter-level high-resolution remote sensing data (such as Google remote sensing images, high-resolution remote sensing data, and Sentinel and Landsat data). The identification results are mostly natural plots of farmland, and the "conceptual" boundaries of farmland planted plots cannot be effectively identified. Higher-resolution image data is required for identification, such as ultra-high-resolution drone images. However, the increase in resolution brings great challenges to the accurate identification of farmland planted plots by traditional models and methods, which are specifically reflected in the following aspects:

[0085] (1) Blurred plot boundaries: The landscape shapes and contours of farmland plots are complex and diverse. The width of the plot boundaries is small and the features are fuzzy, making it difficult to extract complete plots, resulting in positional deviations in the boundaries and low connectivity.

[0086] (2) Plot adhesion: The planting plots are densely distributed and irregularly planted. Since the crop spectra of adjacent plots are similar, adhesion problems occur between adjacent plots.

[0087] (3) Pseudo-boundaries: The spectral and texture characteristics of the same crops within a plot are complex due to differences in growth, sowing and management. Internal heterogeneity creates pseudo-boundary problems, and tree shadows and weeds in complex scenes also bring additional challenges to plot identification.

[0088] In addition, since the "conceptual" boundaries of farmland planting plots have a higher density of redundant information in ultra-high-resolution drone images, it will inevitably lead to an increase in intra-class gaps and spatial complexity, and will also bring the influence of noise. The ultra-high spatial resolution also exacerbates the challenges of capturing long-range spatial information and extracting global semantic information in the local receptive field of the convolutional neural network, resulting in the model's insufficient ability to represent complex and slender objects (corresponding to plot boundaries).

[0089] In view of this, the artificial intelligence automatic extraction system of farmland planting plots from drone images provided in this application aims to extract more refined farmland planting plots from drone images. In this technical solution, a farmland planting plot extraction model (referred to as DCP-MLT model) is constructed based on the general segmentation network structure of encoder (Encoder)-decoder (Decoder). The model introduces a feature decomposition module after feature extraction, and uses DCT transform to convert spatial domain information to frequency domain, thereby mapping grayscale change information to the spectrum to enhance edge features and suppress noise, while alleviating the grayscale changes such as the heterogeneity of different types of plot boundaries, shadows and internal textures of plots under complex backgrounds in ultra-high-resolution drone images. It brings obstacles to spatial domain segmentation, and enhances its deep global context information and shallow spatial information features. Effective expression. It should be noted that in the introduction of the subsequent technical solution, farmland planting plots can be referred to as plots, planting plots, and farmland plots.

[0090] Specifically, the DCP-MLT model consists of three parts: an encoder, a decoder, and a feature decomposition module connected between the encoder and the decoder.

[0091] The encoder consists of multiple coding layers (also called multiple stages), which is used to receive image data and perform spatial feature extraction to obtain a high-dimensional feature map. Therefore, it is also called a feature extraction module.

[0092] In this embodiment, the feature extraction module of the DCP-MLT model can be constructed based on different existing network structures, for example, the VGG series, ResNet series or Transformer series network can be used as the basic structure. This embodiment does not limit the basic model of the feature extraction module.

[0093] For the convenience of description, the technical solution of this embodiment is described in detail below using the pre-trained ResNet-50 backbone network as the feature extraction module.

[0094] In this embodiment, the decoder (Decoder) is built based on the multi-task learning framework, and is therefore also referred to as a multi-task learning integrated decoding module. The multi-task learning integrated decoding module is used to convert the first region feature map and the first boundary feature map output by the feature decomposition module into specific outputs of multiple tasks. Among them, the multi-task learning integrated decoding module includes one core task and two subtasks, a total of three decoding tasks. Furthermore, the extraction (Region) of farmland planting plots (for the convenience of description, the internal area of ​​the farmland planting plots will be referred to as the plot area below) is set as the core task of the multi-task, and the plot boundary (Boundary) identification and distance (Distance) estimation are set as two auxiliary subtasks of the multi-task.

[0095] To complete the above three tasks, the multi-task learning integrated decoding module can combine deep supervision technology to construct multiple sub-decoding modules. The specific construction method will be described in detail in the subsequent embodiments. In addition, the multi-task learning integrated decoding module can also use the existing three parallel and structurally identical sub-decoding modules to complete different subtasks respectively. The specific construction method can refer to the prior art (for example, the decoder of the Psi-Net model or the decoder structure of the BsiNet model can be referred to). For the sake of simplicity, this embodiment will not be repeated.

[0096] After the feature extraction module and before the multi-task learning integrated decoding module, it also includes: feature decomposition module. This module realizes frequency domain conversion by introducing DCT transformation, so it is also called DCT module.

[0097] For example, Figure 2 It is a structural diagram of the DCP-MLT model. Figure 2 As shown in the figure, the DCP-MLT model adopts a Unet-like model structure as a whole, which consists of an encoder (feature extraction module) and a decoder (multi-task learning integrated decoding module), and a feature decomposition module is introduced between the encoder and the decoder.

[0098] Specifically, the feature decomposition module includes a first DCT transformation unit, a feature separation unit and an inverse transformation unit; the first DCT transformation unit is used to receive a high-dimensional feature map, and perform a first DCT transformation on the high-dimensional feature map to obtain a feature map after the first DCT transformation; the feature separation unit is used to perform feature separation on the high-dimensional feature map to obtain high-frequency components and low-frequency components accordingly; the inverse transformation unit is used to multiply the high-frequency component with the feature map after the first DCT transformation, and then perform an inverse transformation, and the result is used as the first boundary feature map of the farmland planting plot; the low-frequency component is multiplied with the feature map after the first DCT transformation, and then an inverse transformation is performed, and the result is used as the first regional feature map of the farmland planting, and the first boundary feature map and the first regional feature map are output to the multi-task learning integrated decoding module.

[0099] The two-dimensional DCT transform provided by the feature decomposition module can decouple the high-dimensional feature map obtained after feature extraction, and then use the global frequency domain information to extract the global high-frequency features (high-frequency components) and low-frequency features (low-frequency components). The high-frequency components are used to enhance the feature representation of the plot boundary, and the low-frequency components are used to enhance the high-level semantic features of the plot area, thereby improving the expression ability of the plot boundary and plot area information. In addition, by sequentially performing DCT transformation on the high-dimensional feature map, separating the features into high-frequency components and low-frequency components, and inverse transformation operations, it is possible to separate irrelevant information such as noise and pseudo-boundaries in the plot area, thereby improving the accuracy of farmland planting plot extraction.

[0100] In detail, the feature decomposition module regards the input high-dimensional feature map as a two-dimensional (or three-dimensional) discrete signal. These discrete signals have information of different frequencies. Through the two-dimensional DCT transform (also called DCT transform, DCT) in the first DCT transform unit, the image information can be feature decomposed, and the frequency domain transform method is used to extract and separate local low-frequency features, high-frequency features and global context information.

[0101] Assume that the high-dimensional feature map output by the feature extraction module is denoted as X5, and the first DCT transform unit performs DCT transform on X5, and the formula is as follows:

[0102]

[0103] Where, X dct It is the feature map after the first DCT transformation, DCT{.} represents the two-dimensional DCT transformation, R is a real number, and C5×H5×W5 represents the channels, height and width of X5.

[0104] In order to obtain high-frequency (plot boundary) and low-frequency (plot area) information in high-dimensional features, the feature separation unit specifically includes: a second DCT transformation subunit, a high-frequency information extraction subunit and a low-frequency information extraction subunit; the second DCT transformation subunit is used to use a 1×1 convolution kernel to extract global spatial information from the high-dimensional feature map before frequency domain conversion, and perform a second DCT transformation on the extracted spatial information to obtain a feature map after the second DCT transformation; the low-frequency information extraction unit uses a Sigmoid function to perform a nonlinear transformation on the feature map after the second DCT transformation to obtain a nonlinear transformation result; it is determined whether the nonlinear transformation result is greater than or equal to 0.5, and if so, the nonlinear transformation result is multiplied by a first constant matrix to obtain a low-frequency component; the high-frequency information extraction subunit is used to obtain the difference between the first constant matrix and the low-frequency component to obtain a high-frequency component; the first constant matrix is ​​a two-dimensional matrix with the same height and width as the high-dimensional feature map and a value of 1.

[0105] According to the above description, the expression corresponding to the feature separation unit is:

[0106]

[0107] X_hig=C-X_low (2)

[0108] In the formula, is the feature map after the second DCT transformation, X_low is the low-frequency component, X_hig is the high-frequency component, C is the first constant matrix, which is a two-dimensional matrix with a height and width of H5 and W5 respectively, and each element in the matrix is ​​a constant 1, sig{.} represents the Sigmoid function.

[0109] After the high-dimensional feature map is separated, the inverse transformation unit first performs filtering operations on the low-frequency component and the high-frequency component in the frequency domain, that is, element-wise multiplication, and then performs inverse transformation to obtain the first regional feature map and the first boundary feature map. The corresponding expressions are as follows:

[0110]

[0111] Where, X r , X b They are the first region feature map and the first boundary feature map, IDCT{.} represents the inverse transform, Represents element-wise multiplication.

[0112] In the feature decomposition module provided in this embodiment, the frequency domain method of DCT is used to perform frequency domain decomposition on the high-dimensional feature map, and the low-frequency component expressing the plot area and the high-frequency component expressing the plot boundary are extracted, and then the low-frequency component and the high-frequency component are respectively multiplied element by element with the feature map after the first DCT transformation to enhance the low-frequency and high-frequency information in the high-dimensional feature. After the inverse transformation, two more discriminative feature maps are obtained, namely the first region feature map and the first boundary feature map. The first boundary feature map focuses on expressing the plot boundary feature, and the first region feature map focuses on expressing the plot region feature. At the same time, the low-frequency component and the high-frequency component are respectively multiplied element by element with the feature map after the first DCT transformation, which can also effectively reduce the noise information in the high-dimensional feature map and enhance the feature expression capability.

[0113] In some optional embodiments, the feature extraction module includes multiple sequentially connected coding layers; the feature extraction module is used to receive image data and input it into multiple sequentially connected coding layers for feature extraction, thereby obtaining feature maps of each coding layer; the feature extraction module is also used to: use the feature map of the last coding layer as a high-dimensional feature map and output it to the first DCT transformation unit.

[0114] As an example, Figure 2 The encoder structure including 4 coding layers is shown. Each coding layer can be composed of convolution blocks and identity connection blocks, and each coding layer can have one or more convolution blocks and identity connections. The first coding layer receives the drone image data and performs feature extraction, and then inputs the result to the next coding layer, which performs further feature extraction, and so on, until the last coding layer, finally obtaining a high-dimensional feature map and outputting it to the first DCT transform unit.

[0115] It can be understood that the encoder structure including four encoding layers is only used as an exemplary description. In actual applications, the number of layers of the encoder may also be 5, 6, etc., which is not limited in this embodiment.

[0116] In some optional embodiments, the multi-task learning integrated decoding module is to reconstruct multiple sub-decoding modules in combination with deep supervision technology, including multiple dual-branch channel attention DBAM blocks (DBAM Block) connected in sequence; the number of DBAM blocks is equal to the number of coding layers; each DBAM block adopts a dual-branch structure, including: regional branches and boundary branches, and the regional branches and boundary branches are set in parallel.

[0117] The region branch is used to sequentially perform 2x upsampling, 1×1 convolution, Sigmoid activation, and jump connection with the feature map of the corresponding encoding layer on the first region feature map to obtain the second region feature map.

[0118] The boundary branch is used to sequentially perform 2x upsampling, 1×1 convolution, Sigmoid activation, and jump connection with the feature map of the corresponding encoding layer on the first boundary feature map to obtain the second boundary feature map.

[0119] The second regional feature map and the second boundary feature map are provided to the next DBAM block, which processes the second regional feature map and the second boundary feature map using the regional branch and the boundary branch of the DBAM block and then provides them to the next DBAM block again until all DBAM blocks are processed.

[0120] In the multi-task learning integrated decoding module provided in this embodiment, multiple sub-decoding modules are reconstructed in combination with deep supervision technology. The DBAM block is the core component of the module. The DBAM block adopts a dual-branch structure. The first boundary feature map (high-frequency data stream) and the first regional feature map (low-frequency data stream) are enhanced respectively through regional branches and boundary branches, so that different branches focus on their respective regions of interest, further improving the expression ability of features.

[0121] In this embodiment, the first half of the regional branch and the boundary branch use the same network structure, both of which include 2x upsampling, 1×1 convolution, and Sigmoid activation. Considering that the continuous use of spatial attention will over-suppress feature extraction, resulting in feature instability and thus hindering model training, this embodiment uses a residual method to optimize this attention mechanism. That is, after the Sigmoid activation of the regional branch and the boundary branch, it also includes: the step of jump-connecting the feature maps obtained at different stages of the encoder with the corresponding decoding module, so that the low-frequency data stream and the high-frequency data stream can be combined with the features of the corresponding coding layer respectively. The corresponding expressions are as follows:

[0122]

[0123] In the formula, S r Represents the regional feature map of Sigmoid activation, S b represents the boundary feature map activated by Sigmoid, Up(.) represents the upsampling operation, that is, bilinear interpolation, c is the number of channels of the feature map, Represents a 1×1 convolution operation, Conv2D t×t (.) indicates a two-dimensional convolution operation with a convolution kernel of 3, Concat(.) indicates a channel dimension concatenation operation, i is the serial number of the current DBAM block, and f i represents the feature map of the i-th coding layer. That is, the sequence number of the coding layer must be the same as the sequence number of the DBAM block, and there is a corresponding relationship between the coding layer and the DBAM block. is the second region feature map of the current DBAM block, is the second boundary feature map of the current DBAM block. It contains the spatial attention features of low-frequency data streams and high-frequency data streams respectively, so it is also called the spatial attention map corresponding to the low-frequency data stream and the high-frequency data stream.

[0124] By using residual connections combined with spatial attention mechanisms, the stability of feature propagation can be enhanced and the noise of shallow features can be effectively suppressed. In addition, by fusing the feature map of the coding layer with the feature map after Sigmoid activation (including regional feature maps and boundary feature maps), it is also possible to avoid the information loss that may be caused by DCT transformation in the frequency domain feature decomposition process, fully retain the original features, and further enhance the feature expression capability.

[0125] Considering that in the land segmentation task, the similarity between the land features and the background features (such as grass, ridges, and roads) in the drone image often leads to inaccurate segmentation results, and the blurred boundaries between the land plots lead to the land segmentation results not matching the actual land plot boundaries. Therefore, in some optional embodiments, the DBAM block also includes: an information embedding block (also called a DBAMmodule), which is used to embed boundary information into regional information; the information embedding block uses the ASPP (Atrous Spatial Pyramid Pooling) method to enhance the features of the second boundary feature map, and sums the enhanced second boundary feature map with the second regional feature map to obtain a third regional feature map; the third regional feature map replaces the second regional feature map and is provided to the next DBAM block for processing. By introducing the ASPP method, ASPP is used to increase the receptive field of each feature point and capture multi-scale features, and the enhanced boundary feature map is embedded into the regional feature map, further improving the accuracy of farmland planting plot segmentation.

[0126] Specifically, the information embedding block adopts a dual-branch structure, including a channel attention branch and a spatial attention branch; the channel attention branch is set in parallel with the spatial attention branch; the spatial attention branch is used to After average pooling, the image is input into four parallel two-dimensional dilated convolution kernels. Each two-dimensional dilated convolution applies a different expansion rate to obtain four two-dimensional feature maps of different scales. The four two-dimensional feature maps are combined and then subjected to a 1×1 convolution operation. The Sigmoid function is used for nonlinear transformation to obtain an enhanced boundary space attention map. The channel attention branch is used to focus on the second boundary feature map. Global pooling is performed to obtain a one-dimensional feature map, and the one-dimensional feature map is processed using four one-dimensional hole convolution kernels with different expansion rates to obtain four one-dimensional feature maps. The four one-dimensional feature maps are combined, and then a 1×1 convolution operation is performed. The Sigmoid function is used for nonlinear transformation to obtain an enhanced boundary channel attention map; the information embedding block is also used to sum the enhanced boundary space attention map, the enhanced boundary channel attention map and the second region feature map, and then a 3×3 convolution operation is performed to obtain the third region feature map.

[0127] Reference Figure 3 , the spatial attention branch first performs an average pooling (AP) operation, and then performs dilated convolution through four parallel convolution branches (four parallel two-dimensional hole convolution kernels) to capture spatial information of different scales. Each convolution branch uses a different dilation rate d i , expansion rate d iFor example, it can be set to 1, 3, 5, and 7. After that, the results of each dilated convolution are combined and the outputs of these branches are merged into a spatial attention map of the current DBAM block using 1×1 convolution and Sigmoid activation function, that is, the enhanced boundary spatial attention map, denoted as The corresponding expressions are as follows:

[0128]

[0129] In the formula, AP(.) represents average pooling, ASPP2D(.) represents two-dimensional dilated convolution, is the spatial attention result, with a shape height and width of 2H×2W, and i is the serial number of the current DBAM block.

[0130] The mechanism of the channel attention branch is similar to that of the spatial attention branch, with the second boundary feature map As input, the 1D feature result of global average pooling is performed, and the channel of the feature map is processed using a 1D convolution. Similarly, different expansion rates d i The convolution kernels of are used to capture multi-scale features, which are 1, 3, 5 and 7 respectively. After combining the results, the channel attention map is generated through 1x1 convolution and Sigmoid activation function, that is, the enhanced boundary channel attention map, denoted as Its expression is as follows:

[0131]

[0132] In the formula, ASPP1D(.) represents a one-dimensional dilated convolution, is the enhanced boundary channel attention map, and c represents the number of channels.

[0133] Finally, the information embedding block is also used to add the enhanced boundary space attention map and the enhanced boundary channel attention map, which is expressed as follows:

[0134]

[0135] In the formula, represents the third region feature map, and i is the sequence number of the current DBAM block.

[0136] The structure of the information embedding block is based on the second boundary feature map As input, average pooling (AP, Avg-Pool) and global average pooling (GAP) operations are first performed to extract the corresponding spatial features and global features, respectively, to achieve feature enhancement of global and local details, respectively.

[0137] In addition, the DBAM block is also used to: Perform a 3×3 convolution operation to obtain the third boundary feature map And use the third boundary feature map As a new second boundary feature map, it participates in the calculation of subsequent steps, as shown in the second formula of formula (8).

[0138] In this embodiment, by introducing the ASPP method in the information embedding block, context information of different scales is effectively combined, the receptive field is expanded, and it is conducive to the effective expression of the boundary features of ultra-high resolution drone images. In addition, the dual-branch structure is adopted to effectively enhance the boundary features through the spatial and channel attention mechanism, reduce feature redundancy and enhance valuable features.

[0139] In drone images with ultra-high spatial resolution, the land boundary is a physical entity with a certain width composed of multiple centimeter-level pixels. Compared with the linear characteristics of the land boundary in satellite images, the land boundary in drone images has a surface attribute of width, that is, the surface area is a transition area composed of multiple gradient pixels between the interior of the land and the boundary, which is called the transition area.

[0140] Figure 4 A schematic diagram of analysis results of a transition region provided according to some embodiments of the present application. Figure 5 A scenario simulation diagram of a transition area provided according to some embodiments of the present application. Figure 5 In the figure, light green is the plot, pink is the plot boundary, and the orange area is the transition area. The transition area is between the plot area and the plot boundary. The blue area represents the boundary line and its buffer zone, which is used to determine the plot boundary and transition area. Figure 4 , the left side is the plot in the image, the red line is the profile line of the plot, the pixel values ​​of the drone image are collected along the profile line, and the profile of the plot and boundary are drawn to obtain the curve graph of part I in the upper right part of the figure. As can be seen from the figure, the pixel values ​​of the plot-plot boundary-plot show a trend of rising-plateau-declining in the curve graph, where the rising and falling areas are defined as transition areas. In order to analyze the transition zone more intuitively, the relative change rate R(X) of each pixel passed by the profile line can be calculated based on the collected pixel values, and the curve graph of part II in the lower right part of the figure is obtained. As can be seen from the figure, the relative change rate R(X) in the transition zone is significantly greater than the relative change rate R(X) inside the plot area and outside the plot boundary. Pixel values ​​were collected for the transition zones of different geographical locations (such as ridges, drainage ditches, bare soil, hard soil roads, and planting areas) and different crops (such as rice planting plots and grasslands) in UAV images, and the above method was used to analyze the transition zones. It can be proved that transition zones are ubiquitous in ultra-high-resolution UAV images, that is, the transition zone theory is universal.

[0141] The calculation formula of the relative change rate R(X) is as follows:

[0142]

[0143] Where R(X) represents the absolute value of the relative rate of change at pixel X, represents the pixel value at F(X), F(X-1) represents the adjacent pixel value at X, and |·| represents the absolute value.

[0144] For the machine learning model, each pixel in the transition area may be classified as a plot area or a plot boundary. In other words, there is uncertainty as to which category (plot area or plot boundary) the pixels located in the transition area should be classified into.

[0145] For the transition region between the plot area and the plot boundary in the centimeter-level UAV image, some optional embodiments provide an uncertainty perception module for transition region (UPMTR), which is located after the SDM module and is used to process the feature map of the third region (for example, by comparing with a preset uncertainty threshold to determine whether each pixel in each feature map belongs to the foreground or the background), and obtain an uncertainty map FU for measuring the foreground plot. M and uncertainty map BU of background plot M At the same time, the second boundary feature map is processed to obtain the uncertainty map FU used to measure the boundary of the foreground block B and background uncertainty map BU B ;

[0146] The uncertainty perception module is also used to enhance uncertainty perception according to the following rules:

[0147]

[0148] In the formula, T is the preset uncertainty threshold, and α is the proportional factor, also called weight. The uncertainty result value of the transition area is weighted by formula (9). The stronger the uncertainty of each pixel in the feature map, the greater its weight (i.e., the value of α). By enhancing the uncertainty perception of the transition area, the model pays more attention to the uncertain area, and realizes the interaction and update of the information of the plot and the plot boundary. Then, the updated plot and plot boundary uncertainty area results are updated to the model results for prediction of the corresponding task.

[0149] Figure 6 An exemplary structure of the uncertainty perception module is shown, Figure 6As shown in FIG, the feature maps output by the SDM module include a regional feature map (the third regional feature map), a distance feature map, and a boundary feature map (the second boundary feature map), wherein the distance feature map does not participate in the various processing flows in the uncertainty perception module. The value of each pixel in the regional feature map represents the probability of it belonging to the foreground or background. Therefore, the uncertainty maps FU of the foreground (belonging to the plot area) and the background (not belonging to the plot area) can be extracted according to the value of each pixel. M , BU M Similarly, the value of each pixel in the boundary feature map also represents the probability of it belonging to the foreground or background. The uncertainty map FU of the foreground (belonging to the plot boundary) and the background (not belonging to the plot boundary) can be extracted according to the value of each pixel. B , BU B , and then uncertainty perception enhancement is performed according to the rules shown in formula (10).

[0150] It should also be noted that in the previous multi-task learning framework, the various output branches in the multi-task output head are set in parallel. In this embodiment, the uncertainty perception module and the SDM module are used to associate different task branches and output the supervision loss. The transition area is supervised by two tasks, the plot area and the plot boundary, and the dependency between the two tasks is adaptively learned, which can effectively improve the feature response of the transition area, reduce uncertainty, effectively analyze plots and plot boundaries, and jointly supervise the realization of information interaction and update of the two tasks, which can improve the accuracy of the two tasks at the same time.

[0151] In some optional embodiments, the multi-task learning integrated decoding module also includes: a spatial group enhancement SGE (Spatial Group-wise Enhance) module, a second uncertainty perception module and a field boundary connectivity module (Field Boundary Connectivity Module, FBCM); the SGE module is used to receive the feature map output by the last DBAM block, and use the multi-task output head to obtain the plot area map, plot boundary map and distance map; the second uncertainty perception module is used to enhance the uncertainty of the plot area map output by the SGE module to obtain the final plot area extraction result; the plot boundary connectivity module includes 4 strip convolution kernels in different directions, which are used to enhance the connectivity of the plot boundary map output by the SGE module to obtain the final plot boundary extraction result; the final plot area extraction result, plot boundary extraction result and distance map are used as the result of farmland planting plot extraction.

[0152] The SGE module is used to group and locally enhance the feature map output by the last DBAM block in the spatial dimension, and then obtain the plot region map (Region), plot boundary map (Boundary) and distance map (Distance) through the multi-task output head.

[0153] The second uncertainty perception module is used to enhance the uncertainty of the plot area map output by the SGE module to obtain the final plot area extraction result. It should be noted that the second uncertainty perception module has the same structure and processing flow as the uncertainty perception module described in the aforementioned embodiment. The only difference is the position in the DCP-MLT model. The second uncertainty perception module is located after the SGE module and is used to enhance the uncertainty of the plot area map output by the SGE module to obtain the final plot area extraction result. The specific uncertainty enhancement method can be executed with reference to the uncertainty perception module in the aforementioned embodiment, and this embodiment will not be repeated here one by one.

[0154] In farmland landscapes, the discontinuity of plot boundaries is caused not only by insufficient model representation, but also by scattered trees in the farmland, weeds growing randomly between plots (which have a high similarity to crops), and "continuous" planting of crops in the gaps between plots. This will lead to discontinuity of plot boundaries. Unlike the model representation, in this scenario, there is substantial occlusion or breakage at the plot boundaries, which interferes with the connectivity of the plot boundaries. The disconnection of plot boundaries caused by discontinuity is not obvious in the pixel-level results, but it will have a great impact on the connectivity of the plot boundary vectors. In order to alleviate this problem, in this embodiment, a plot boundary connectivity module is set before outputting the plot boundary map. In view of the special strip-like shape characteristics of the plot boundaries, strip convolution is used to model the strip-like plot boundaries. At the same time, considering the discontinuous characteristics of the plot boundaries, strip convolution with holes is used to explicitly model its special texture features, that is, a constant value, namely the hole, is used to simulate the discontinuous boundaries ( Figure 7 In the figure, light red, light green, light blue, and light purple are holes simulated by constant values).

[0155] The structure of the plot boundary connectivity module is as follows Figure 7 As shown, the implementation details are as follows:

[0156] Strip convolution is a strip convolution kernel in four different directions (horizontal, vertical, left diagonal and right diagonal). Vertical and horizontal strips can be directly generated by designing the shape of the convolution kernel. The diagonal strip needs to shift the input feature map one pixel to the left or right row by row to achieve diagonal sliding.

[0157] For example, suppose X∈R H×W×CRepresents the input feature map of the block boundary connectivity module, where H, W, and C represent the height, width, and number of channels, respectively. First, X is input to channels of different scales n times (for example, it can be 3 times by default) to achieve multi-channel replication of X so that it can be input to different processing paths; then the channel-replicated X is input to convolution blocks of different scales, each of which has a different convolution kernel size and / or a different direction (four different directions) of the void ratio, through the convolution kernel w at different scales and directions. r Convolution operations are performed to capture feature information of different scales and directions; then, the output feature maps of the hole strip convolutions in different directions (such as horizontal, vertical, diagonal, etc.) in each convolution block are connected through channels to obtain output feature maps of channels of different scales. Finally, the output feature maps of channels at different scales are connected to obtain the output of the plot boundary connectivity module. The corresponding expression is as follows:

[0158]

[0159]

[0160] In the formula, r represents the expansion rate, which is used to determine the position of the hole strip convolution kernel; k represents the size of the convolution kernel; w r represents the strip convolution kernel, which is a random matrix with holes, used to control the distribution position of local constants; i, j represent the position index of the input feature map X; D h , D v , D[h,v] represents the direction tensor, and h,v∈0,1; l represents w r The position index of ; G represents the function of the channel; is a random matrix; m is the channel index; n is the total number of channels.

[0161] Finally, the output F(X) of the plot boundary connectivity module is subjected to convolutional layers, batch normalization layers, and activation functions and then combined as a residual unit to obtain better robustness.

[0162] Refer again Figure 2 After 4 encoding layers, the resolution of the feature map is restored to half of the input image. The spatial details of the feature map are very rich, and the scale and geometric features of the plot boundaries and the types and shapes of boundary discontinuities are more complex. In the decoding stage, the strip convolution with holes generated by random sequences is used to adaptively learn more types of spatial detail information of discontinuous plot boundaries, so that the convolution kernel can match more boundary conditions like a template, better reduce discontinuities, improve the model's ability to capture boundary features of different scales, and output fitting results with a higher probability of plot boundary discontinuities.

[0163] In the training process, the back propagation algorithm is used to optimize the weight parameters of the DCP-MLT model. In addition, considering that in a deep and complex neural network structure, if the final output layer is supervised only by the loss function, it is easy to cause the gradient to disappear during back propagation, resulting in unstable optimization of the parameters of the middle layer of the network and affecting the network performance, in some optional embodiments, the DCP-MLT model also introduces deep supervision (Deep Supervision, DS) strategy, that is, in the multi-task learning integrated decoding module, it also includes: multiple integrated sub-decoding SDM modules; the number of SDM modules is one less than the DBAM blocks, that is, except for the last DBAM block, each DBAM block in the multiple DBAM blocks corresponds to an SDM module; the SDM module is used to obtain the third region feature map and the second boundary feature map of the DBAM block corresponding to it, and use the third region feature map and the second boundary feature map to generate the regional deep supervision loss, boundary deep supervision loss and distance deep supervision loss of the DBAM block (the regional, boundary and distance deep supervision losses of each DBAM block are collectively referred to as deep supervision losses, denoted as O1, O2, O3..., and the subscripts 1, 2, 3... correspond to the serial numbers of the DBAM blocks); the multi-task learning integrated decoding module is also used to weight the regional deep supervision loss, boundary deep supervision loss and distance deep supervision loss of each DBAM block to obtain the overall loss function.

[0164] Through the regional deep supervision loss, boundary deep supervision loss and distance deep supervision loss of each DBAM block, the intermediate layer no longer relies solely on the gradient gradually back-propagated by the final output layer, but is also supervised by the output graphs of different spatial resolutions of each intermediate DBAM to alleviate the gradient vanishing problem and improve the performance of the network. Since each DBAM block has a supervisory function in combination with its corresponding SDM module, the structure of DBAM block + SDM module is called a supervision branch.

[0165] Specifically, for the third region feature map, the second boundary feature map, and the distance feature map output by the DBAM block in each supervision branch, the image with the same spatial resolution is obtained by upsampling it and associating it with the corresponding true value label to calculate the loss, and the regional depth supervision loss, boundary depth supervision loss, and distance depth supervision loss of the DBAM block are obtained. By receiving direct feedback from the true value, the intermediate layer produces a better representation of the target area (especially the edge of the plot).

[0166] On this basis, the multi-task learning integrated decoding module is also used to optimize the DCP-MLT model, including: weighting the regional depth supervision loss, boundary depth supervision loss and distance depth supervision loss of each DBAM block to obtain the overall loss function.

[0167] The joint optimization strategy for multi-task deep supervised learning for optimizing the DCP-MLT model is described in detail below.

[0168] In this embodiment, the weighted binary cross entropy loss L wbce and dice loss L dice The combined loss function L jointly optimizes the model parameters of the plot area and the plot boundary to solve the performance problems caused by the high depth and complexity of the network model, while alleviating the problems of class imbalance and training instability. The corresponding expression is as follows:

[0169]

[0170] In the formula, represents the predicted value, and y represents the true value.

[0171] Weighted binary cross entropy loss L wbce It can effectively quantify the difference between the predicted result and the true value of each pixel, and the dice loss L dice The overlap between the predicted results and the true values ​​can be measured. The combination of the two can effectively alleviate the problems of category imbalance and training instability.

[0172] Furthermore, the weighted binary cross entropy loss L wbce The expression is as follows:

[0173]

[0174] In the formula, n represents the total number of pixels, i represents the sequence number of the pixel, ∈ represents the weight of the pixel, and y i , They represent the predicted value and true value of the i-th pixel respectively.

[0175] Considering that the number of pixels at the plot boundary is very small compared to the plot area, in order to balance the loss between positive and negative samples, ω is set to 7 for the plot boundary task and ω is set to 1 for the plot segmentation task.

[0176] Dice loss L dice The expression is as follows:

[0177]

[0178] In the formula, ∈ is a smoothing term with a value of 1. The purpose of setting this parameter is to prevent the error of dividing by 0. The other parameters have the same meaning as those in formula (12) and will not be described here.

[0179] Among the three tasks of the DCP-MLT model, the extraction of plot area and plot boundary belongs to the classification task, while the distance estimation belongs to the regression task. This embodiment uses the mean square error loss L mseTo quantify the error between the true distance and the predicted distance, the expression is as follows:

[0180]

[0181] Finally, the overall loss function is as follows:

[0182]

[0183] In the formula, ∈ N represents the loss weight, They represent the regional depth supervision loss, boundary depth supervision loss, and distance depth supervision loss of the DBAM block respectively; They represent the regional depth supervision loss, boundary depth supervision loss, and distance depth supervision loss output by the last DBAM block, respectively.

[0184] In the above formula, ω N It can be preset, for example, N Set to [0.3, 0.3, 0.5].

[0185] By jointly optimizing the model through the loss functions of the three tasks, the local and overall information of the plots in the network can be more comprehensively supervised.

[0186] In addition, in order to verify the artificial intelligence automatic extraction system of farmland planting plots from drone images, it also includes: building a dataset module and a comparative test analysis module.

[0187] There is currently no public data set for model training based on drone imagery to identify farmland plots. Therefore, this application provides a dataset building module to build the first set of drone farmland plot identification data sets with ultra-high spatial resolution, which has the characteristics of ultra-high spatial resolution, diverse plot types, and wide coverage. The detailed data set parameters are shown in Table 1. Table 1 is as follows:

[0188] Table 1. UAV farmland plot recognition dataset

[0189]

[0190]

[0191] The drone images and the corresponding labeled images were cropped into image blocks of 1024×1024 size, and the overlap rate was set to 25%. In order to improve the robustness of the model, randomly cropped images of 512×512 size were also included in the model training. Data enhancement methods included horizontal and vertical flipping, random rotation, and elastic transformation.

[0192] In the comparative test analysis module, several relevant state-of-the-art (SOTA) algorithms are selected for comparison. They are mainly divided into two types. The first type is single-task semantic segmentation algorithms, which include U-Net, PspNet, DeepLab V3+, HRNet, and SegFormer. The second type is multi-task segmentation algorithms, which have achieved SOTA performance in land extraction or building extraction tasks, including BsiNet, SEANet, CBR-Net, and HD-Net.

[0193] The experimental parameters and environment settings are as follows: PyTorch framework, computer: 4 NVIDIA RTX 2080Ti GPUs, using Adam optimizer, weight decay is 10 -5 The initial learning rate is set to 0.0001, the learning rate adjustment strategy is "CosineAnnealingWarmRestarts", and the parameters T0, T mult ,eta min The training was performed for 200 epochs until convergence. The best performing model weights on the validation set were saved for subsequent testing and evaluation.

[0194] The test results are shown in Table 2. Table 2 is as follows:

[0195]

[0196]

[0197] As shown in Table 2, the test results show that the accuracy evaluation indicators F1, IoU and mIoU of the DCP-MLT model in the plot area (Region) are 94.33, 92.88 and 82.72 respectively, which are 12.58%, 12.99% and 10.38% higher than the suboptimal results. At the same time, the accuracy evaluation indicators F1, IoU and mIoU of the DDCP-MLT model in the plot boundary (Boundary) are 70.57%, 60.94% and 78.15% respectively, which are 27.47%, 29.8% and 15.16% significantly improved compared with the suboptimal results. It can be seen that the DCP-MLT model provided in this embodiment uses DCT transform for feature separation, and then combines the segmentation of the plot area and the segmentation of the plot boundary as two different tasks to enhance each other, which can effectively improve the overall performance of the model.

[0198] Method Embodiment

[0199] This embodiment provides a method for automatically extracting farmland planting plots from drone images using artificial intelligence. The method is executed by using the system for automatically extracting farmland planting plots from drone images provided in any of the above embodiments. Figure 8 As shown, the method includes: Step S101 to Step S105. Specifically:

[0200] Step S101: Acquire ultra-high resolution image data of farmland planting plots, and input the image data into a feature extraction module;

[0201] Step S102: extracting features from the image data to obtain a high-dimensional feature map;

[0202] Step S103: performing a first DCT transformation on the high-dimensional feature map to obtain a feature map after the first DCT transformation; performing feature separation on the high-dimensional feature map to obtain high-frequency components and low-frequency components;

[0203] Step S104: multiplying the high-frequency component with the feature map after the first DCT transformation, and then performing an inverse transformation, and the result obtained is used as the first boundary feature map of the farmland planting plot; multiplying the low-frequency component with the feature map after the first DCT transformation, and then performing an inverse transformation, and the result obtained is used as the first regional feature map of the farmland planting;

[0204] Step S105: Based on the framework of multi-task learning, feature decoding is performed on the first region feature map and the first boundary feature map respectively to obtain the result of farmland planting plot extraction.

[0205] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. An artificial intelligence automatic extraction system for farmland planting plots from drone images, characterized in that: include: Input module, feature extraction module, feature decomposition module, multi-task learning integrated decoding module; The input module is used to obtain drone image data of the farmland planting plots and input the image data into the feature extraction module; The feature extraction module is used to receive the image data and perform feature extraction on the image data to obtain a high-dimensional feature map; The feature decomposition module includes a first DCT transformation unit, a feature separation unit and an inverse transformation unit; The first DCT transform unit is used to receive the high-dimensional feature map, and perform a first DCT transform on the high-dimensional feature map to obtain a feature map after the first DCT transform; The feature separation unit is used to perform feature separation on the high-dimensional feature map to obtain high-frequency components and low-frequency components; The inverse transformation unit is used to multiply the high-frequency component with the feature map after the first DCT transformation, and then perform an inverse transformation, and the result is used as the first boundary feature map of the farmland planting plot; multiply the low-frequency component with the feature map after the first DCT transformation, and then perform an inverse transformation, and the result is used as the first regional feature map of the farmland planting, and output the first boundary feature map and the first regional feature map to the multi-task learning integrated decoding module; The multi-task learning integrated decoding module, based on the framework of multi-task learning, performs feature decoding on the first regional feature map and the first boundary feature map respectively to obtain the result of farmland planting plot extraction.

2. The system according to claim 1, characterized in that The feature extraction module comprises a plurality of sequentially connected encoding layers; The feature extraction module is used to receive the image data and input it into a plurality of sequentially connected coding layers for feature extraction, thereby obtaining a feature map of each coding layer; The feature extraction module is further used to: use the feature map of the last coding layer as the high-dimensional feature map, and output it to the first DCT transformation unit.

3. The system according to claim 1, characterized in that The feature separation unit includes: a second DCT transformation subunit, a high-frequency information extraction subunit and a low-frequency information extraction subunit; The second DCT transformation subunit uses a 1×1 convolution kernel to extract global spatial information from the high-dimensional feature map, and performs a second DCT transformation on the extracted spatial information to obtain a feature map after the second DCT transformation; The low-frequency information extraction unit uses a Sigmoid function to perform a nonlinear transformation on the feature map after the second DCT transformation to obtain a nonlinear transformation result; determines whether the nonlinear transformation result is greater than or equal to 0.5, and if so, multiplies the nonlinear transformation result by a first constant matrix to obtain the low-frequency component; The high-frequency information extraction subunit is used to obtain the difference between the first constant matrix and the low-frequency component to obtain the high-frequency component; The first constant matrix is ​​a two-dimensional matrix having the same height and width as the high-dimensional feature map and a value of 1.

4. The system according to claim 2, characterized in that The multi-task learning integrated decoding module includes a plurality of dual-branch channel attention DBAM blocks connected in sequence; the number of DBAM blocks is equal to the number of encoding layers; The DBAM block adopts a dual-branch structure, including: a regional branch and a boundary branch, wherein the regional branch and the boundary branch are arranged in parallel; The region branch is used to sequentially perform 2-fold upsampling, 1×1 convolution, Sigmoid activation, and jump connection with the feature map of the corresponding encoding layer on the first region feature map to obtain a second region feature map; The boundary branch is used to sequentially perform 2-fold upsampling, 1×1 convolution, Sigmoid activation, and jump connection with the feature map of the corresponding encoding layer on the first boundary feature map to obtain a second boundary feature map; The second regional feature map and the second boundary feature map are provided to the next DBAM block, and the next DBAM block uses the regional branch and the boundary branch of the DBAM block to process the second regional feature map and the second boundary feature map, and then provides them to the next DBAM block again until all DBAM blocks are processed.

5. The system according to claim 4, characterized in that The DBAM block further includes: an information embedding block, the information embedding block being used to embed boundary information into the region information; The information embedding block uses the ASPP method to perform feature enhancement on the second boundary feature map, and sums the enhanced second boundary feature map with the second region feature map to obtain a third region feature map; The third region feature map replaces the second region feature map and is provided to the next DBAM block for processing.

6. The system according to claim 5, characterized in that The information embedding block adopts a dual-branch structure, including a channel attention branch and a spatial attention branch; The channel attention branch is arranged in parallel with the spatial attention branch; The spatial attention branch is used to perform average pooling on the second boundary feature map and then input it into four parallel two-dimensional hole convolution kernels, each two-dimensional hole convolution applies a different expansion rate to obtain four two-dimensional feature maps of different scales, combine the four two-dimensional feature maps, perform a 1×1 convolution operation, and use a Sigmoid function for nonlinear transformation to obtain an enhanced boundary spatial attention map; The channel attention branch is used to perform global pooling on the second boundary feature map to obtain a one-dimensional feature map, use four one-dimensional hole convolution kernels with different expansion rates to process the one-dimensional feature map to obtain four one-dimensional feature maps, combine the four one-dimensional feature maps, perform a 1×1 convolution operation, and perform nonlinear transformation with a Sigmoid function to obtain an enhanced boundary channel attention map; The information embedding block is also used to sum the enhanced boundary space attention map, the enhanced boundary channel attention map and the second region feature map, and then perform a 3×3 convolution operation to obtain the third region feature map.

7. The system according to claim 6, characterized in that The multi-task learning integrated decoding module further includes: a plurality of integrated sub-decoding SDM modules; except for the last DBAM block, each DBAM block in the plurality of DBAM blocks corresponds to one SDM module; An SDM module is used to obtain the third regional feature map and the second boundary feature map of the DBAM block corresponding thereto, and generate a regional depth supervision loss, a boundary depth supervision loss and a distance depth supervision loss of the DBAM block by using the third regional feature map and the second boundary feature map; The multi-task learning integrated decoding module is also used to weight the regional depth supervision loss, the boundary depth supervision loss and the distance depth supervision loss of each DBAM block to obtain an overall loss function.

8. The system according to claim 7, characterized in that After the SDM module, it also includes: an uncertainty perception module; The uncertainty perception module is used to process the third regional feature map to obtain an uncertainty map for measuring the foreground land parcel. Uncertainty map of the background plot At the same time, the second boundary feature map is processed to obtain an uncertainty map for measuring the boundary of the foreground block. and background uncertainty map ; The uncertainty perception module is further used to enhance uncertainty perception according to the following rules: , , In the formula, is the preset uncertainty threshold, is the scale factor.

9. The system according to claim 8, characterized in that The multi-task learning integrated decoding module also includes: a spatial group enhancement SGE module, a second uncertainty perception module and a plot boundary connectivity module; The SGE module is used to receive the feature map output by the last DBAM block, and obtain the plot area map, the plot boundary map and the distance map by using the multi-task output head; The second uncertainty perception module is used to enhance the uncertainty of the land parcel area map output by the SGE module to obtain a final land parcel area extraction result; The plot boundary connectivity module includes four strip convolution kernels in different directions, which are used to enhance the connectivity of the plot boundary map output by the SGE module to obtain the final plot boundary extraction result; The final plot area extraction result, the plot boundary extraction result and the distance map are used as the result of farmland planting plot extraction.

10. A method for automatically extracting farmland planting plots from drone images using artificial intelligence, the method being performed by the system for automatically extracting farmland planting plots from drone images according to any one of claims 1 to 9, the method comprising: Step S101: Acquire ultra-high resolution image data of farmland planting plots, and input the image data into the feature extraction module; Step S102: extracting features from the image data to obtain a high-dimensional feature map; Step S103: performing a first DCT transformation on the high-dimensional feature map to obtain a feature map after the first DCT transformation; Performing feature separation on the high-dimensional feature map to obtain high-frequency components and low-frequency components; Step S104: multiplying the high-frequency component with the feature map after the first DCT transformation, and then performing an inverse transformation, and the result obtained is used as the first boundary feature map of the farmland planting plot; multiplying the low-frequency component with the feature map after the first DCT transformation, and then performing an inverse transformation, and the result obtained is used as the first regional feature map of the farmland planting; Step S105: Based on the framework of multi-task learning, feature decoding is performed on the first regional feature map and the first boundary feature map respectively to obtain the result of farmland planting plot extraction.

Citation Information

Patent Citations

  • Crop disease identification method based on FCSA-OfficientNetV2

    CN114863278A

  • High-standard farmland plot vectorization extraction method based on multi-task learning

    CN115346137A