Cultivated land non-agrochemical pattern spot extraction method and device, equipment and storage medium
By combining hyperspectral image preprocessing and deep learning models with superpixel segmentation and spatiotemporal conditional random field models, the problems of low efficiency and limited accuracy in traditional monitoring of farmland non-agricultural use are solved. This enables accurate identification and attributed output of changes in non-agricultural use, and is suitable for dynamic monitoring of large areas.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-19
- Publication Date
- 2026-04-03
AI Technical Summary
Traditional monitoring of farmland conversion to non-agricultural uses relies on manual verification, which is inefficient, costly, and difficult to achieve dynamic supervision. Existing satellite remote sensing technology suffers from problems such as insufficient spectral-temporal separation processing, insufficient feature expression, difficulty in identification, and complex patch output, resulting in a high false alarm rate and limited identification accuracy in change detection.
By employing hyperspectral image preprocessing, four-dimensional data tensor construction, a hybrid model of temporal deep learning and Transformer, superpixel segmentation, and spatiotemporal conditional random field model, time-series spectral features and spatial correlation information are integrated to generate non-agricultural target patches with area, location, and time of change.
It improves the processing efficiency and identification accuracy of non-agricultural monitoring, is suitable for non-agricultural feature detection in complex environments over large areas, supports attributed patch output with area, location and change time, reduces manual interpretation costs, and improves the robustness and automation of monitoring.
Smart Images

Figure CN121789044A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of non-agricultural land parcel extraction technology, specifically relating to methods, apparatus, equipment, and storage media for extracting non-agricultural land parcels. Background Technology
[0002] The conversion of arable land to non-agricultural uses refers to the unauthorized use of arable land for non-agricultural construction, including industrial construction, commercial development, afforestation, and landscaping. This activity damages the topsoil structure, leading to a decline in land quality. Especially when arable land is converted to construction land, its productivity suffers irreversible damage, posing a potential threat to food security. Traditional monitoring of arable land conversion relies mainly on manual verification, which is inefficient, costly, and lacks dynamic monitoring capabilities. In recent years, the rapid development of satellite remote sensing technology has provided a new solution for monitoring arable land conversion. This technology offers advantages such as wide coverage, high monitoring frequency, and low cost, enabling proactive, grid-based, and precise monitoring management, significantly improving monitoring efficiency and accuracy. However, existing technologies still face some challenges: (1) Spectral-temporal separation processing: Traditional methods (such as SVM-RF) process single-phase data independently, ignoring temporal correlation, resulting in a high false alarm rate for change detection.
[0003] (2) Insufficient feature representation: Manually designed features (such as NDVI time-series curves) are difficult to capture nonlinear spectral variations, especially sensitive to shadows and mixed pixels.
[0004] (3) There are many types of non-agricultural uses of arable land, which are difficult to identify. They are mainly manifested in mining land, urban / residential land expansion, and aquaculture (such as fish ponds).
[0005] (4) The output process of the map features is complicated and it is impossible to extract map features with accurate attributes (area, land type, etc.). Summary of the Invention
[0006] In view of the shortcomings of the existing technology, the purpose of this invention is to provide a method, apparatus, equipment and storage medium for extracting non-agricultural patches of arable land.
[0007] According to one aspect of this application, a method for extracting non-agricultural land parcels is disclosed, the method comprising: After acquiring hyperspectral images of the target area at different time phases, the hyperspectral images of each time phase are preprocessed to obtain the target hyperspectral image of each time phase. Each target hyperspectral image is represented based on a three-dimensional data block. Multiple three-dimensional data blocks from multiple time phases are stacked along the time axis to obtain a four-dimensional data tensor of the target region. The four-dimensional data tensor is represented based on the four-dimensional data set. The four-dimensional data tensor of the target area is input into a pre-trained temporal deep learning model for classification and recognition, and the land cover classification result map of the target area in different time phases is obtained. The temporal deep learning model is a hybrid model of three-dimensional convolutional neural network and Transformer. Obtain the rules for converting land to non-agricultural use; Based on the land cover classification results map of different time phases and the non-agricultural land use conversion rules, determine the non-agricultural land use change areas in the target area that meet the non-agricultural land use conversion rules; Superpixel segmentation is performed on the non-agricultural change region to generate initial patches; The spatiotemporal conditional random field model is used to optimize the spatiotemporal consistency of the boundaries and categories of the initial patches to generate non-agricultural target patches, wherein each non-agricultural target patch includes area, location and change time.
[0008] In some embodiments, after acquiring hyperspectral images of the target area at different time phases, preprocessing the hyperspectral images of each time phase to obtain the target hyperspectral image for each time phase includes: After acquiring hyperspectral images of the target area at different time phases, the hyperspectral images of each time phase are sequentially processed by radiometric and atmospheric correction, geometric registration and fusion, and temporal consistency based on a pre-built coupling correction model to obtain the target hyperspectral image for each time phase.
[0009] In some embodiments, the hyperspectral image of each time phase is sequentially subjected to radiometric and atmospheric correction processing, geometric registration and fusion processing, and temporal consistency processing based on a pre-built coupling correction model to obtain the target hyperspectral image of each time phase, including: The hyperspectral images of each time phase are subjected to radiometric and atmospheric correction processing by combining a coupled correction model to convert the original brightness values of the hyperspectral images into the surface reflectance of the ground objects, thereby obtaining an initial optimized hyperspectral image. The coupled correction model is a hybrid model of the fast atmospheric correction QUAC model and the atmospheric radiative transfer 6S model. Determine the reference image; Based on the SIFT algorithm, matching point pairs between the initial optimized hyperspectral image and the reference image at each time phase are determined to eliminate the spatiotemporal misalignment between the initial optimized hyperspectral images at different time phases, thus obtaining spatiotemporally optimized hyperspectral images. Phenological alignment of vegetation; The cloud-covered image region on the spatiotemporally optimized hyperspectral image is determined, and the cloud-covered image region is repaired for cloud coverage defects based on a generative adversarial network to obtain the target hyperspectral image for each time phase.
[0010] In some embodiments, the step of inputting the four-dimensional data tensor of the target region into a pre-trained temporal deep learning model for classification and recognition, and obtaining land cover classification result maps of the target region at different time phases, includes: The four-dimensional data tensor of the target region is input into a pre-trained temporal deep learning model, so that the temporal deep learning model performs band filtering, temporal feature extraction, and spatial feature extraction on the four-dimensional data tensor. Band filtering of the four-dimensional data tensor includes: The difference in class probability distribution can be measured by the JM (Jeffries-Matusita) distance or importance can be assessed based on random forest, and redundant bands can be removed. Extracting temporal features from the four-dimensional data tensor includes: Extract the statistics of the spectral index time series of each pixel in the four-dimensional data tensor within a preset time window. The statistics include the mean, variance, trend slope, and coefficient of variation (CV). The vegetation index time series was fitted based on the Logistic growth function, and phenological parameters including the start time, end time, length and peak value of the growing season were extracted from the fitted curve. The spectral index or reflectance time series of each pixel is input into a Long Short-Term Memory (LSTM) network or a Temporal Convolutional Network (TCN), and its hidden layer state is extracted as a temporal feature. Spatial feature extraction of the four-dimensional data tensor includes: Construct multi-scale morphological profiles of different types of land cover; Calculate the mean and variance of each pixel in the Gabor response to generate a texture energy map; Superpixel objects are generated using the Simple Linear Iterative Clustering (SLIC) algorithm, and then spectral statistical features and geometric features are calculated within each superpixel object. The spatial features are determined based on the spectral statistical features and the geometric features.
[0011] In some embodiments, superpixel segmentation of the non-agricultural change region to generate initial patches includes: Based on the superpixel segmentation algorithm, the image corresponding to the non-agricultural change area is divided into a uniformly distributed initial cluster center grid. The distance from each pixel in the image to each initial cluster center is determined, and each pixel is assigned to the nearest cluster center. The distance metric combines the color similarity and spatial proximity of the pixels. Update the cluster centers based on the average position of all pixels assigned to each cluster; Repeat the allocation and update steps until the change in cluster centers is less than a preset threshold or the maximum number of iterations is reached, and generate superpixel segmentation results. The superpixel segmentation result is superimposed on the non-agricultural change region, and superpixels containing more than a preset threshold number or proportion of changed pixels are retained. The retained superpixels are used as the initial patch.
[0012] In some embodiments, the step of optimizing the spatiotemporal consistency of the boundaries and categories of the initial patch using a spatiotemporal conditional random field model to generate the non-agricultural target patch includes: The initial patch and multiple temporal image data are input into the spatiotemporal conditional random field model; The initial set of map features is optimized using the Spatiotemporal Conditional Random Field (SCRF) model to generate non-agricultural target map features. The energy function of the SCRF model consists of a univariate term, a spatial smoothing term, and a temporal consistency term. The univariate term is determined based on the class probability output by a deep learning classification model. The spatial smoothing term is used to constrain adjacent and spectrally similar map features to have consistent class labels. The temporal consistency term is used to constrain the class labels of map features at the same spatial location to conform to a preset land cover evolution law in different time phases.
[0013] According to another aspect of this application, a device for extracting non-agricultural land parcels is also disclosed, the device comprising: The hyperspectral image acquisition module is used to acquire hyperspectral images of the target area at different time phases, and then preprocess the hyperspectral images of each time phase to obtain the target hyperspectral image of each time phase. Each target hyperspectral image is represented based on a three-dimensional data block. The stacking module is used to stack multiple three-dimensional data blocks of multiple time phases along the time axis to obtain a four-dimensional data tensor of the target region. The four-dimensional data tensor is represented based on the four-dimensional data group. The input module is used to input the four-dimensional data tensor of the target area into a pre-trained temporal deep learning model for classification and recognition, so as to obtain the land cover classification result map of the target area in different time phases. The temporal deep learning model is a hybrid model of three-dimensional convolutional neural network and Transformer. The conversion rule acquisition module is used to acquire the conversion rules for non-agricultural land use. The module for determining non-agricultural land use change areas is used to determine non-agricultural land use change areas in the target area that satisfy the non-agricultural land use conversion rules based on land cover classification result maps of different time phases and the non-agricultural land use conversion rules. The initial patch generation module is used to perform superpixel segmentation on the non-agricultural change region to generate initial patches; The non-agricultural target patch generation module is used to optimize the spatiotemporal consistency of the boundary and category of the initial patch using a spatiotemporal conditional random field model to generate non-agricultural target patches, wherein each non-agricultural target patch includes at least area, location and change time.
[0014] According to another aspect of this application, an electronic device is also disclosed, the electronic device including a memory and at least one processor, the memory storing instructions; the at least one processor invokes the instructions in the memory to cause the electronic device to perform the various steps of the method for extracting non-agricultural land parcels as described in any of the preceding claims.
[0015] According to another aspect of this application, a computer-readable storage medium is also disclosed, on which instructions are stored, which, when executed by a processor, implement the various steps of the method for extracting non-agricultural patches of cultivated land as described in any of the preceding claims.
[0016] The present invention includes, but is not limited to, the following beneficial effects: (1) This scheme preprocesses hyperspectral images from multiple periods, uses a deep learning framework that integrates three-dimensional convolutional neural networks and Transformers to complete the automatic classification of land features, and uses superpixel segmentation technology and spatiotemporal topological relationships to optimize the final patch output. This technology improves the problems of low processing efficiency and limited recognition accuracy in the dynamic monitoring process of existing technologies by integrating time series spectral features and spatial correlation information. It shows obvious advantages in the fields of agricultural crop identification and urban development monitoring, and is suitable for handling non-agricultural land feature detection tasks in complex environments; (2) This scheme realizes the change from single-temporal to multi-temporal, uses four-dimensional tensors to construct the farmland evolution process, and realizes the capture of non-agricultural land use changes from non-agricultural identification; (3) This scheme supports attribute-based patch output based on area, location, and change time. Unlike traditional deep learning classification maps, the patch output of this scheme is more in line with the business needs of natural resource supervision and can meet the needs of law enforcement and supervision; (4) The rule reasoning mechanism of this scheme for non-agricultural use improves the accuracy of judgment. The introduction of non-agricultural land conversion rules explicitly encodes the behavior of land features changing from cultivated land to construction land, aquaculture, industrial and mining land, etc. It not only requires the category change to meet the logical consistency of the preceding and following phases and avoids interference from short-term outliers, but also eliminates pseudo-changes such as seasonal yellowing of cultivated land, short-term bare land, and cloud shadows, greatly improving the accuracy and interpretability of non-agricultural changes; (5) This invention integrates atmospheric correction, phenological alignment, temporal modeling, superpixel patch generation and CRF optimization into a unified process, which can maintain stable accuracy in large-area and multi-temporal monitoring. Especially in complex scenarios of rapid urban expansion and industrial construction, this invention shows higher robustness and automation, which can significantly reduce the cost of manual interpretation. The overall processing flow is more suitable for engineering applications of large-area and continuous monitoring. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.
[0018] Figure 1 This is a flowchart of the method for extracting non-agricultural land patches according to an embodiment of this application; Figure 2 This is a flowchart illustrating one method for generating initial patches according to an embodiment of this application. Figure 3 This is a flowchart illustrating how to obtain target hyperspectral images for each time phase according to an embodiment of this application; Figure 4 This is a flowchart illustrating one method for generating non-agricultural target patches according to an embodiment of this application; Figure 5 This is a schematic diagram illustrating the stacking of three-dimensional data blocks along the time axis in an embodiment of this application. Figure 6 This is a flowchart illustrating the target patch output process according to an embodiment of this application; Figure 7 This is a structural block diagram of the farmland non-agricultural patch extraction device according to an embodiment of this application; Figure 8 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application; Figure 9 This is a schematic diagram of a four-dimensional tensor in an embodiment of the application; Figure 10 This is a schematic diagram illustrating the repair principle of an embodiment of this application. Detailed Implementation
[0019] The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0020] For ease of understanding, the specific process of the embodiments of the present invention will be described below. Figure 1 This is a flowchart of a method for extracting non-agricultural land parcels according to an embodiment of this application. The method includes: S100. After acquiring hyperspectral images of the target area at different time phases, preprocess the hyperspectral images of each time phase to obtain the target hyperspectral image of each time phase.
[0021] Specifically, after acquiring hyperspectral images of the target area at different time phases, the hyperspectral images of each time phase can be processed sequentially using a pre-built coupling correction model, including radiometric and atmospheric correction, geometric registration and fusion, and temporal consistency processing, to obtain the target hyperspectral image for each time phase.
[0022] Each target hyperspectral image is represented based on a three-dimensional data block. This three-dimensional data block is a standardized data structure used to represent the target hyperspectral image, integrating the multidimensional information of the image into a unified data framework for easier subsequent storage, transmission, and analysis. Let the three-dimensional data block be D, then it can be represented as: Where H represents the image height (number of pixels), W represents the image width (number of pixels), and B represents the number of spectral bands. Preprocessing of hyperspectral images for each temporal phase can involve eliminating sensor noise, illumination differences, and atmospheric interference; commonly used algorithms include FLAASH and the 6S atmospheric radiative transfer model. This is used to achieve a unified radiometric benchmark for multi-temporal data, ensuring spectral comparability.
[0023] S102. Stack multiple three-dimensional data blocks of multiple time phases along the time axis to obtain a four-dimensional data tensor of the target region.
[0024] Among them, such as Figure 9 As shown, the four-dimensional data tensor is represented based on a four-dimensional data set (time, height, width, band). Specifically, before stacking, the images of each time phase undergo spatial alignment, size uniformity, and band alignment. Spatial alignment means that the pixels of all images correspond geographically, which can be achieved through geometric registration. Size uniformity means that all images have the same height (H) and width (W). Band alignment means that all images have the exact same number of bands (B) and band definitions. Furthermore, as... Figure 5 As shown, N three-dimensional data blocks of shape (H, W, B) are arranged along a time dimension to obtain a four-dimensional data tensor of shape (T, H, W, B). Here, T represents the number of time phases, i.e., the length of the time dimension. Obtaining such a tensor facilitates subsequent data retrieval. For example, to obtain the reflectance values of a pixel located at coordinates (x=100, y=200) in June 2023 across all spectral bands, simply index the four-dimensional tensor as: [time phase index corresponds to June 2023, 200, 100,:], which will yield a spectral band vector of length B.
[0025] S104. Input the four-dimensional data tensor of the target area into a pre-trained temporal deep learning model for classification and recognition, and obtain the land cover classification result map of the target area in different time phases.
[0026] Specifically, the temporal deep learning model is a hybrid model of 3D convolutional neural networks and Transformers. Specifically, when a 4D tensor is input into the model, it utilizes its complex hierarchical structure (such as 3D convolutional layers and attention mechanisms) to extract and combine features layer by layer. The model does not view a single pixel in isolation at a particular time phase, but simultaneously observes the spatial context, such as the texture and shape surrounding a pixel (e.g., a regular rectangle might indicate a building, while an irregular one might indicate woodland), spectral features (the spectral curve composed of hundreds of bands of the pixel acts like a "fingerprint" of the land, distinguishing between farmland, cement, and vegetation), and temporal evolution (the reflectance change pattern of the pixel over T time phases). For example, the spectrum of farmland exhibits regular greenness changes during the growing season, but once it becomes a construction site, its spectrum suddenly becomes similar to bare soil or building materials and remains stable over a long period. Furthermore, the model can learn the rationality and temporal logic of land feature changes. For example, if a location exhibits typical farmland characteristics for the first 90% of its time series, but suddenly transforms into a homogeneous area with high reflectivity in the last few phases, the model is more likely to classify it as newly constructed buildings rather than bare land. The model can distinguish between seasonal changes (such as winter fallowing of farmland) and permanent transformations (farmland hardening). Ultimately, the model makes a classification judgment for each spatial location (H, W) in the four-dimensional tensor at each time phase (T). This judgment is based on the probability derived from a comprehensive analysis of the complete spatiotemporal spectral characteristics of that point, and the model selects the category with the highest probability as the output label. Furthermore, the system ultimately generates classification maps corresponding to the number T of input time phases; each map is a snapshot of the land cover categories. For example, if a four-dimensional data tensor with shape (T, H, W, B) is input, the output will be T land cover classification result maps, each a two-dimensional matrix with shape (H, W). The value at each position (i,j) in the matrix is a category label (e.g., 1=farmland, 2=buildings, 3=woodland, 4=water bodies, etc.), indicating the type of land cover that pixel was identified as in the corresponding time phase. It can be understood that the conversion of farmland to non-agricultural use is a process-oriented change; that is, the conversion of farmland to non-agricultural use is itself a four-dimensional event in spatiotemporal spectrum. Any reduction in dimensionality will lose the key information necessary to understand this event. A four-dimensional tensor can naturally and completely encapsulate all this information and can be directly understood and processed by advanced deep learning architectures such as Transformer, enabling the representation of the joint change trajectory of the same pixel in time, spectrum, and space. Furthermore, combined with Transformer, it can simultaneously focus on key time phases before, during, and after the change, making it suitable for abrupt, non-gradual changes such as the conversion of farmland to construction land.
[0027] In the specific classification and recognition process, the temporal deep learning model performs band filtering, temporal feature extraction, and spatial feature extraction on the four-dimensional data tensor. Band filtering of the four-dimensional data tensor includes: The difference in class probability distribution can be measured by the JM (Jeffries-Matusita) distance or importance can be assessed based on random forest, and redundant bands can be removed.
[0028] For example, the difference in probability distribution between two categories can be measured based on the JM distance, with a value range of [0, √2]. A larger value indicates a stronger band discrimination ability. The overall separability of band k is obtained by averaging over all category pairs.
[0029] Band contributions can also be assessed using Gini importance or permutation importance from random forests. Specifically, importance scores can be determined based on normalization, where the normalization formula is: ; In the formula, k represents the wave, I(k) is the importance or arrangement importance of wave band k, max(I) is the maximum importance value, and Inorm(k) is the normalized importance score of wave band k.
[0030] Furthermore, temporal feature extraction from four-dimensional data tensors includes: Extract the statistics of the spectral index time series of each pixel in the four-dimensional data tensor within a preset time window. The statistics include mean, variance, trend slope and coefficient of variation (CV).
[0031] The vegetation index time series was fitted using the Logistic growth function, and phenological parameters, including the start time, end time, length, and peak value of the growing season, were extracted from the fitted curve.
[0032] The spectral index or reflectance time series of each pixel is input into a Long Short-Term Memory (LSTM) network or a Temporal Convolutional Network (TCN), and its hidden layer states are extracted as temporal features.
[0033] Spatial feature extraction from four-dimensional data tensors includes: Construct multi-scale morphological profiles of different types of land features.
[0034] Specifically, a structuring element template can be moved across the image to perform erosion followed by dilation, i.e., first shrinking bright areas and then expanding them, or first dilation followed by erosion. It can be understood that erosion followed by dilation can smooth contours and separate fine connections, while dilation followed by erosion can fill small holes and connect adjacent areas. The structuring element can refer to a small disk or square of a predetermined size. Furthermore, a series of structuring elements of varying radii, from small to large, can be used to continuously perform either erosion followed by dilation or dilation followed by erosion on the image. The result of each scale's operation is a new image; subtracting the original image from the result of each scale yields a sequence of difference images that constitute a morphological profile.
[0035] Calculate the mean and variance of each pixel in the Gabor response to generate a texture energy map.
[0036] Specifically, based on Gabor filters, the mean and variance of each pixel's response across all Gabor filters are determined. The mean reflects the dominant direction and roughness of the texture at that point, while the variance reflects the contrast or saliency of the texture. The statistical values of all filter responses are combined to form the Gabor texture feature vector for that pixel, which, when visualized, becomes the texture energy map.
[0037] Superpixel objects are generated using the Simple Linear Iterative Clustering (SLIC) algorithm, and then spectral statistical features and geometric features are calculated within each superpixel object.
[0038] Specifically, a simple linear iterative clustering algorithm is used to automatically segment the input image into multiple small regions, each of which is a superpixel object. For each generated superpixel object, spectral statistical features and geometric features are calculated. The spectral statistical features include the statistical values of all superpixel objects in each band, such as the mean and standard deviation. The geometric features are the shape attributes of the superpixel object itself, such as area, perimeter, and aspect ratio.
[0039] Spatial features are determined based on spectral statistical characteristics and geometric features.
[0040] Specifically, spectral statistical features and geometric shape features are integrated into a unified high-dimensional feature vector according to preset rules. These preset rules can refer to vector concatenation or weighted combination rules, etc. The fused feature vector is then defined as the spatial feature of the analysis unit, such as a superpixel object.
[0041] S106. Obtain the rules for converting land to non-agricultural use.
[0042] For example, rules for converting land to non-agricultural use can be determined based on industry standards and classification systems, or based on domain expert knowledge and historical cases.
[0043] S108. Based on the land cover classification results map of different time phases and the non-agricultural land use conversion rules, determine the non-agricultural land use change areas in the target area that meet the non-agricultural land use conversion rules.
[0044] Specifically, the land cover classification result map is a pixel-level land cover category map with time labels, indicating the specific land cover type (such as cultivated land, construction land, forest land, etc.) of each pixel in each time phase (e.g., 2020, 2021, 2022). When executing the land cover classification result maps based on different time phases and the non-agricultural land use conversion rules to determine the non-agricultural land use change areas in the target area that meet the rules, the system will sequentially compare the classification maps of each pair of consecutive time phases (e.g., comparing the 2021 map with the 2020 map, and then comparing the 2022 map with the 2021 map). For each spatial pixel location, the system can perform the following judgment: obtain the classification result of the pixel in the previous time phase (Ti) and the classification result of the pixel in the next time phase (Tj). The two land use codes, such as "cultivated land" in the previous phase and "construction land" in the subsequent phase, are substituted into the non-agricultural land conversion rules defined in step S106 for matching. If the land use type in the previous phase ∈ the cultivated land set and the land use type in the subsequent phase ∈ the non-agricultural land set, and the change satisfies the constraints such as persistence, then it is determined that the pixel has undergone a "non-agricultural change" during this period. The persistence constraint can refer to the target duration of the change, such as 3 months, 4 months, or half a year. This example achieves automatic screening of cultivated land non-agriculturalization issues in large areas, replacing the inefficient traditional manual visual interpretation method, greatly improving the coverage and efficiency of monitoring. Furthermore, through strict rule matching, it can effectively exclude non-target changes such as cultivated land → unused land (abandoned land) and cultivated land → other agricultural land, accurately focusing regulatory attention on the core issue of non-agriculturalization and improving the accuracy of law enforcement. Understandably, the rule-based reasoning mechanism for non-agricultural land use in this scheme improves the accuracy of the judgment. By introducing rules for the conversion of land use to non-agricultural land, the scheme explicitly encodes the conversion of land features from arable land to construction land, aquaculture, industrial and mining land, etc. It not only requires that the category change meet the logical consistency of the preceding and following time phases and avoids interference from short-term outliers, but also eliminates pseudo-changes such as seasonal yellowing of arable land, short-term bare land, and cloud shadows, which greatly improves the accuracy and interpretability of non-agricultural land use changes.
[0045] S110. Perform superpixel segmentation on the non-agricultural change area to generate initial patches.
[0046] In some embodiments, Figure 2 This is a flowchart illustrating the generation of an initial patch according to an embodiment of this application. The flowchart includes the following steps: S200: Based on the superpixel segmentation algorithm, the image corresponding to the non-agricultural change area is divided into a uniformly distributed initial cluster center grid.
[0047] S202. Determine the distance from each pixel in the image to each initial cluster center, and assign each pixel to the nearest cluster center. The distance measure is the color similarity and spatial proximity of the fused pixels.
[0048] S204. Update the cluster centers based on the average position of all pixels assigned to each cluster.
[0049] S206. Repeat the assignment and update steps until the change in cluster centers is less than the preset threshold or the maximum number of iterations is reached, and generate superpixel segmentation results.
[0050] S208. Overlay the superpixel segmentation result with the non-agricultural change region, retain the superpixels whose number or proportion of changed pixels exceeds a preset threshold, and use the retained superpixels as the initial patch.
[0051] Specifically, the analysis of steps S200 to S208 is as follows: The non-agricultural land conversion change area can be represented as a binary raster image. Pixels with a value of 1 represent areas determined to have changed from cultivated land to non-agricultural land, while pixels with a value of 0 represent unchanged areas. Further, the reference image, rather than the binary conversion image, is segmented, for example using the SLIC algorithm. After segmentation, K initial cluster centers are evenly distributed on the reference image. For each pixel, within a certain search range, its combined distance to neighboring cluster centers in color and coordinate space is calculated, and it is assigned to the nearest cluster center. This process is iterated until convergence. Finally, the image is segmented into K compact regions with boundaries closely matching the contours of ground features and relatively uniform internal color and texture, i.e., superpixels. The reference image refers to the base image used for superpixel segmentation, which has high spatial resolution and rich spectral texture information. Its role is to provide the necessary visual basis for the superpixel segmentation algorithm to generate high-quality, accurately defined patches. Further, the generated superpixel boundary layer is precisely overlaid with the binary raster image of the differential non-agricultural land conversion change area. The number of pixels with a value of 1 within a superpixel is counted, and the percentage of these changed pixels is calculated: Percentage of Change = Number of Changed Pixels within the Superpixel / Total Number of Pixels in the Superpixel. Furthermore, a percentage threshold is set, such as 0.4 or 0.5. If a certain percentage of a region within a superpixel, such as 40% or 50%, is considered a changed region, the entire superpixel is considered a potential non-agricultural patch. The boundaries of all preserved superpixels are converted from raster format to vector polygon format such as Shapefile; each generated vector polygon is an initial patch.
[0052] S112. The spatiotemporal conditional random field model is used to optimize the spatiotemporal consistency of the boundaries and categories of the initial patches to generate non-agricultural target patches.
[0053] Specifically, this step can be implemented in the following steps: Step 1: Construct the graph structure: (1) Define nodes: Treat each initial patch as a node in the graph.
[0054] (2) Define spatial edge: If two polygons are spatially adjacent, i.e. have a common boundary, then a spatial edge is established between them. A spatial edge means that the categories or boundaries of the two nodes will affect each other.
[0055] (3) Define time edges: For the same geographical location, link the states of its patches in different time phases, such as 2021, 2022, and 2023, to form time edges. A time edge means that the past state of a place will affect its present state.
[0056] (4) Determine the graph structure: a topological graph containing spatial and temporal connections, where each node has a category label and boundaries.
[0057] Step 2: Define the energy function Total energy E = Energy of data terms + Energy of spatial smoothing terms + Energy of time consistency terms Specifically, the data term ensures that the optimization does not deviate from the strong evidence identified by the AI model. The spatial smoothing term corrects isolated, potentially misclassified small patches, making them consistent with the surrounding large, homogeneous areas. It ensures that the final boundaries fall on places where the ground features in the image truly change abruptly, such as field ridges, roads, and water edges. The temporal consistency term identifies single-temporal classification errors caused by clouds, shadows, and phenology.
[0058] Step 3: Model Reasoning and Optimization Solution Image segmentation, belief propagation and other optimization algorithms are used to iteratively search for possible solutions in the space. In this iterative process, the model continuously tries to adjust the labels of each patch and evaluate the overall energy, eventually converging to a stable state that is globally or locally optimal.
[0059] Step 4: Generate optimized target polygons During the optimization process, the spatial smoothing constraint automatically aligns the patch boundaries with the true edges in the image. After optimization, each patch is assigned a final category label, such as non-agricultural construction land.
[0060] Step 5: Attribute Calculation and Output (1) Determine the area: Based on the optimized vector boundary, it is directly calculated by GIS software.
[0061] (2) Determine the location: Calculate the coordinates of the center point of the patch or record its boundary coordinate sequence.
[0062] (3) Determine the change time: Analyze the optimal category sequence obtained by the map patch under the constraint of the time consistency term, and then clearly determine the specific time phase of the "conversion of cultivated land to non-agricultural land". Combine the image date to generate the change time attribute.
[0063] The final output of the non-agricultural target patches is a structured geographic map with location, area, and time of change.
[0064] Furthermore, after determining the non-agricultural target patches, the method may also include, for example: Figure 6 The steps shown are as follows: S600: Using non-agricultural target patches as input data, automatically calculate the patch area and center point coordinates based on the vector boundary of the non-agricultural target patches.
[0065] Specifically, based on the geometric calculation function of the GIS software in the system, the actual surface area of each polygon in its projected coordinate system can be automatically calculated, as well as the geometric center of each polygon can be automatically calculated and its latitude and longitude coordinates can be output.
[0066] S602. Spatial association is performed between the non-agricultural target map patches and place name data or administrative division data to extract the location local place name or administrative division information corresponding to the map patches.
[0067] Specifically, the system overlays and analyzes the map patch layer with the administrative division layer and place name layer to automatically determine which village-level administrative division unit each map patch falls within and records its complete affiliation (e.g., "Zhejiang Province / Hangzhou City / Xiaoshan District / Ningwei Subdistrict / Xinhua Village"). Understandably, to provide more precise positioning, the system can find the key place name closest to the center of the map patch and, combined with its location and distance, generate descriptive text. For example, "Located in Group 3 of Xinhua Village, approximately 150 meters east of Ningshui Road." This results in two location attributes for each map patch: administrative division information and its local place name.
[0068] S604. Combining the land cover classification results and change determination results, assign attribute information such as the original land cover type, the transformed land cover type, and the change time to the map patches.
[0069] Specifically, this step extracts the complete category sequence of the map features from the input data based on their spatial location. This includes the original land use type, the transformed land use type, and the change time attribute. The original land use type is the stable land feature category before the change occurred, usually a subcategory of cultivated land, such as paddy fields. The transformed land use type is the final land feature category after the change stabilizes, such as industrial land or rural residential land. The change time is the period during which the system determines the change occurred, such as between June 2023 and September 2023.
[0070] S606. Generate a non-agricultural patch attribute table according to the preset field structure to realize the structured storage of patch information.
[0071] Specifically, a standardized attribute table field structure can be pre-defined, mapping all the information generated in steps S600-S604 to the respective fields. This results in a clearly structured non-agricultural patch attribute table. This table can be associated with the patch vector file or used independently as input for statistical analysis. It signifies the transformation of unstructured information into structured data. S608. Export the non-agricultural patch results and their attribute tables, and statistically summarize the patch areas to form results data that can be used for database construction and statistical analysis.
[0072] Specifically, the final non-agricultural land parcel vector files, such as Shapefiles, and attribute tables, such as Excel or CSV, are exported to a standard format to form a statistical table of non-agricultural land parcels.
[0073] For example, the statistical table of farmland converted to non-agricultural uses is shown in Table 1: Table 1. Statistics of Non-Agricultural Use of Cultivated Land
[0074] Understandably, this solution preprocesses multi-period hyperspectral images, employs a deep learning framework integrating 3D convolutional neural networks and Transformers to automatically classify land features, and utilizes superpixel segmentation technology and spatiotemporal topological relationships to optimize the final patch output. This technology, by integrating time-series spectral features and spatial correlation information, improves upon existing technologies in dynamic monitoring by addressing issues such as low processing efficiency and limited recognition accuracy. It demonstrates significant advantages in fields such as agricultural crop identification and urban development monitoring, and is suitable for handling non-agricultural feature detection tasks in complex environments. Furthermore, this solution achieves a shift from single-temporal to multi-temporal data, utilizing four-dimensional tensors to construct the farmland evolution process, enabling the capture of changes in farmland use from non-agriculturalization identification. Moreover, this solution supports attribute-based patch output based on area, location, and time of change. Unlike traditional deep learning classification maps, the patch output of this solution better aligns with the operational needs of natural resource supervision, satisfying law enforcement and regulatory requirements. Furthermore, simple segmentation is prone to isolated noise pixels, cloud shadows, bare land, and misjudgments of phenological changes. Spatiotemporal conditional random fields can achieve spatial smoothing terms, suppress isolated small patches, and maintain temporal consistency terms, eliminating single-phase anomalies. Then, by combining superpixel segmentation with spatiotemporal conditional random field optimization, the collaborative constraints effectively reduce jagged boundaries and edge drift phenomena.
[0075] In some embodiments, Figure 3 A flowchart for obtaining target hyperspectral images for each temporal phase, such as Figure 3 The steps shown are as follows: S300: Combined with the coupling correction model, the hyperspectral images of each time phase are subjected to radiometric and atmospheric correction processing to convert the original brightness values of the hyperspectral images into the surface reflectance of the ground objects, thus obtaining the initial optimized hyperspectral images.
[0076] The coupled correction model is a hybrid of the Quick Atmospheric Correction (QUAC) model and the Atmospheric Radiative Transfer (6S) model. When processing hyperspectral images, the coupled correction model uses the 6S model for baseline calibration. In key time phases or regions, when accurate atmospheric parameters are available, the 6S model is used for high-precision correction, generating a baseline reflectance image. Furthermore, for other time phases lacking atmospheric parameters, the QUAC model is used for rapid correction. Then, through statistical relationships between images (such as histogram matching and pseudo-invariant feature points), the QUAC correction results are normalized to the reflectance baseline scale established by the 6S model. This hybrid strategy significantly improves processing efficiency while ensuring overall correction accuracy and guarantees the physical consistency and comparability of reflectance values across multiple time phases.
[0077] S302. Determine the reference image.
[0078] For example, the highest quality image from all temporal phases can be selected as the unified spatial reference frame for subsequent geometric registration. The highest quality image can be defined as one free of clouds, snow, and fog, with clear imagery, high geolocation accuracy, and no significant distortion. Typically, the mid-growing season or the phase with the most prominent target land cover features is chosen to better reflect the typical land cover types and landscape characteristics of the study area. Understandably, all other temporal phases will be registered to this reference image, ensuring spatial alignment of the entire temporal data stack.
[0079] S304. Based on the SIFT algorithm, determine the matching point pairs between the initial optimized hyperspectral image and the reference image for each time phase, so as to eliminate the spatiotemporal misalignment between the initial optimized hyperspectral images of different time phases and obtain the spatiotemporally optimized hyperspectral image.
[0080] Specifically, the matching point pairs between the initial optimized hyperspectral image and the reference image for each temporal phase are determined based on the scale-invariant feature transform (SIFT) algorithm. This process can include the following steps: SIFT is applied to both the reference image and the image to be registered. SIFT extracts keypoints that remain stable despite rotation, scaling, and brightness changes, and generates feature descriptors with a target dimension of, for example, 128 dimensions. By calculating the Euclidean distance between descriptors, the most similar (closest) and second most similar keypoints are found for each keypoint in the reference image on the image to be registered. A ratio test is used to eliminate fuzzy matches, resulting in a set of matching point pairs. Using these matching point pairs, an affine or polynomial transformation is applied to determine the geometric transformation model. This model describes the translation, rotation, scaling, and other deformations that occur in the image to be registered relative to the reference image. Based on this transformation model, the image to be registered is resampled, and its pixel grid is reprojected onto the coordinate grid of the reference image to generate the final registered image, i.e., the spatiotemporally optimized hyperspectral image.
[0081] S306. Perform phenological alignment on vegetation.
[0082] Understandably, since vegetation growth varies seasonally across different phenological stages, directly comparing images from different years on the same date with different phenological stages—for example, one year showing early spring and another late spring—can lead to misjudgments. Therefore, phenological alignment is necessary. Specifically, images from all time periods can be uniformly adjusted to a standard phenological stage using interpolation methods. For example, this could be unified to the peak vegetation growth period, heading stage, or leaf fall stage. Key phenological dates for each year can be fitted using long-term vegetation index curves, such as NDVI curves. Simultaneously, a growth baseline curve is constructed, and then the vegetation index curves of each year's images are stretched or compressed on the time axis using algorithms such as dynamic time warping to align their key phenological feature points with the baseline curve.
[0083] Understandably, after phenological alignment, images from different years reflect the state of vegetation at the same growth stage. This gives time-series comparisons clear biological and agronomic significance, enabling more accurate detection of spectral differences caused by land cover changes, such as farmland converted to buildings, rather than normal interannual fluctuations caused by phenological differences.
[0084] It is understood that this invention integrates atmospheric correction, phenological alignment, temporal modeling, superpixel patch generation, and CRF optimization into a unified process, which can maintain stable accuracy in large-area, multi-temporal monitoring. Especially in complex scenarios such as urban expansion and rapid industrial construction, this invention exhibits higher robustness and automation, which can significantly reduce the cost of manual interpretation. The overall processing flow is more suitable for engineering applications of large-area, continuous monitoring.
[0085] S308. Determine the cloud-covered image region on the spatiotemporally optimized hyperspectral image, and repair the cloud-covered image region based on the generative adversarial network to obtain the target hyperspectral image for each time phase.
[0086] Specifically, such as Figure 10 As shown, a Generative Adversarial Network (GAN) consists of a generator and a discriminator. The generator is typically a U-Net-type encoder-decoder structure that captures contextual information and performs spatial reconstruction. Its input is an image covered by clouds or a cloud mask, and its goal is to generate a complete, sharp image. The discriminator's goal is to determine whether the image is real or generated; its input could be a real, sharp image or an image forged by the generator. The two components compete against each other and co-evolve during training. This process can be formalized as a minimax game. After training, the cloud-covered image and its corresponding cloud mask are input into the generator, which outputs a complete image. Finally, the cloud-covered areas repaired by the generator replace the original cloud areas, while the sharp areas are preserved, thus obtaining the target hyperspectral image.
[0087] In some embodiments, Figure 4 This is a flowchart illustrating the generation of non-agricultural target patches according to an embodiment of this application. (See attached document.) Figure 4 It includes the following steps: S400. Input the initial patch and multiple temporal image data into the spatiotemporal conditional random field model.
[0088] S402. Optimize the category labels and geometric boundaries of the initial patch set using a spatiotemporal conditional random field model to generate non-agricultural target patches.
[0089] The energy function of the spatiotemporal conditional random field model consists of a univariate term, a spatial smoothing term, and a temporal consistency term. The univariate term is determined based on the class probability output by a deep learning classification model. The spatial smoothing term is used to constrain adjacent and spectrally similar patches to have consistent class labels. The temporal consistency term is used to constrain the class labels of patches at the same spatial location in different time phases to conform to the preset land cover evolution law.
[0090] Examples are illustrated below with typical ground features: The main land feature identification targets are shown in Table 1, distinguishing between construction land, forest and grassland, abandoned land and water bodies; distinguishing between urban / residential construction, production construction and cemetery construction in construction land; distinguishing between forest land and grassland in forest and grassland; distinguishing between cultivated land and forest and grassland during phenological periods; and distinguishing between cultivated land and abandoned land, and forest and grassland and abandoned land during phenological periods.
[0091] (1) Perform spectral index normalization calculation: NDVI, NDWI, NDBI.
[0092] (2) Based on threshold segmentation and feature extraction, first extract water bodies (NDWI), then separate vegetation (NDVI), and finally distinguish between buildings and abandoned land (NDBI+SWIR).
[0093] Table 2. Land Feature Types
[0094] (3) Based on the characteristics of high reflectivity roof materials (such as NDBI>0.2), regular grid spatial distribution, and nighttime light data to assist in verification, distinguish between urban / residential land, production and construction land and cemetery.
[0095] (4) Differentiate between forest land and grassland by combining the vegetation red edge index (NDRE). Forest land has stronger absorption of the red edge band and lower NDRE value (0.3-0.6), while grassland has higher NDRE value (0.5-0.8).
[0096] ; (5) Use time series to distinguish between cultivated land and forest / grassland and abandoned land during phenological periods.
[0097] Farmland: Periodic NDVI spikes, dramatic seasonal fluctuations, and subject to human control.
[0098] Forest and grassland: NDVI is stable or varies naturally with the seasons, without periodic peaks.
[0099] Abandoned land: NDVI is rising year by year, and seasonal fluctuations are transitioning from "arable land type" to "natural vegetation type".
[0100] Table 3 Comparison of Three Types of Land Use Characteristic Parameters
[0101] After data is input into the model, it is classified temporally and constrained to maintain spatiotemporal consistency using Markov random fields (MRF). Change detection and updates are performed based on differential imagery (CVA, change vector analysis), while the model is incrementally learned and updated to adapt to the dynamic evolution of ground features.
[0102] Examples of image segmentation are as follows: For example, the ESP2 tool can be used to select the optimal segmentation scale, and superpixel segmentation (SLIC) can be used to generate homogeneous patches, with parameters optimized using ROC curves.
[0103] (1) Initialization: The image is divided into an initial grid of uniformly distributed cluster centers.
[0104] (2) Assignment: Each pixel is assigned to the nearest cluster center based on a distance metric that combines spatial and color information.
[0105] (3) Update: Recalculate the cluster center based on the average position of all pixels assigned to the cluster.
[0106] (4) Iteration: Repeat steps (2) and (3) until convergence, i.e., the cluster centers do not change significantly between iterations.
[0107] The process of generating homogeneous patches through multi-scale segmentation involves image segmentation, scale parameter optimization, and region merging. Its core is based on region growing algorithms and the criterion of minimizing local heterogeneity. The following is a step-by-step explanation and key formulas: 1) Define the heterogeneity criterion. The goal of segmentation is to minimize the heterogeneity within the patches. H This approach maximizes heterogeneity between patches. Heterogeneity is typically composed of spectrum (color) and shape (compactness / smoothness). ; in: Δ h color: spectral heterogeneity (e.g., weighted sum of standard deviations of bands).
[0108] Δ h shape: Shape heterogeneity (compactness Δ) h compactness or smoothness Δ h smooth).
[0109] w color, w shape is the weight ( w color+ w shape=1).
[0110] Spectral heterogeneity formula (taking multi-band images as an example): ; σ b Let w be the standard deviation of the b-th band. b For band weights.
[0111] Formula for shape heterogeneity: ; Compactness Δh compact =lA (l is the perimeter of the boundary, A is the area), smoothness Δh smooth Measure the deviation of the boundary from the ideal rectangle.
[0112] The edge detection (Canny operator) and region growing method are combined to improve boundary accuracy.
[0113] Morphological filtering (opening and closing operations) can be used to remove speckles (small noise points or isolated regions).
[0114] 1) Open operation removes speckles, removes bright speckles, and retains larger bright areas.
[0115] Erosion: Erosion of an image with a structuring element (such as a circle or square) will eliminate fragmentation (if the structuring element is larger than the fragmentation).
[0116] Dilation: Restores the original size of the object (but the fragments will not reappear because they have been eroded).
[0117] 2) Close the hole to remove dark spots and retain larger dark areas.
[0118] Dilation: Dilates the image using structuring elements, filling in dark fragments.
[0119] Erosion: Restores the original size of the object (but holes will not reappear since they have been filled). 3) Merge oversegmented regions based on semantic rules (area, shape index).
[0120] Further time series patch association was performed. It can construct spatiotemporal topological relationships of land features (adjacency, inclusion, intersection, etc. of land features in time and space), track the evolution trajectory of land features (the morphological and attribute changes of the same land feature at different points in time), and use sequence analysis (Markov chain) to predict future evolution trends.
[0121] (1) Use ArcGIS's geometric verification tool to fix problems such as self-intersection, gaps, and overlaps of polygons.
[0122] (2) Construct spatiotemporal topological relationships.
[0123] Constructing spatial topological relationships using ArcGIS's spatial connectivity and overlay analysis tools: Adjacency: Determines whether polygons share a boundary (e.g., share an edge or a vertex).
[0124] Inclusion: Whether the parent polygon completely contains the child polygon (such as an island in a lake).
[0125] Intersection: Whether the patches partially overlap (e.g., urban expansion covering farmland).
[0126] A temporal topology is constructed by associating data from different time phases using timestamps and unique IDs.
[0127] Continuity: The existence status of the same patch at adjacent time points (e.g., continuing, disappearing, or being added).
[0128] Splitting and Merging: The splitting (e.g., land subdivision) or merging (e.g., land consolidation) of plots from one time period to the next.
[0129] (3) Image patch transformation recognition The identification of cultivated land parcels transforming into non-agricultural parcels extends the update time for extracting non-agricultural parcels. Images from two separate classification periods are overlaid, and the classification results are compared to extract changed areas. Simultaneously, the latest image is overlaid with historical cultivated land parcels to check whether the original farmland boundaries have been encroached upon by newly constructed roads, houses, factories, etc.
[0130] (4) Patch matching and evolution tracking Attribute-based matching: If a plot has a unique identifier (such as a land parcel ID), it is directly associated by ID. If there is no unique ID, matching is performed using attribute similarity (such as land use type, area).
[0131] Matching based on spatial relationships: The method of maximizing overlapping area matches the current phase patch with the patch with the largest overlapping area in the next phase.
[0132] Trajectory construction: Linking matched patches in chronological order to form an evolutionary chain (e.g., ).
[0133] (4) Trajectory pattern mining was implemented using ArcGIS Pro's spatiotemporal pattern mining tool, Change Detection Wizard, and Python library programming to predict future map feature evolution trends.
[0134] The ownership of conflict areas was determined through evidence-based arbitration, comparing real-time remote sensing images with data from the Third National Land Survey.
[0135] Further classification and verification Stratified sampling generates a confusion matrix, and the overall accuracy (OA), Kappa coefficient, and F1 score are calculated.
[0136] The time-series results need to be verified to check the false negative / false positive rate of the changed pixels.
[0137] Further assessment of the map features is required. Compare manually interpreted or high-resolution imagery to calculate the positional error (RMSE): The coordinates of points with the same ground feature are extracted from the manual interpretation results and high-resolution imagery to ensure a one-to-one correspondence between points. The formula is: ; Where (x) i , y i (X) represents the coordinates of the i-th point as manually interpreted; i , Y i) represents the coordinates of the i-th point corresponding to the high-resolution image; n is the number of points.
[0138] The Dice coefficient is used to assess the degree of regional overlap between two binary images, converting manually interpreted and high-resolution images into binary images of the same resolution (foreground=1, background=0). The calculation formula is as follows: ; Where A is the manually interpreted binary image; B is the high-resolution binary image; |·| is the pixel count.
[0139] Further format conversion is required: Using Python and ArcGIS, we can perform format conversion (raster to vector), attribute calculation, extraction, and export of segmented polygons.
[0140] After identifying and verifying the area and land type attributes, a streamlined extraction process is performed: (1) After image segmentation, output raster data, use ArcGIS's built-in conversion tools and attribute table functions to calculate area attributes and repair geometrically processed self-intersecting or hollow patches.
[0141] (2) Extract land category attributes from existing fields or through spatial association.
[0142] Spatial Join: Overlays map features with land cover layers, matching intersecting or contained land cover attributes. Intersect: Preserves the intersection areas of map features and land cover layers, while inheriting attributes.
[0143] (3) Save as a standardized format (such as CSV table + Shapefile), including attributes such as ID (serial number), LandType (land type), Area (area).
[0144] The process of (data input - format conversion - attribute table export) is automated by using a for loop. Then, the conversion from text format to Excel spreadsheet is automated by calling a function.
[0145] According to another aspect of this application, a device for extracting non-agricultural land parcels is also disclosed, such as... Figure 7 As shown, the device includes: The hyperspectral image acquisition module is used to acquire hyperspectral images of the target area at different time phases, and then preprocess the hyperspectral images of each time phase to obtain the target hyperspectral image of each time phase. Each target hyperspectral image is represented based on a three-dimensional data block. The stacking module is used to stack multiple three-dimensional data blocks from multiple time phases along the time axis to obtain a four-dimensional data tensor of the target region. The four-dimensional data tensor is represented based on a four-dimensional data set (time, height, width, band). The input module is used to input the four-dimensional data tensor of the target area into a pre-trained temporal deep learning model for classification and recognition, and to obtain the land cover classification result map of the target area in different time phases. The temporal deep learning model is a hybrid model of three-dimensional convolutional neural network and Transformer. The conversion rule acquisition module is used to acquire the conversion rules for non-agricultural land use. The module for determining non-agricultural land use change areas is used to determine non-agricultural land use change areas in the target area that meet the non-agricultural land use conversion rules, based on land cover classification result maps of different time phases and non-agricultural land use conversion rules. The initial patch generation module is used to perform superpixel segmentation on non-agricultural change areas and generate initial patches; The non-agricultural target patch generation module is used to optimize the spatiotemporal consistency of the boundaries and categories of the initial patches using a spatiotemporal conditional random field model, and generate non-agricultural target patches, wherein each non-agricultural target patch includes at least area, location and change time.
[0146] The application of the relevant modules of the device in this example can be referred to the relevant introduction of the method principle above, and will not be repeated here.
[0147] This solution preprocesses multi-period hyperspectral imagery, employs a deep learning framework integrating 3D convolutional neural networks and Transformers to automatically classify ground features, and optimizes the final patch output using superpixel segmentation and spatiotemporal topological relationships. By integrating time-series spectral features and spatial correlation information, this technology improves upon existing techniques in dynamic monitoring, addressing issues such as low processing efficiency and limited recognition accuracy. It demonstrates significant advantages in fields such as agricultural crop identification and urban development monitoring, and is particularly suitable for handling non-agricultural ground feature detection tasks in complex environments.
[0148] above Figure 7 The device for extracting non-agricultural land parcels in this embodiment of the invention will be described in detail from the perspective of modular functional entities. The electronic equipment in this embodiment of the invention will be described in detail from the perspective of hardware processing.
[0149] Figure 8This is a schematic diagram of the structure of an electronic device 800 provided in an embodiment of the present invention. The electronic device 800 can vary significantly due to different configurations or performance characteristics. It may include one or more central processing units (CPUs) 810 (e.g., one or more processors) and a memory 820, and one or more storage media 830 (e.g., one or more mass storage devices) for storing application programs 833 or data 832. The memory 820 and storage media 830 can be temporary or persistent storage. The program stored in the storage media 830 may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the electronic device 800. Furthermore, the processor 810 may be configured to communicate with the storage media 830 and execute the series of instruction operations in the storage media 830 on the electronic device 800.
[0150] Electronic device 800 may also include one or more power supplies 840, one or more wired or wireless network interfaces 850, one or more input / output interfaces 860, and / or one or more operating systems 831, such as Windows Server, MacOSX, Unix, Linux, FreeBSD, etc. Those skilled in the art will understand that... Figure 8 The illustrated electronic device structure does not constitute a limitation on electronic devices and may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.
[0151] The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when executed on a computer, cause the computer to perform the steps of the method for extracting non-agricultural land parcels.
[0152] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system, device, or unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0153] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0154] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for extracting non-agricultural land parcels, characterized in that, The method includes: After acquiring hyperspectral images of the target area at different time phases, the hyperspectral images of each time phase are preprocessed to obtain the target hyperspectral image of each time phase. Each target hyperspectral image is represented based on a three-dimensional data block. Multiple three-dimensional data blocks from multiple time phases are stacked along the time axis to obtain a four-dimensional data tensor of the target region. The four-dimensional data tensor is represented based on the four-dimensional data set. The four-dimensional data tensor of the target area is input into a pre-trained temporal deep learning model for classification and recognition, and the land cover classification result map of the target area in different time phases is obtained. The temporal deep learning model is a hybrid model of three-dimensional convolutional neural network and Transformer. Obtain the rules for converting land to non-agricultural use; Based on the land cover classification results map of different time phases and the non-agricultural land use conversion rules, determine the non-agricultural land use change areas in the target area that meet the non-agricultural land use conversion rules; Superpixel segmentation is performed on the non-agricultural change region to generate initial patches; The spatiotemporal conditional random field model is used to optimize the spatiotemporal consistency of the boundaries and categories of the initial patches to generate non-agricultural target patches, wherein each non-agricultural target patch includes area, location and change time.
2. The method for extracting non-agricultural land parcels according to claim 1, characterized in that, After acquiring hyperspectral images of the target area at different time phases, preprocessing the hyperspectral images of each time phase to obtain the target hyperspectral image for each time phase includes: After acquiring hyperspectral images of the target area at different time phases, the hyperspectral images of each time phase are sequentially processed by radiometric and atmospheric correction, geometric registration and fusion, and temporal consistency based on a pre-built coupling correction model to obtain the target hyperspectral image for each time phase.
3. The method for extracting non-agricultural land parcels according to claim 2, characterized in that, The pre-built coupled correction model sequentially performs radiometric and atmospheric correction, geometric registration and fusion, and temporal consistency processing on the hyperspectral images of each time phase, resulting in the target hyperspectral image for each time phase, including: The hyperspectral images of each time phase are subjected to radiometric and atmospheric correction processing by combining a coupled correction model to convert the original brightness values of the hyperspectral images into the surface reflectance of the ground objects, thereby obtaining an initial optimized hyperspectral image. The coupled correction model is a hybrid model of the fast atmospheric correction QUAC model and the atmospheric radiative transfer 6S model. Determine the reference image; Based on the SIFT algorithm, matching point pairs between the initial optimized hyperspectral image and the reference image at each time phase are determined to eliminate the spatiotemporal misalignment between the initial optimized hyperspectral images at different time phases, thus obtaining spatiotemporally optimized hyperspectral images. Phenological alignment of vegetation; The cloud-covered image region on the spatiotemporally optimized hyperspectral image is determined, and the cloud-covered image region is repaired for cloud coverage defects based on a generative adversarial network to obtain the target hyperspectral image for each time phase.
4. The method for extracting non-agricultural land parcels according to claim 1, characterized in that, The step of inputting the four-dimensional data tensor of the target region into a pre-trained temporal deep learning model for classification and recognition includes: The four-dimensional data tensor of the target region is input into a pre-trained temporal deep learning model, so that the temporal deep learning model performs band filtering, temporal feature extraction, and spatial feature extraction on the four-dimensional data tensor. Band filtering of the four-dimensional data tensor includes: The difference in class probability distribution can be measured by the JM (Jeffries-Matusita) distance or importance can be assessed based on random forest, and redundant bands can be removed. Extracting temporal features from the four-dimensional data tensor includes: Extract the statistics of the spectral index time series of each pixel in the four-dimensional data tensor within a preset time window. The statistics include the mean, variance, trend slope, and coefficient of variation (CV). The vegetation index time series was fitted based on the Logistic growth function, and phenological parameters including the start time, end time, length and peak value of the growing season were extracted from the fitted curve. The spectral index or reflectance time series of each pixel is input into a Long Short-Term Memory (LSTM) network or a Temporal Convolutional Network (TCN), and its hidden layer state is extracted as a temporal feature. Spatial feature extraction of the four-dimensional data tensor includes: Construct multi-scale morphological profiles of different types of land cover; Calculate the mean and variance of each pixel in the Gabor response to generate a texture energy map; Superpixel objects are generated using the Simple Linear Iterative Clustering (SLIC) algorithm, and then spectral statistical features and geometric features are calculated within each superpixel object. The spatial features are determined based on the spectral statistical features and the geometric features.
5. The method for extracting non-agricultural land parcels according to claim 1, characterized in that... The non-agricultural change region is segmented into superpixels to generate initial patches, including: Based on the superpixel segmentation algorithm, the image corresponding to the non-agricultural change area is divided into a uniformly distributed initial cluster center grid. The distance from each pixel in the image to each initial cluster center is determined, and each pixel is assigned to the nearest cluster center. The distance metric combines the color similarity and spatial proximity of the pixels. Update the cluster centers based on the average position of all pixels assigned to each cluster; Repeat the allocation and update steps until the change in cluster centers is less than a preset threshold or the maximum number of iterations is reached, and generate superpixel segmentation results. The superpixel segmentation result is superimposed on the non-agricultural change region, and superpixels containing more than a preset threshold number or proportion of changed pixels are retained. The retained superpixels are used as the initial patch.
6. The method for extracting non-agricultural land parcels according to claim 5, characterized in that, The step of optimizing the spatiotemporal consistency of the boundaries and categories of the initial patch using a spatiotemporal conditional random field model to generate non-agricultural target patches includes: The initial patch and multiple temporal image data are input into the spatiotemporal conditional random field model; The initial set of map features is optimized using the Spatiotemporal Conditional Random Field (SCRF) model to generate non-agricultural target map features. The energy function of the SCRF model consists of a univariate term, a spatial smoothing term, and a temporal consistency term. The univariate term is determined based on the class probability output by a deep learning classification model. The spatial smoothing term is used to constrain adjacent and spectrally similar map features to have consistent class labels. The temporal consistency term is used to constrain the class labels of map features at the same spatial location to conform to a preset land cover evolution law in different time phases.
7. A device for extracting non-agricultural land parcels, characterized in that, The device includes: The hyperspectral image acquisition module is used to acquire hyperspectral images of the target area at different time phases, and then preprocess the hyperspectral images of each time phase to obtain the target hyperspectral image of each time phase. Each target hyperspectral image is represented based on a three-dimensional data block. The stacking module is used to stack multiple three-dimensional data blocks of multiple time phases along the time axis to obtain a four-dimensional data tensor of the target region. The four-dimensional data tensor is characterized based on a four-dimensional data set (time, height, width, band). The input module is used to input the four-dimensional data tensor of the target area into a pre-trained temporal deep learning model for classification and recognition, so as to obtain the land cover classification result map of the target area in different time phases. The temporal deep learning model is a hybrid model of three-dimensional convolutional neural network and Transformer. The conversion rule acquisition module is used to acquire the conversion rules for non-agricultural land use. The module for determining non-agricultural land use change areas is used to determine non-agricultural land use change areas in the target area that satisfy the non-agricultural land use conversion rules based on land cover classification result maps of different time phases and the non-agricultural land use conversion rules. The initial patch generation module is used to perform superpixel segmentation on the differentially varied regions to generate initial patches. The non-agricultural target patch generation module is used to optimize the spatiotemporal consistency of the boundary and category of the initial patch using a spatiotemporal conditional random field model to generate non-agricultural target patches, wherein each non-agricultural target patch includes at least area, location and change time.
8. An electronic device, characterized in that, The electronic device includes a memory and at least one processor, the memory storing instructions; the at least one processor invokes the instructions in the memory to cause the electronic device to perform the various steps of the method for extracting non-agricultural land parcels as described in any one of claims 1-6.
9. A computer-readable storage medium storing instructions thereon, characterized in that, When the instructions are executed by the processor, they implement each step of the method for extracting non-agricultural patches of cultivated land as described in any one of claims 1-6.