A deep learning enhanced cultivated land pattern history reconstruction system and method
The deep learning-enhanced historical reconstruction system for cultivated land patterns, combined with multi-source data and inverse simulation strategies, solves the problems of accuracy and continuity in the reconstruction of early cultivated land spatial patterns, achieving high-precision reconstruction of cultivated land patterns and providing a scientific basis for agricultural management and ecological protection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NORTHEAST INST OF GEOGRAPHY & AGRIECOLOGY C A S
- Filing Date
- 2026-01-28
- Publication Date
- 2026-05-19
AI Technical Summary
Existing technologies are insufficient in spatial resolution, temporal continuity, and spatial heterogeneity of farmland expansion when reconstructing early farmland spatial patterns, making it difficult to meet the needs of high-precision land use change research and regional agricultural management.
A deep learning-enhanced historical reconstruction system for cultivated land patterns acquires land use, natural environment, and socio-economic data through a multi-source data input module. It combines the inverse simulation strategies of the deep learning models U-net and PLUS to reconstruct cultivated land area and patterns year by year. By adjusting the operation mechanism of the CARS module using cultivated land existence probability maps and inverse simulation rules, it achieves high-precision reconstruction of cultivated land area.
It has achieved high-precision reconstruction of arable land patterns, providing scientific support for ecological protection and sustainable agricultural development in land areas, accurately tracing the historical evolution of the quantity and quality of land resources, and quantitatively assessing the contribution rate of arable land expansion to land degradation.
Smart Images

Figure CN121597964B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of historical farmland data reconstruction technology, specifically involving a deep learning-enhanced system and method for historical reconstruction of farmland patterns. Background Technology
[0002] Currently, most studies on land use change in the black soil region focus on periods monitored by modern remote sensing technology. Accurate reconstruction of the spatial pattern of cultivated land, especially in periods without remote sensing monitoring, remains significantly insufficient. Traditional methods for reconstructing historical cultivated land generally rely on statistical data interpolation and simple spatial allocation models. These methods have limitations in spatial resolution, temporal continuity, and characterization of the spatial heterogeneity of cultivated land expansion, easily leading to distorted reconstruction results and failing to meet the needs of high-precision land use change research and regional agricultural management.
[0003] Therefore, there is an urgent need to develop a method that can improve spatial accuracy while ensuring the continuity of time series, and combine regional geographic process simulation technology to deeply depict the evolution characteristics of cultivated land patterns in different periods. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention proposes a deep learning-enhanced system and method for historical reconstruction of arable land patterns.
[0005] Firstly, the system includes a multi-source data input module, a historical cultivated land area reconstruction module, and a deep learning-enhanced spatial simulation module;
[0006] Multi-data input module: acquires land use / cover remote sensing monitoring data, natural environment data, and socio-economic data of the area to be reconstructed; obtains remote sensing cultivated land map through land use / cover data, obtains natural environment driving factors through natural environment data, and obtains socio-economic driving factors through socio-economic data.
[0007] Historical cultivated land area reconstruction module: Construct a grain output-cultivated land area regression model for the study area, and reconstruct the cultivated land area of the area to be reconstructed based on the grain output-cultivated land area regression model. By calculating the difference in cultivated land area between adjacent years, the net increase or decrease in cultivated land pixels per year is obtained.
[0008] Deep learning-enhanced spatial simulation module: Through a reverse simulation strategy based on the PLUS model, the range of years to be reconstructed is identified, and the cultivated land area is reconstructed year by year.
[0009] Furthermore, natural environmental drivers include climate, soil, topography, and hydrology, while socioeconomic drivers include population, economy, and roads.
[0010] Furthermore, the formula for the grain yield-arable land area regression model is as follows:
[0011] ;
[0012] in, For arable land area, For grain production, For regression coefficients, For the intercept term, This is the error term.
[0013] Furthermore, in the deep learning-enhanced spatial simulation module, when reconstructing the cultivated land area year by year, from... Starting in 2010, it was gradually restructured year by year. The year, specifically:
[0014] S41. Generation of the probability map of cultivated land existence: Using the trained U-net model, input... The corresponding natural environment driving factors and socio-economic driving factors for each year were obtained. Probability of arable land existence corresponding to the year Probability map of arable land In China, use Represents the corresponding pixel The transition probability;
[0015] S42. Reverse Simulation Execution: The reverse simulation is performed using the CARS module of the PLUS model, starting from... Starting in [year], the operating mechanism of the CARS module was adjusted and [further adjustments were made]. The reconstruction of the cultivated land pattern in the region awaiting reconstruction;
[0016] S43. Refactoring by Year: Repeat steps S41-S42 until completion. The arable land pattern in the region awaiting reconstruction will be restructured.
[0017] Furthermore, the operating mechanism of the CARS module is adjusted as follows:
[0018] Macroeconomic aggregate reverse constraint: First, the historical arable land area reconstruction module calculates... Annual net increase / decrease in cultivated land pixels As a hard termination condition for the iteration at that time, deep learning was used in the evaluation formula of CARS. The transition probability replaces the traditional inference function, and the specific formula is as follows:
[0019] ;
[0020] in, It is arable land like a yuan Maintain the comprehensive conversion potential of arable land. Farmland pixels predicted based on the U-net model Transition probability, It is arable land like a yuan Neighborhood competition weight, It is a constraint of land use demand;
[0021] Reverse simulation rule process: Loading Land Use Map (Year) The conversion rule is defined as farmland removal, and the CARS module is based on all farmland pixels. Maintain the comprehensive conversion potential of arable land Generate all farmland pixels according to A table sorted in ascending order of values, then removing the top values from the sorted table one by one. One pixel, in The status was changed to non-arable land, and the update was obtained. Land Use Map (Year) .
[0022] Furthermore, the U-net model structure is as follows: it adopts an encoder-decoder symmetric structure and includes skip links; the encoder consists of four downsampling stages, each containing two... Convolutional layer, followed by batch normalization and LeakyReLU activation, with the end using... Max pooling; each stage of the decoder begins with The transposed convolution is upsampled, halving the feature channels, and then concatenated with the corresponding layer features from the encoder, followed by two more passes. Convolutional layers fuse features; the output layer uses... Convolution and Sigmoid activation output the farmland probability for each pixel.
[0023] Furthermore, the training of the U-net model is specifically as follows:
[0024] The input to the U-net model is organized into an image format, specifically the input feature X: integrating all spatialized driving factor layers of the baseline year, with the input being a single... tensor, and
[0025] These represent the number of rows and columns of the raster in the study area, respectively. The number of all spatialized driving factor layers, including two main parts: natural environmental factors and socio-economic factors;
[0026] The training label data consisted of a binary farmland distribution map from the baseline year, which was used as the training label. The farmland distribution map was then standardized for resolution to align with the driving factor data, resulting in a dimensional dataset. The binary mask.
[0027] Secondly, the specific method is as follows:
[0028] Acquire land use / cover remote sensing monitoring data, natural environment data, and socio-economic data of the area to be reconstructed for arable land pattern; obtain remote sensing arable land maps through land use / cover data, obtain natural environment driving factors through natural environment data, and obtain socio-economic driving factors through socio-economic data.
[0029] A regression model of grain output-arable land area was constructed for the study area. Based on the regression model of grain output-arable land area, the regional arable land area was reconstructed in the area to be reconstructed. By calculating the difference in arable land area between adjacent years, the net increase or decrease in the number of arable land pixels per year was obtained.
[0030] By using a reverse simulation strategy based on the PLUS model, the range of years to be reconstructed is identified, and the cultivated land area is reconstructed year by year.
[0031] The beneficial effects of the system described in this invention are as follows: It proposes a geospatial simulation framework enhanced by deep learning, which combines multi-source remote sensing images, historical archives and statistical data to achieve high-precision reconstruction of the historical pattern of cultivated land in land areas, providing scientific support for ecological protection, cultivated land resource management and sustainable agricultural development in land areas; it can accurately trace the historical evolution trajectory of the quantity and quality of land resources and quantitatively assess the contribution rate of cultivated land expansion to land degradation in different historical stages. Attached Figure Description
[0032] Figure 1 This is a flowchart illustrating the deep learning-enhanced historical reconstruction system for farmland patterns in this embodiment of the invention.
[0033] Figure 2 This is a flowchart of the deep learning-enhanced historical reconstruction system for farmland patterns after adding a system verification and uncertainty analysis module in this embodiment of the invention.
[0034] Figure 3 This is a schematic diagram of the distribution pattern and expansion pattern of cultivated land from 1949 to 1979 in an embodiment of the present invention. Detailed Implementation
[0035] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.
[0036] Example 1
[0037] This embodiment provides a deep learning-enhanced historical reconstruction system for cultivated land patterns, including a multi-source data input module, a historical cultivated land area reconstruction module, and a deep learning-enhanced spatial simulation module.
[0038] Multi-source data input module: acquires land use / cover remote sensing monitoring data, natural environment data, and socio-economic data of the area to be reconstructed; obtains remote sensing cultivated land maps through land use / cover data, obtains natural environment driving factors through natural environment data, and obtains socio-economic driving factors through socio-economic data.
[0039] Natural environmental drivers include climate, soil, topography, and hydrology, while socioeconomic drivers include population, economy, and roads.
[0040] Historical Cultivated Land Area Reconstruction Module: Construct a regression model of grain output-cultivated land area in the study area. Based on the grain output-cultivated land area regression model, reconstruct the cultivated land area in the area to be reconstructed. By calculating the difference in cultivated land area between adjacent years, obtain the annual net increase or decrease in cultivated land pixels.
[0041] The formula for the grain yield-arable land area regression model is:
[0042] ;
[0043] in, For arable land area, For grain production, For regression coefficients, For the intercept term, This is the error term.
[0044] Deep learning-enhanced spatial simulation module: Using a reverse simulation strategy based on the PLUS model, the cultivated land area is reconstructed year by year, starting from the last year of the reconstruction.
[0045] In the deep learning-enhanced spatial simulation module, the reconstruction of cultivated land area year by year is specifically performed as follows:
[0046] When reconstructing the cultivated land area year by year, from Starting in 2010, it was gradually restructured year by year. The year, specifically:
[0047] S41. Generation of the probability map of cultivated land existence: Using the trained U-net model, input... The corresponding natural environment driving factors and socio-economic driving factors for each year were obtained. Probability of arable land existence corresponding to the year Probability map of arable land In China, use Represents the corresponding pixel The transition probability;
[0048] S42. Reverse Simulation Execution: The reverse simulation is performed using the CARS module of the PLUS model, starting from... Starting in [year], the operating mechanism of the CARS module was adjusted and [further adjustments were made]. The reconstruction of the cultivated land pattern in the region awaiting reconstruction;
[0049] S43. Refactoring by Year: Repeat steps S41-S42 until completion. The arable land pattern in the region awaiting reconstruction will be restructured.
[0050] The specific adjustments to the CARS module's operating mechanism are as follows:
[0051] Macroeconomic aggregate reverse constraint: First, the historical arable land area reconstruction module calculates... Annual net increase / decrease in cultivated land pixels As a hard termination condition for the iteration at that time, deep learning was used in the evaluation formula of CARS. The transition probability replaces the traditional inference function, and the specific formula is as follows:
[0052] ;
[0053] in, It is arable land like a yuan Maintain the comprehensive conversion potential of arable land. Farmland pixels predicted based on the U-net model Transition probability, It is arable land like a yuan Neighborhood competition weight, It is a constraint of land use demand;
[0054] Reverse simulation rule process: Loading Land Use Map (Year) The conversion rule is defined as farmland removal, and the CARS module is based on all farmland pixels. Maintain the comprehensive conversion potential of arable land Generate all farmland pixels according to A table sorted in ascending order of values, then removing the top values from the sorted table one by one. One pixel, in The status was changed to non-arable land, and the update was obtained. Land Use Map (Year) .
[0055] The U-net model structure is as follows: it adopts a symmetrical encoder-decoder structure and includes skip links; the encoder consists of four downsampling stages, each containing two... Convolutional layer, followed by batch normalization and LeakyReLU activation, with the end using... Max pooling; each stage of the decoder begins with The transposed convolution is upsampled, halving the feature channels, and then concatenated with the corresponding layer features from the encoder, followed by two more passes. Convolutional layers fuse features; the output layer uses... Convolution and Sigmoid activation output the farmland probability for each pixel.
[0056] The specific steps for training the U-net model are as follows:
[0057] The input to the U-net model is organized into an image format, specifically the input feature X: integrating all spatialized driving factor layers of the baseline year, with the input being a single... tensor, and
[0058] These represent the number of rows and columns of the raster in the study area, respectively. The number of all spatialized driving factor layers, including two main parts: natural environmental factors and socio-economic factors;
[0059] The training label data consisted of a binary farmland distribution map from the baseline year, which was used as the training label. The farmland distribution map was then standardized for resolution to align with the driving factor data, resulting in a dimensional dataset. The binary mask.
[0060] Example 2
[0061] This embodiment further defines Embodiment 1, employing a deep learning-enhanced historical reconstruction system for cultivated land patterns to achieve high-precision reconstruction of the cultivated land patterns in the Northeast Black Soil Region from 1949 to 1979. (Refer to...) Figure 1 The specific steps are as follows.
[0062] Multi-source data input module:
[0063] 1. Land use / cover data
[0064] The land use data in this embodiment was acquired from two aspects: first, to provide high-precision baseline years and training sample data for the simulation framework; and second, to provide verification evidence for the reliability of the reconstruction results by collecting independent historical data. The land use / cover remote sensing monitoring dataset comes from the Resource Science and Data Center (RESDC) of the Chinese Academy of Sciences. This dataset is based on Landsat imagery and has high classification accuracy and spatiotemporal consistency. Image data from 1980 was acquired for subsequent research.
[0065] To independently verify the accuracy of the reconstruction results of this invention, various historical geographical data from the study period (1949-1979) were collected. Historical map data primarily consisted of 1:250,000 topographic maps of China drawn by the U.S. Army Cartography Bureau in the 1940s and 1950s. These maps detailed settlements, roads, and large areas of paddy fields and dry land. Through high-precision scanning, georegistration, and visual interpretation, the cultivated land information in these maps was vectorized, forming reference maps of cultivated land distribution at key time points. Simultaneously, local gazetteers, agricultural records, national county-level agricultural statistics, and construction archives of some state-owned farms in sample counties and cities in Northeast China were systematically consulted. Key textual descriptions and statistical data regarding "reclaimed land area," "location of reclamation areas," and "construction of water conservancy facilities" were extracted. This information provided corroborating evidence for determining the location and chronology of the reconstructed cultivated land expansion.
[0066] 2. Driving factors of the natural environment
[0067] The natural environment driving factor system adopted in this embodiment includes four main aspects: climate, soil, topography, and hydrology. Climate driving factors encompass nine specific indicators, including average annual temperature, growing season temperature, total precipitation, and average annual wind speed. These data were obtained by integrating multiple authoritative datasets, including those from the China Meteorological Administration Data Service Center (CMDSC), the CRU TS dataset, the National Climate Data Center (NCDC), and the Center for Hydrometeorology and Remote Sensing at the University of California, Irvine. Soil driving factors involve eight physicochemical properties, including soil texture, bulk density, organic carbon content, and pH value. The data primarily comes from the World Soil Database (HWSD V2.0) and the China Soil Dataset for Land Surface Simulation (CSDLv2). Elevation data in the topography driving factors were directly extracted from the SRTM 90m Digital Elevation Database (v4.1), while slope data was calculated based on this elevation data. Finally, hydrological driving factors, including rivers, lakes, and reservoirs of levels 1-4 within the study area, were obtained from a global reservoir and lake surface area change dataset covering 1980-2022. The detailed parameters and source information of each driving factor are shown in Table 1.
[0068] Table 1. Natural Environment Data and Sources for Northeast China:
[0069]
[0070] 3. Socioeconomic driving factors
[0071] Socioeconomic drivers fall into three categories: population, economy, and roads. Population variables include settlements, population density, labor force, census data, and nighttime light. The raw data for settlements, labor force, and previous censuses primarily come from the Chinese censuses (1953, 1964, 1982, 1990, 2000) and historical / reconstructed datasets (such as HYDE 3.2, GWP, Popdynamics, etc.). Based on this, a multi-factor weighting method (Dasymetric Mapping) is used, employing topography, land use type (non-built-up areas), and distance from waterways as spatial weights to spatialize the county-level total population, generating 1 km resolution population density rasteres for two key time points: 1953 and 1964. Subsequently, linear interpolation is performed on these two baseline time points to obtain the annual population density dynamic layer from 1949 to 1979. Nighttime light data uses global NPP (Natural Population Percentage). VIIRS (and its VIIRS) Like historical reconstruction datasets, these serve as supplementary indicators of population / activity intensity (typical resolution approximately 500 m). Economic variables include GDP, industrial output, proportion of non-agricultural land, and electricity consumption. Data sources include the Ministry of Agriculture and Rural Affairs' "National Agricultural Statistics," the Resource and Environmental Science Data Center of the Chinese Academy of Sciences, FAO STAT, and existing 1 km × 1 km gridded real GDP and electricity consumption datasets (typical coverage periods 1950–2020 (statistics), 1992–2019 (grid electricity / GDP)). Some economic indicators are based on county-level statistics and rasterized to the research resolution. Road variables include national / provincial / county roads, railways / expressways, and road accessibility indicators, primarily sourced from the China's Surface Transport System Database (time coverage approximately 1992–2020). These are used to extract road network levels, road density, and accessibility measures from any raster to the nearest road / railway. All of the above data types undergo spatiotemporal resolution and time-period consistency processing to match the unified raster used in the research analysis. Detailed parameters and source information for each driving factor are shown in Table 2.
[0072] Table 2 Socioeconomic data for Northeast China and their sources:
[0073]
[0074] Historical arable land area reconstruction module:
[0075] This embodiment first uses linear regression to reconstruct the total cultivated land area at the county level. This embodiment assumes that, during historical periods of relatively stable agricultural technology, there is a robust positive correlation between total grain output and the area of land cultivated. To verify and utilize this relationship, agricultural statistics for counties (cities) in the three northeastern provinces were collected for years with relatively reliable data (1949-1960 and after 1980), extracting two sets of time-series data: total grain output and officially reported cultivated land area. An ordinary least squares (OLS) regression model was constructed and verified, with the following formula:
[0076] ;
[0077] in, For arable land area, For grain production, For regression coefficients, For the intercept term, This is the error term.
[0078] Deep learning-enhanced spatial simulation module:
[0079] 1. U-Net model training and farmland existence probability map generation
[0080] This embodiment uses the U-Net semantic segmentation network to learn the ideal spatial distribution of cultivated land in the context of a comprehensive geographical environment. Its training does not rely on historical "transformation events" but learns the global mapping relationship between land use "state" and environmental factors.
[0081] Input variables for the U-Net model: The U-Net model requires the input to be organized into an image format. Input feature X: Integrates all spatialized driving factor layers for the baseline year; the input is a single... The tensor. This represents the number of rows and columns of the raster in the study area. The number of spatialized driving factor layers is given. Spatialization involves creating a unified raster layer from the data, comprising two main parts: natural environmental factors (see Table 1) and socioeconomic factors (see Table 2). The time-series variables will be selected based on the ease of acquisition, using driving factor data as close as possible to the year being simulated. For example, when simulating the state in 1970, data such as population density and GDP corresponding to 1970 will be used, representing dynamic data. For slowly changing factors such as topography and soil, static data will be used. Label data Y preparation: A binarized cultivated land distribution map of the baseline year is used as the training label. This map is generated by extracting the "cultivated land" category from a high-precision land use classification map. To ensure alignment with the driving factor data, the cultivated land distribution map undergoes the same spatial reference and resolution standardization processing, resulting in a dimension of [missing data]. The binary mask.
[0082] Network Architecture and Optimization: 1) Network Architecture: A classic encoder-decoder symmetric structure is adopted, including skip links. The encoder consists of four downsampling stages, each containing two 3D models. Three convolutional layers, followed by batch normalization and LeakyReLU activation, with 2 at the end. 2. Max pooling; each stage of the decoder begins with 2. 2. Transposed convolution is used for upsampling, halving the feature channels, and then concatenated with the corresponding layer features from the encoder. This is then processed through two 3... Three convolutional layers fuse features. The output layer uses one convolutional layer. 1) Convolution and Sigmoid activation output the probability of farmland for each pixel. 2) Loss function design: To address the spatial imbalance in farmland distribution, a weighted combination of Dice loss and binary cross-entropy loss is used, with the following formula:
[0083] ;
[0084] in This represents the total loss value. For real labels, To predict probabilities, To prevent division by zero in the smoothing term, To balance the weights, This represents the effective number of pixels for cultivated land.
[0085] Model training and evaluation: Using the Adam optimizer, the initial learning rate was set to... to A cosine annealing strategy was employed to dynamically reduce the learning rate during the performance plateau period, promoting model convergence to a better local minimum. The training epoch was set to 100 epochs, and an early stopping strategy was used: training was terminated when the validation set loss no longer decreased within 10-20 consecutive epochs to avoid overfitting. At the beginning of training, the driving factor dataset was spatially divided into a training set (70%), a validation set (15%), and a test set (15%). The test set was used to calculate the overall accuracy, Kappa coefficient, and intersection-over-union ratio of farmland categories to objectively evaluate the model's generalization ability. Models with a Kappa coefficient greater than 0.75 were used for subsequent simulations.
[0086] Farmland diagnostic probability map generation using the U-Net model: After completing model training, for any target year in the inverse simulation... All driving factor data for that year The data is input into the U-Net model. After one forward propagation, the model outputs a diagnostic probability map of arable land presence for the target year. The value of each pixel in this image. This represents the year. Given the specific environmental context, the probability of this location being spatially reasonable as arable land.
[0087] 2. Definition of inverse simulation rules based on diagnostic probability
[0088] This embodiment uses the U-Net model to optimize the inverse simulation strategy of the traditional CA model. Based on a high-precision remote sensing land use map from 1980, it reconstructs the historical evolution path of cultivated land in the Northeast Black Soil Region by tracing back year by year to 1949. Unlike traditional forward simulation, the key to inverse simulation is determining which cultivated land cells should be removed "first" (i.e., restored to their previous non-cultivated state) when time is reversed. The basic assumption of this embodiment is that during periods of cultivated land expansion, the worst-condition and most unsuitable marginal lands are usually the last to be cultivated; therefore, when retracing time, these lands should also be "reverted" to non-cultivated land first.
[0089] The reverse conversion rule based on comprehensive conversion potential: In contrast to conventional forward simulation which seeks the region with the highest development potential for "expansion," this embodiment defines the conversion rule as "farmland removal" in reverse simulation. The dilemma of traditional methods in reverse simulation lies in their reliance on expansion samples from the "past to the future." It can only answer "where is it easier to become farmland," but cannot directly answer the inverse question, "which farmlands should be reverted to non-farmland first." Directly inverting the expansion samples is logically crude and fails to learn the spatial patterns specific to the contraction process. This embodiment no longer directly calculates the "development potential for transformation" but instead calculates a "probability of the current state's rationality based on the global context." Given all current geographic environment drivers, the model predicts the probability that a pixel "should" be in a cultivated land state; it's an end-to-end, data-driven state diagnosis probability. If the current pixel is cultivated land, but... A low value indicates that, based on its geographical environment (potentially steep slopes or remoteness), the model determines it "shouldn't" be arable land. Therefore, during reverse contraction, it will be preferentially removed. This simulation logic perfectly aligns with the reverse logic of "last land reclamation": the last marginal land to be reclaimed is precisely that land with poorer environments and lower suitability for expansion in the forward model, yet it is still reclaimed. The U-Net model, by learning the complex linear relationship between environmental characteristics and land use status, can identify these ecologically mismatched reclaimed arable lands.
[0090] 3. Reverse simulation execution and spatial pattern evolution
[0091] Starting with the land use status quo in 1980, and following the principle of "reclaiming land first and then reclaiming it later," priority was given to removing arable land pixels with the lowest conversion potential. In each iteration, the removal ratio was dynamically adjusted based on macro-level constraints to ensure that the annual arable land area was consistent with the reconstruction results. Through a neighborhood competition mechanism, considering patch shape, size, and spatial correlation, the rationality of the landscape pattern was maintained.
[0092] The simulation process starts with the land use map of 1980 and iterates backward to 1949, using years as the time step. In each year... In the reverse iteration (i.e. from) Annual Simulation (Regarding the state of the year), the operating mechanism of the CARS module has been adjusted as follows:
[0093] Macroeconomic aggregate reverse constraint: First, the historical arable land area reconstruction module calculates... Annual net increase / decrease in cultivated land pixels As a hard termination condition for the iteration that year, this negative demand will serve as a hard termination condition for the iteration that year.
[0094] Comprehensive conversion potential calculation: for the current year For each arable land cell on the map, its comprehensive conversion potential is calculated using deep learning. After replacing the traditional inference function with the probability, the specific evaluation formula for CARS is as follows:
[0095] ;
[0096] in, It is arable land like a yuan Maintain the comprehensive conversion potential of arable land. Farmland pixels predicted based on the U-net model Transition probability, It is the neighborhood competition weight, represented by the cultivated land pixel. Within the central area, the proportion of arable land, This refers to land use demand constraints. It represents two years (...). and The difference in the number of pixels in cultivated land, and the numerical value. equal.
[0097] Reverse simulation algorithm flow: Loading Land Use Map (Year) and corresponding And the number of farmland pixels to be removed in this iteration Based on the calculation results of the above formula, all cultivated land pixels are generated according to... A table sorted in ascending order of values, then removing the top values from the sorted table one by one. One pixel, in The status was changed to "non-arable land". The land use map was updated. Then, using this as the input for the next time period, repeat the above steps until the simulation of 1949 is completed, and output the reconstructed farmland pattern sequence.
[0098] The reverse evolution of spatial patterns: This removal process repeats itself until the total cultivated land area on the current map is equal to... The annual target values are exactly the same. In this process, the neighborhood effect mechanism of the PLUS model is interpreted in reverse: isolated arable land pixels located on the edge of patches and with low suitability will be removed earlier due to their unstable spatial structure and poor natural endowment. This makes the simulated land use "shrinkage" process present a realistic spatial pattern of retreat from marginal areas to core high-quality areas, rather than a random reduction.
[0099] Example 3
[0100] This embodiment further defines Embodiment 2 by adding a system verification and uncertainty analysis module to the existing deep learning-enhanced historical reconstruction system of cultivated land patterns. Figure 2 As shown. To ensure the scientific validity and reliability of the historical farmland pattern reconstruction results proposed in this invention, this embodiment establishes a comprehensive verification and uncertainty quantification system driven by multi-dimensional and multi-source data. This system not only evaluates the consistency between the reconstruction results and independent historical evidence, but also systematically quantifies the sources of uncertainty in each link of the system chain and their impact on the final result.
[0101] 1. Accuracy Verification Method
[0102] 1) Quantitative spatiotemporal consistency verification
[0103] This embodiment uses a third-party historical map, independent of the system input data, as a reference. A confusion matrix is generated through spatial overlay analysis to quantitatively evaluate the spatiotemporal consistency of the reconstruction results. This method aims to assess the degree of agreement between the reconstructed spatial distribution of cultivated land and the actual records at a specific historical point in time at the pixel scale. The core evaluation indicators are overall accuracy and the Kappa coefficient, which measure the improvement of the classification results compared to random classification and more objectively reflect the level of consistency. Their calculation formulas are as follows:
[0104] ; ; In the formula for Kappa coefficient; Total number of categories; This represents the total number of pixels; The th in the confusion matrix Line 1 The number of cells in the column (i.e., the number of cells that are correctly classified); and The first Line 1 The total number of pixels in the column.
[0105] 2) Qualitative corroborating evidence and case verification
[0106] To verify the system's responsiveness to land use changes driven by key historical events, this embodiment introduces a qualitative corroborating method based on historical documents. Descriptive records of large-scale land reclamation in specific regions (such as state-owned farms or reclamation areas) during specific periods are extracted from local gazetteers and other documents, and compared spatiotemporally with the annual dynamic evolution results of cultivated land generated in Embodiment 2 to verify the logical rationality and authenticity of the reconstructed pattern. The verification logic can be visualized as follows:
[0107] ;
[0108] In the formula, The qualitative verification score is 1 for a match and 0 for a non-match. These represent events recorded in literature and events simulated in models, respectively. L and T These represent the spatial location and time at which the event occurred, respectively. This is within an acceptable range of time error.
[0109] 2. Quantification of Uncertainty
[0110] This embodiment considers two main sources of error and assigns probability distributions to them to simulate the propagation effect of uncertainty, thereby providing a reliable confidence assessment of the reconstruction results. The main sources of error include historical statistical data errors (following a normal distribution with a standard deviation of 5%-10%) and model parameter uncertainties (random forest tree number ~ Uniform(500, 1000), U-Net learning rate ~ LogNormal(0.001, 0.5), neighborhood competition weights ~ Normal(0.5, 0.1)). This embodiment runs 1000 iterations of Monte Carlo Simulation (MCS). In each iteration, perturbation values are randomly sampled from the error distribution, and area reconstruction and inverse simulation are re-executed to generate a perturbed farmland pattern map, quantifying the impact of uncertainty on the output results.
[0111] Example 4
[0112] This embodiment is a further limitation of embodiment 2.
[0113] Results output and application:
[0114] The direct result of Example 2 is the generation of a dynamic reconstruction dataset of historical cultivated land patterns covering the Northeast Black Soil Region from 1949 to 1979 during the period without remote sensing. This dataset is output in the standard GeoTIFF format with multiple layers and multiple attributes, along with accompanying metadata files, ensuring data exchangeability and repeatability. Figure 3 As shown, the reconstruction results obtained through this system clearly reveal the spatiotemporal trajectory of cultivated land expansion from the core plain area to the peripheral hilly area, providing reliable data support for understanding the historical process of land use change in the black soil region. This method has strong versatility and portability, and can be applied to long-term land use reconstruction studies in other regions. Figure 3 In the diagram, A, B, C, and D represent the distribution of cultivated land in 1949, 1959, 1969, and 1979, respectively. Using this system, this embodiment reconstructs the county-level cultivated land distribution in the Northeast Black Soil Region from 1949 to 1979. The accuracy is further validated by combining this data with existing historical atlases from different periods. The overall accuracy is 0.84, and the kappa coefficient is 0.76, which fully meets the needs of further application analysis. The confusion matrix is shown in Table 3, indicating that there are relatively few misclassifications of cultivated land overall.
[0115] Table 3. Confusion matrix of total land use in selected areas at different times: (unit: km²) 2 )
[0116]
[0117] Furthermore, its estimation results were compared with those of three existing technologies, HYDE3.2, Jia et al. (2024) and Yu et al. (2021), during the period of 1949–1979 (see Table 4).
[0118] Source for HYDE3.2: Klein Goldewijk, K., Beusen, A., Doelman, J., andStehfest, E.: Anthropogenic land use estimates for the Holocene-HYDE 3.2, Earth Syst. Sci. Data, 9, 927–953.
[0119] In Jia et al. (2024): Yu, Z. and Lu, C.: Historical cropland expansion and abandonment in the continental US During 1850 to 2016, GlobalEcol. Biogeogr., 27, 322–333.
[0120] In Yu et al. (2021): Jia, R., Fang, X., Yang, Y., Yokozawa, M., andYe, Yu.: A 28 time-points cropland area change dataset in Northeast China from 1000 to 2020, figshare [data set].
[0121] The results show that all datasets reflect the continuous expansion trend of cultivated land area during this period, but the results of Example 2, based on the inverse evolution model, are significantly more reliable. Specifically, the estimates in Example 2 are significantly higher than those in the global-scale HYDE3.2 dataset, and are generally higher than those in the study by Jia et al. (2024). Compared with the dataset of Yu et al. (2021), the results in Example 2 are closest in magnitude: slightly lower in 1949 and 1959, but surpassed in 1969 (262,100 square kilometers in Example 2, 260,900 square kilometers in Yu et al.), indicating differences in the details of dynamic changes between the two.
[0122] Table 4. Comparison of estimation results of cultivated land area from the 1940s to the 1970s using different technologies: (Unit: 10,000 km²) 2 )
[0123]
[0124] Example 5
[0125] This embodiment illustrates the application of the technical effects of the deep learning-enhanced historical reconstruction system for farmland patterns described in this invention.
[0126] This system provides a long-term, continuous, and reliable data foundation for solving key issues in resources, environment, and sustainable development in the black soil region. Its main applications include, but are not limited to:
[0127] This system provides a long-term, continuous, and reliable data foundation for solving key issues in resources, environment, and sustainable development in the black soil region. Its main applications include, but are not limited to:
[0128] 1) Assessment of Black Soil Resource Evolution and Degradation: By overlaying the dataset obtained by this system with the spatial distribution map of soil properties (such as organic matter content and thickness), the historical evolution trajectory of the quantity and quality of black soil resources can be accurately traced back, and the contribution rate of cultivated land expansion to black soil degradation in different historical stages can be quantitatively assessed, providing a historical perspective attribution analysis for the implementation of the "storing grain in the land" strategy.
[0129] 2) Agricultural ecosystem service simulation: As key input data, the datasets obtained by this system can drive ecological process models (such as InVEST and SWAT) to reconstruct the dynamic changes in key ecosystem service functions such as carbon fixation, water conservation, and soil retention in the black soil region over the past 70 years, and reveal the long-term impact of agricultural development on the regional ecological security pattern.
[0130] 3) Post-evaluation of the effects of land use policies: By accurately linking the spatiotemporal hotspots of arable land expansion with the implementation time and region of major agricultural policies and ecological projects, the actual regulatory effects of various policies on arable land use patterns can be scientifically evaluated from a spatially explicit perspective, providing empirical evidence for the optimization of future national spatial planning and ecological compensation policies.
[0131] 4) Provide a benchmark for future scenario prediction: The long-term arable land pattern reconstructed by this system is an indispensable historical benchmark data for calibrating and validating global change models and land use change models, which can significantly improve the accuracy and reliability of predictions of future arable land use changes and their ecological and environmental effects in the black soil region.
Claims
1. A deep learning-enhanced system for historical reconstruction of arable land patterns, characterized in that, The system includes a multi-source data input module, a historical cultivated land area reconstruction module, and a deep learning-enhanced spatial simulation module. Multi-data input module: acquires land use / cover remote sensing monitoring data, natural environment data, and socio-economic data of the area to be reconstructed; obtains remote sensing cultivated land map through land use / cover data, obtains natural environment driving factors through natural environment data, and obtains socio-economic driving factors through socio-economic data. Historical cultivated land area reconstruction module: Construct a grain output-cultivated land area regression model for the study area, and reconstruct the cultivated land area of the area to be reconstructed based on the grain output-cultivated land area regression model. By calculating the difference in cultivated land area between adjacent years, the net increase or decrease in cultivated land pixels per year is obtained. Deep learning-enhanced spatial simulation module: Through the inverse simulation strategy based on the PLUS model, the range of years to be reconstructed is identified, and the cultivated land area is reconstructed year by year. When reconstructing the cultivated land area year by year, from Starting in 2010, it was gradually restructured year by year. The year, specifically: S41. Generation of the probability map of cultivated land existence: Using the trained U-net model, input... The corresponding natural environment driving factors and socio-economic driving factors for each year were obtained. Probability of arable land existence corresponding to the year Probability map of arable land In China, use Represents the corresponding pixel The transition probability; S42. Reverse Simulation Execution: The reverse simulation is performed using the CARS module of the PLUS model, starting from... Starting in [year], the operating mechanism of the CARS module was adjusted and [further adjustments were made]. The reconstruction of the cultivated land pattern in the region awaiting reconstruction; S43. Refactoring by Year: Repeat steps S41-S42 until completion. The arable land pattern in the region awaiting reconstruction will be restructured.
2. The deep learning-enhanced historical reconstruction system for cultivated land patterns according to claim 1, characterized in that, Natural environmental drivers include climate, soil, topography, and hydrology, while socioeconomic drivers include population, economy, and roads.
3. The deep learning-enhanced historical reconstruction system for cultivated land patterns according to claim 2, characterized in that, The formula for the grain yield-arable land area regression model is: ; in, For arable land area, For grain production, For regression coefficients, For the intercept term, This is the error term.
4. The deep learning-enhanced historical reconstruction system for cultivated land patterns according to claim 3, characterized in that, The specific adjustments to the CARS module's operating mechanism are as follows: Macroeconomic aggregate reverse constraint: First, the historical arable land area reconstruction module calculates... Annual net increase / decrease in cultivated land pixels As a hard termination condition for the iteration at that time, deep learning was used in the evaluation formula of CARS. The transition probability replaces the traditional inference function, and the specific formula is as follows: ; in, It is arable land like a yuan Maintain the comprehensive conversion potential of arable land. Farmland pixels predicted based on the U-net model Transition probability, It is arable land like a yuan Neighborhood competition weight, It is a constraint of land use demand; Reverse simulation rule process: Loading Land Use Map (Year) The conversion rule is defined as farmland removal, and the CARS module is based on all farmland pixels. Maintain the comprehensive conversion potential of arable land Generate all farmland pixels according to A table sorted in ascending order of values, then removing the top values from the sorted table one by one. One pixel, in The status was changed to non-arable land, and the update was obtained. Land Use Map (Year) .
5. The deep learning-enhanced historical reconstruction system for cultivated land patterns according to claim 4, characterized in that, The U-net model structure is as follows: it adopts a symmetrical encoder-decoder structure and includes skip links; the encoder consists of four downsampling stages, each containing two... Convolutional layer, followed by batch normalization and LeakyReLU activation, with the end using... Max pooling; each stage of the decoder begins with The transposed convolution is upsampled, halving the feature channels, and then concatenated with the corresponding layer features from the encoder, followed by two more passes. Convolutional layers fuse features; the output layer uses... Convolution and Sigmoid activation output the farmland probability for each pixel.
6. The deep learning-enhanced historical reconstruction system for cultivated land patterns according to claim 5, characterized in that, The specific steps for training the U-net model are as follows: The input to the U-net model is organized into an image format, specifically the input feature X: integrating all spatialized driving factor layers of the baseline year, with the input being a single... tensor, and These represent the number of rows and columns of the raster in the study area, respectively. The number of all spatialized driving factor layers, including two main parts: natural environmental factors and socio-economic factors; The training label data consisted of a binary farmland distribution map from the baseline year, which was used as the training label. The farmland distribution map was then standardized for resolution to align with the driving factor data, resulting in a dimensional dataset. The binary mask.
7. A deep learning-enhanced method for historical reconstruction of arable land patterns, characterized in that, The method uses the deep learning-enhanced historical reconstruction system for cultivated land patterns as described in claim 6, specifically as follows: Acquire land use / cover remote sensing monitoring data, natural environment data, and socio-economic data of the area to be reconstructed for arable land pattern; obtain remote sensing arable land maps through land use / cover data, obtain natural environment driving factors through natural environment data, and obtain socio-economic driving factors through socio-economic data. A regression model of grain output-arable land area was constructed for the study area. Based on the regression model of grain output-arable land area, the regional arable land area was reconstructed in the area to be reconstructed. By calculating the difference in arable land area between adjacent years, the net increase or decrease in the number of arable land pixels per year was obtained. By using a reverse simulation strategy based on the PLUS model, the range of years to be reconstructed is identified, and the cultivated land area is reconstructed year by year.