A land use change scenario simulation method considering driving factor time lag
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SOUTHWEST PETROLEUM UNIV
- Filing Date
- 2026-05-14
- Publication Date
- 2026-08-07
AI Technical Summary
[0005]本发明提供了一种顾及驱动因子时滞的土地利用变化情景模拟方法,拟解决现有技术未对驱动因子的时滞特征进行针对性处理,导致难以刻画社会经济因子从发生变化到最终导致土地覆盖更替的时间延迟过程
将每个栅格像元转换为不同目标土地类型的综合转换概率作为轮盘赌的权重,每个目标土地类型对应轮盘中的一个区域,区域大小与对应土地类型的综合转换概率成正比,通过随机选取轮盘区域的方式,确定对应栅格像元最终的土地类型。
Smart Images

Figure CN122528631A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of land use change simulation technology, and more specifically, to a land use change scenario simulation method that takes into account the time lag of driving factors. Background Technology
[0002] Land use change is a core area of research on global environmental change and sustainable development. Accurately simulating and predicting future land use evolution trends is of great practical significance for scientifically delineating ecological protection red lines and urban development boundaries, optimizing the spatial layout of the country, achieving a balance of regional resource and environmental carrying capacity, and refining land use management. This need is even more urgent in the context of my country's in-depth urbanization process and the proposal of dual-carbon goals.
[0003] Land use change is a complex and dynamic process driven by both the natural environment and socio-economic factors. Natural environmental factors include topography, meteorology, and soil, while socio-economic factors include population, GDP, and road network. Scientifically, the impact of these factors on land use is not entirely immediate; there are often significant time lags and cumulative effects. For example, the development of transportation infrastructure can take several years to drive the development of surrounding land, and the effects of ecological restoration policies can only be seen through long-term biological succession.
[0004] Currently, existing land use simulation models (such as traditional cellular automata (CA), FLUS models, and PLUS models) are mainly based on static mapping logic for simulation. Their core flaw lies in ignoring the time-lag effect of driving factors on land use change, specifically manifested as follows: Most existing models use single-phase or simplified time-series data as input, without specifically addressing the time lag characteristics of driving factors, making it difficult to characterize the time delay process from changes in socioeconomic factors to the eventual land cover replacement. Summary of the Invention
[0005] This invention provides a land use change scenario simulation method that takes into account the time lag of driving factors, aiming to solve the problem that the existing technology does not specifically process the time lag characteristics of driving factors, making it difficult to characterize the time delay process from changes in socio-economic factors to the final land cover replacement.
[0006] A land use change scenario simulation method that takes into account the time lag of driving factors includes the following steps: S1: Standardize and perform time-delay feature engineering on the collected multi-source heterogeneous data, extract the driving factors from the multi-source heterogeneous data, and construct the time-delay sequence set of the driving factors as the basic input sequence to characterize the time accumulation effect. S2: Using an artificial neural network with the aforementioned basic input sequence as input, the high-dimensional nonlinear relationship between driving factors and land use conversion probability is mined to obtain the land use suitability probability and the relative contribution of each driving factor. S3: Using a dynamic Bayesian network, the land use suitability probability is used as the prior probability of land state. Combined with the significant lag period of each driving factor and the extracted driving factors, a causal relationship between driving factors and land state is constructed in the dynamic Bayesian topology. The time cumulative effect and causal lag logic of driving factors are quantitatively captured to obtain the dynamic transfer probability of considering the time lag effect. S4: Using the dynamic transition probability as the core input, the comprehensive transition probability is obtained by coupling calculation with the neighborhood effect of the adaptive cellular automaton, the adaptive inertia coefficient set by the land use type characteristics, and the land type conversion cost. Then, through multi-scenario iterative allocation, the spatiotemporal pattern simulation of land use change is completed.
[0007] This invention extracts driving factors by standardizing and time-lag feature engineering multi-source heterogeneous raw data, and constructs a time-lag sequence set of driving factors to achieve targeted processing of the time-lag characteristics of driving factors. Subsequently, an artificial neural network is used to mine the nonlinear relationship between driving factors and land use conversion probability with the time-lag sequence set as input, quantifying the relative contribution of each driving factor. A dynamic Bayesian network is used, with the land use suitability probability as the prior probability, combined with the driving factors and their significant lag periods, to construct the causal relationship between driving factors and land state, quantitatively capturing the time cumulative effect and causal lag logic of driving factors, and characterizing the time delay process from changes in socio-economic factors to land cover replacement, outputting the dynamic transfer probability that takes into account the time lag effect. Finally, an adaptive cellular automaton is used, with the dynamic transfer probability as the core input, coupled with neighborhood interaction, adaptive inertia coefficient and conversion cost to calculate the comprehensive conversion probability, and the land use spatiotemporal pattern simulation is completed through multi-scenario iterative allocation.
[0008] Preferably, the multi-source heterogeneous data includes land use remote sensing data, natural geographic data, and socio-economic statistical data.
[0009] Preferably, the standardization process and time delay feature engineering include the following steps: By using a GIS platform, land use remote sensing data, natural geographic data, and socio-economic statistical data are unified to the same coordinate system and preset raster spatial resolution to obtain standardized multi-source heterogeneous data. Then, natural driving factors and socio-economic driving factors are extracted from standardized multi-source heterogeneous data, with natural driving factors extracted from standardized natural geographic data and socio-economic driving factors extracted from standardized socio-economic statistical data. For each extracted driving factor, the mutual information between the sequence of each driving factor at different lag times and the current land state is calculated to determine the significant lag period of the impact of each driving factor on land use. The lag period corresponding to the maximum value of the mutual information is the significant lag period of the corresponding driving factor. Based on the determined significant lag period, a time lag sequence set of the corresponding driving factor is constructed, and the time lag sequence set is used as the basic input sequence.
[0010] Preferably, the natural geographic data includes at least one of topographic data, meteorological data, and soil data, and the corresponding natural driving factors include at least one of topographic factors, meteorological factors, and soil factors.
[0011] Preferably, the socioeconomic statistics include at least one of population data, GDP data, and road network data, and the corresponding socioeconomic driving factors include at least one of population factors, GDP factors, and road network factors.
[0012] Preferably, step S2 includes the following steps: The basic input sequence is used as the input of the artificial neural network, and the artificial neural network processes the data to output the static suitability probability of the mutual conversion of various land uses. The Garson algorithm is used to analyze the relative contribution percentage of each driving factor to the probability of land use type conversion by calculating the absolute weight product from the input layer to the hidden layer and from the hidden layer to the output layer in the artificial neural network, thus obtaining the relative contribution of each driving factor.
[0013] Preferably, the relative contribution is a standardized value between 0 and 1, and the relative contribution is sorted in descending order. Based on a preset threshold, driving factors with relative contributions greater than the threshold are selected as core driving factors to adjust the dynamic Bayesian network structure.
[0014] Preferably, step S3 includes the following steps: Dynamic Bayesian Network Construction: Define two types of nodes: land state nodes at each time step and driving factor observation nodes at each time step. In the dynamic Bayesian network, explicitly establish directed edges from the observation node corresponding to the significant lag period of each driving factor to the land state node at the current time step, representing the impact of each driving factor on the current land transformation before the significant lag period. Data from all nodes in the dynamic Bayesian network is collected, collinearity is tested and nodes with collinearity exceeding the threshold are removed. The remaining node data is discretized using the optimal binning and discrete classification standard methods and then organized into training and test sets according to time slices. Using the land use suitability probability as the prior probability of land state nodes, the expectation-maximization algorithm is used to handle missing data, and the conditional probability of each node under different combinations of parent node states is calculated to obtain the dynamic transition probability of time lag effect.
[0015] Preferably, step S4 includes the following steps: The dynamic transfer probability of considering the time lag effect, the neighborhood interaction coefficient corresponding to the horizontal neighborhood factor of the adaptive cellular automaton, the adaptive inertia coefficient set based on the natural attributes of land use type, and the land type conversion cost set based on the difficulty of land use type conversion are coupled and calculated to obtain the comprehensive conversion probability of each raster cell from the initial land type to the target land type at each time. Based on actual land use planning needs, at least three scenario parameters are set, and an adaptive inertial mechanism is adopted to adjust the total demand balance of each land type. Using the roulette wheel algorithm, based on the obtained comprehensive conversion probability, the land state of each grid cell is changed, thereby simulating the spatiotemporal pattern of land use under different scenarios.
[0016] Preferably, the execution steps of the roulette wheel betting algorithm are as follows: The overall conversion probability of each raster cell to different target land types is used as the weight of the roulette wheel. Each target land type corresponds to a region in the roulette wheel, and the size of the region is proportional to the overall conversion probability of the corresponding land type. The final land type of the corresponding raster cell is determined by randomly selecting a region in the roulette wheel.
[0017] The beneficial effects of this invention include: This invention constructs a multi-source heterogeneous driving factor system, utilizes artificial neural networks to extract the nonlinear contribution weights of natural environment and socio-economic factors to land use transformation; by introducing a high-order Markov chain structure of dynamic Bayesian network, it establishes a time accumulation function of driving factors at different time scales to capture the lag response probability of socio-economic policies and climate fluctuations to land use change; finally, the lag-corrected dynamic transfer probability is injected into the spatial evolution framework of cellular automata (CA) to achieve allocation under local spatial competition, which solves the limitations of the instantaneous response assumption between driving factors and land change, and can quantitatively identify the lag influence cycle and intensity of different factors.
[0018] This embodiment integrates the time-delay effect of driving factors into the model building process for the first time, combining the nonlinear mapping capability of ANN, the time-series probabilistic reasoning capability of DBN, and the spatial evolution capability of CA. It breaks through the key limitations of traditional models that ignore the time-delay effect of driving factors, lack cumulative dynamics representation, and have weak uncertainty reasoning capability, and realizes the leap from static mapping to dynamic causality. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a flowchart illustrating the overall process framework of an embodiment of the present invention. Detailed Implementation
[0021] To make the technical problems, solutions, and beneficial effects of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0022] See Figure 1 As shown, this embodiment provides a land use change scenario simulation method that takes into account the time lag of driving factors. It utilizes the powerful nonlinear feature mapping capability of Artificial Neural Networks (ANNs) to construct a multi-temporal set of driving factors, extracting the initial contribution weights of natural environment and socio-economic factors to land conversion. Secondly, it introduces a Dynamic Bayesian Network (DBN) with advantages in time series reasoning. By constructing high-order Markov chains and directed arcs across time slices, it quantitatively characterizes the cumulative effect and time-lag influence of driving factors in the time dimension, correcting the causal logical feedback missing in traditional models and solving the limitation of the immediate response assumption between driving factors and land change. Specifically, it includes the following steps: S1: Collect land use remote sensing data, natural geographic data, and socio-economic statistics, and use a GIS platform to unify the land use remote sensing data, natural geographic data, and socio-economic statistics to the same coordinate system and spatial resolution (such as 30m grid) to obtain standardized multi-source heterogeneous data. In this embodiment, obtaining standardized multi-source heterogeneous data includes the following steps, exemplarily: Spatial consistency handling: Since the driving factors come from different map projections or resolutions, it is necessary to ensure that they are perfectly aligned in geospatial space; a. Resampling: Unify all raster data (such as meteorological, topographic, and socioeconomic data) to the same spatial resolution as land use data (such as 30m or 100m).
[0023] b. Reprojection: Ensure all data uses the same coordinate system (such as WGS84 or projected coordinate system UTM) to avoid spatial misalignment.
[0024] c. Cropping and Masking: Cropping all factors using the study area boundary ensures that the cell row and column numbers are completely consistent.
[0025] Normalization / Standardization: Min-Max normalization is used to scale the data to the range [0,1].
[0026] Encoding of categorical variables: If the driving data contains categorical variables (e.g., soil type, administrative division), they cannot be directly input as numbers such as 1, 2, 3, etc. In this embodiment, one-hot encoding is used to convert N categories into N binary features to avoid the model mistakenly assuming that there is a size order between the categories.
[0027] Sampling balance: In suitability assessment, it is usually necessary to extract sampling points from the study area for model training; this embodiment adopts spatial stratified sampling to ensure that each land use type (such as forest land, cultivated land, construction land) and its converted / unconverted pixels are sampled evenly; to prevent the model prediction from being biased towards land types with large sample sizes.
[0028] Deterministic transformation and distance analysis: Some driving forces need to be transformed from vectors to rasters. For example, Euclidean distance can be used to perform distance analysis on vector features such as roads, rivers, and town centers to generate a continuous distance field; or density analysis can be used to perform kernel density analysis on population points or points of interest (POIs) to transform them into a continuous intensity raster.
[0029] The natural geographic data includes topographic data, meteorological data, and soil data; the socio-economic statistical data includes population data, GDP data, and road network data. Natural and socioeconomic driving factors are extracted from the standardized multi-source heterogeneous data. Natural driving factors include topographic factors (such as elevation, slope, and aspect), meteorological factors (such as annual precipitation and annual average temperature), and soil factors (such as soil texture and soil organic matter content). Socioeconomic driving factors include population factors (such as population density), GDP factors (such as GDP density), and road network factors (such as road network density).
[0030] Collinearity checking and feature selection: Although artificial neural networks can handle nonlinearity, if there is a strong correlation between the input driving factors (such as "distance from provincial highway" and "distance from main road"), it will increase model redundancy and lead to overfitting. VIF (variance inflation factor) detection is used: VIF < 10 is typically required. Alternatively, a correlation coefficient matrix can be used to remove redundant factors with extremely high correlation (e.g., r > 0.85).
[0031] Based on standardized multi-source heterogeneous data and all extracted driving factors, this study changes the existing model's single-temporal input mode. For each driving factor, a multi-year time series is constructed, and a sliding time window is set (the window size is set according to the actual situation of the target study area and the characteristics of the driving factors, with a value of 1 to 5 years, serving as a candidate lag period). By calculating the mutual information between the sequences of the corresponding driving factors at different lag times and the current land state, the significant lag period of each driving factor's impact on land use is determined. The specific expression is as follows: ; In the formula: Represents the lag sequence of driving factors With the current land status The mutual information value indicates that the larger the mutual information value, the stronger the correlation between the lagged sequence of the driving factor and the current land state. Represents the lagged sequence observations of driving factors With the current land status The joint probability; Represents the lagged sequence observations of driving factors The marginal probability; Indicates the current land status The marginal probability; Indicates time, It has a significant lag period; In this embodiment, for each driving factor, its mutual information value under different candidate lag periods k is calculated, and the lag period k corresponding to the maximum mutual information value is found. The lag period k with the maximum mutual information value is taken as the significant lag period of the corresponding driving factor's impact on land use. It should be noted that in actual engineering calculations, not all mutual information curves of driving factors can exhibit obvious maxima; there are special cases where there is no optimal peak value, mainly falling into two categories: Mutual information is monotonically decreasing and has no intermediate peak: If there is a very strong linear monotonically correlated relationship between the driving factor and the land use status, the influence of the driving factor on land use change will continue to decrease with the increase or decrease of the lag time. The corresponding mutual information value will monotonically decrease with the increase of the lag step k, and only reach the global maximum value when k=1. The curve has no intermediate peak. This phenomenon indicates that the lag effect of such driving factors decays quickly, and only the short-term lag effect is significant. The long-term cumulative effect is negligible.
[0032] Mutual information continuously increases, rises at the end, and has no peak. If the time search window range of the preset candidate lag period is too small, the actual optimal lag period of the driving factor exceeds the preset search range, which will cause the mutual information curve to continuously rise, without any downward inflection point or peak, and the overall performance is that the curve rises at the end. This phenomenon indicates that the time lag period of the corresponding driving factor is longer and the time accumulation effect is more lasting, and the existing time window cannot cover its complete lag effect process.
[0033] For the two types of abnormal operating conditions without obvious maxima mentioned above, this embodiment adopts the following approach to solve them: For cases where mutual information is monotonically decreasing and there is no intermediate peak, the optimal lag period k=1 is directly determined. It is assumed that this type of driving factor satisfies the first-order Markov property, and there is only a short-term 1-year lag effect with no long-term cumulative impact.
[0034] For the condition where mutual information continues to rise but there is no peak at the end: if the current preset time window cannot cover the real lag period, gradually expand the search range of lag period, continuously iterate to calculate the mutual information value corresponding to different lag steps until the mutual information curve shows a clear inflection point and a downward trend, take the lag step corresponding to the peak as the optimal significant lag period of the driving factor, and capture the long-cycle time lag and cumulative effect.
[0035] Based on the determined significant lag periods of each driving factor, a time lag sequence set for each driving factor is constructed as the basic input sequence. Specifically, for a driving factor with a significant lag period of k, its time lag sequence set includes the observations of the driving factor in the current year and the previous k years, as expressed below: ; In the formula: Express the time lag sequence of the i-th driving factor at time t; This represents the observed value of the i-th driving factor at the current time t; Let represent the observed value of the i-th driving factor at time tk; In this embodiment, by constructing a time lag sequence set for each driving factor, the time cumulative effect of the driving factors is fully characterized, providing multi-temporal data support for subsequent exploration of the nonlinear relationship between driving factors and land use conversion probability. This breaks through the limitations of existing static mapping technology and lays the foundation for characterizing the time delay process of socio-economic factors.
[0036] S2: Due to the complex high-dimensional nonlinear relationship between driving factors and land use conversion probabilities, traditional linear methods cannot accurately mine this nonlinear relationship. Therefore, this embodiment introduces an artificial neural network, which has powerful nonlinear mapping capabilities and can fit the high-dimensional nonlinear relationship between driving factors and land use conversion probabilities. Specifically: This embodiment uses a multilayer perceptron (MLP) or a deep neural network (DNN) as the artificial neural network. The number of nodes in the input layer of the artificial neural network is equal to the total length of the lag sequences of all driving factors. The input data is the basic input sequence (i.e., the time lag sequence set of each driving factor), where the time lag sequence set is a multi-temporal feature vector. The output layer is the static suitability probability of the mutual conversion of various land uses. It mainly considers the nonlinear driving relationship of location factors, natural factors, and human activity factors. For example, this embodiment uses MLP as an example. The input includes 10 driving factors, and the combined lag period of the 10 driving factors is 26. Therefore, the number of nodes in the input layer is 26, and the number of hidden layers is set to 2. The number of nodes in the first layer is 50, and the number of nodes in the second layer is 30. The ReLU activation function is used. The number of nodes in the output layer is 6, corresponding to the set land type. The activation function used is the Sigmoid activation function. The output is the static suitability probability of the mutual conversion between various land use types. ; In the formula: This represents the suitability probability of land type m (such as cultivated land, urban construction land, forest land, etc.) in pixel p; This represents the nonlinear driving relationships discovered by the multilayer perceptron after model training; This embodiment represents various locational factors, primarily including distance from roads and distance from town centers. It represents the characteristics of various natural elements, for example, including the lag sequence of topographic factors (elevation, slope, aspect), meteorological factors (annual average temperature, annual precipitation), and soil factors (soil texture grade, soil fertility grade); It represents the characteristics of various human activity elements, including, for example, the lagged series of population factors (resident population density) and GDP factors (GDP per capita); As a further implementation of this embodiment, the Garson algorithm is adopted. By calculating the absolute weight product from the input layer to the hidden layer and from the hidden layer to the output layer in the artificial neural network, the relative contribution percentage of each driving factor to the probability of conversion of a specific land use type is analyzed. Taking the MLP model as an example, the specific process is as follows: For a trained MLP model, the relative importance (i.e., relative contribution) of the i-th driving factor to the k-th land use type is calculated as follows: ; In the formula: represents the relative contribution of the i-th driving factor to the probability of land use type conversion in the k-th region, with a value range of [0,1]. The larger the value, the stronger the influence of the driving factor on land use type conversion; N represents the total number of input layer nodes, i.e., the total number of input natural and socio-economic driving factors; L represents the number of hidden layer nodes. This represents the connection weight between the i-th driving factor node in the input layer and the j-th node in the hidden layer. This represents the connection weight between the m-th driving factor node in the input layer and the j-th node in the hidden layer. This represents the connection weight between the j-th node in the hidden layer and the k-th node in the output layer. In this embodiment, the selected connection weights of the MLP model are extracted, the relative contribution of each driving factor to the six land use types is calculated, and then the relative contributions of all driving factors corresponding to each land use type are sorted in descending order. A threshold (e.g., 0.05) is set, and driving factors with a relative contribution greater than the preset threshold are selected as core driving factors to adjust the structure of the dynamic Bayesian network (DBN). By eliminating driving factors with low contribution, network redundancy is avoided and the model's computational efficiency is improved.
[0037] In this embodiment, the nonlinear relationship between high-dimensional driving factors and land use conversion rate is explored through MLP; the Garson algorithm is used to quantitatively analyze the relative contribution of each driving factor, screen the core driving factors, simplify the subsequent DBN network structure, and clarify the influence intensity of different driving factors, providing a basis for model optimization.
[0038] S3: Construct a dynamic Bayesian network, using time-slicing to divide the time series into multiple time slices (e.g., a 10-year time series, with each year divided into one time slice). In this embodiment, the dynamic Bayesian network structure defines two types of nodes, including: Land status node: Each time slice corresponds to one land status node, which represents the land use status of the study area at time t, that is, the land use type at time t, such as 1 representing cultivated land, 2 representing urban construction land, 3 representing forest land, 4 representing water area, 5 representing grassland, and 6 representing unused land; Observation nodes for driving factors: Each time slice t corresponds to observation nodes for multiple driving factors, such as resident population density (4-year lag), GDP per capita (4-year lag), distance to the nearest road network (5-year lag), average annual temperature (3-year lag), altitude (2-year lag), and soil fertility level (2-year lag). Observation nodes can be adjusted by determining core driving factors, including deleting driving factors with low contribution and supplementing core driving factors.
[0039] In the DBN topology, the observation nodes from the driving factors are explicitly constructed. To land state node The directed edges, where k represents the significant lag period of each driving factor to Directed edges, for example The resident population density represents the impact of the resident population density four years ago on land conversion at the current time t; for example... (Distance to the nearest road network) The directed edges represent the impact of the road network distribution 5 years ago on the land conversion at the current time t), while establishing the... arrive The directed edges represent the influence of the land state at the previous moment on the land state at the current moment (i.e., the inertia effect of land use).
[0040] In this embodiment, each time slice t contains 1 Nodes and multiple Nodes (e.g., 10), adjacent time slices are connected via... Directed edge connections are established between driving factor nodes and land state nodes via... The directed edges are connected to form a complete temporal causal network, which is used to capture the time lag effect of driving factors and the inertia effect of land use.
[0041] In this embodiment, DBN possesses powerful time-series probabilistic reasoning capabilities. Through high-order Markov chain structures and directed arcs across time slices, it clearly defines the causal relationship between driving factors and land state, rather than the static correlation of traditional models. Furthermore, the explicitly constructed directed edges can directly quantify the impact of driving factors on current land use change before significant lag periods, solving the core problem of traditional models neglecting time-lag effects. Simultaneously, it incorporates… Directed edge connections can capture the inertial effect of land use.
[0042] In this embodiment, by collecting data from all nodes in the DBN topology, namely the land state node data and driving factor observation node data for each time slice, each raster cell corresponds to a complete set of node data sequences; For nodes in the DBN topology, this embodiment uses the variance inflation factor (VIF) method from Python's statsmodels library to perform collinearity tests on the data of multiple driver factor observation nodes, and sets an inflation factor threshold (e.g., 10) to remove nodes with a variance inflation factor greater than the inflation factor threshold; by removing nodes with high collinearity, the independence between driver factors is ensured.
[0043] Then, using the optimal binning and discrete classification standard method, the data of the driving factor observation nodes are discretized, and the continuous values of each driving factor are divided into multiple levels (e.g., 1 to 5 levels, where level 1 represents the smallest value and level 5 represents the largest value); the data of the land status nodes are already discretely encoded, so no additional discretization is needed here.
[0044] The discretized node data is organized into training and testing sets according to time slices, with the training set accounting for 80% and the testing set accounting for 20%, for parameter learning and validation of the DBN model.
[0045] This embodiment uses land use suitability probability. As a land state node The prior probabilities are used to ensure that they reflect the nonlinear contributions of the driving factors, providing a reasonable basis for subsequent probabilistic inference. Specifically, during parameter learning, the Expectation-Maximization (EM) algorithm is used to handle missing data that may exist in the node data (such as missing meteorological data for some years). Through iterative optimization, the conditional probability of each node under different combinations of parent node states is calculated. The specific steps include: Initialize the conditional probability table (CPT) of DBN, and then continuously update the conditional probability table through the E-step (calculate expectation) and M-step (maximize expectation) of the EM algorithm until the algorithm converges, thus obtaining the final conditional probability table.
[0046] This embodiment leverages the powerful missing data processing capabilities of the EM algorithm to effectively address the problem of missing data in practical applications, ensuring the accuracy of parameter learning. Simultaneously, it efficiently estimates the conditional probability table of the DBN, providing core parameters for subsequent dynamic inference.
[0047] Then, based on the trained DBN model and combined with the conditional probability table, the dynamic transition probability considering the time lag effect is calculated, that is, the dynamic probability of pixel p being converted to land type m at the current time t is calculated using the following formula: ; In the formula: This represents the dynamic transfer probability of considering the time lag effect, that is, the dynamic probability of pixel p changing from the initial land type to the target land type m at time t, with a value range of [0,1]. Represents the land state at a given previous moment. and driving factor lag sequence Under the given conditions, the current land status The conditional probabilities are derived from the conditional probability table of DBN; express The joint probability is learned from the DBN parameters; Indicates the state of the land at the previous moment. The marginal probability, that is, the probability that pixel p corresponds to the land type at time t-1; Represents the lag sequence of driving factors The marginal probability is the probability of the observed value of the driving factor at time tk.
[0048] In this embodiment, each raster cell's Data and The data is fed into the DBN model, and combined with the conditional probability table, the dynamic transfer probability of each land use type corresponding to each pixel p is calculated using the above formula. The accuracy of the DBN model is verified using a test set. Through the topological structure components, parameter learning and probabilistic inference of DBN, the time-lag effect of driving factors is captured, which fits the actual dynamic transfer probability, solves the limitation of the instant response assumption of traditional models, and improves the model's ability to infer long-term land use changes.
[0049] S4: The dynamic transition probabilities obtained from DBN derivation are implemented onto a spatial grid. Combined with the spatial evolution capability of adaptive cellular automata (CA), the spatiotemporal patterns of land use under different scenarios are simulated. The specific steps are as follows: The dynamic transition probability obtained from DBN The comprehensive conversion probability of pixel p changing from the initial land type to the target land type m at time t is obtained by coupling the calculation with the neighborhood effect of adaptive CA, the adaptive inertia coefficient, and the land type conversion cost. The specific expression is as follows: ; In the formula: It represents the comprehensive conversion probability of pixel p from the initial land type c to the target land type m at time t, and the value range is [0,1]. It is the core basis for raster state change in CA simulation. This represents the neighborhood influence coefficient of land type m at pixel p at time t, with a value range of [0,1]. It characterizes the degree of influence of the neighboring raster on the conversion of the current pixel p to land type m. For example, in this embodiment, a 3×3 mole neighborhood is used. The larger the number of rasteres with land type m in the neighborhood, the better. The larger; The adaptive inertia coefficient of land type m at time t is represented, with a value range of [0.5-1.5]. It characterizes the stability of land type m. The larger the inertia coefficient, the more difficult it is for land m to be converted. In this embodiment, it is set based on the natural attributes of land use type. For example, the inertia coefficient of water area is the largest and can be 1.5. On the other hand, if unused land is easy to develop, the inertia coefficient is the smallest and can be 0.5. This represents the conversion cost of converting land type c to land type m, with a value range of [0.1-1.0]. The lower the conversion cost, the easier the conversion between the two land types. In this embodiment, the conversion cost is set based on the difficulty of land use conversion, such as 0.3 for arable land to urban construction land, 0.4 for forest land to arable land, and 0.2 for unused land to arable land.
[0050] This embodiment leverages the powerful spatial evolution capabilities of the CA model to simulate spatial competition and clustering effects in land use. By coupling the dynamic transition probability of DBN with the neighborhood effect, adaptive inertia coefficient, and conversion cost of CA, it can combine the time-delay effect in the time dimension with the clustering effect in the spatial dimension, thereby achieving spatiotemporal collaborative simulation of land use change. The setting of the adaptive inertia coefficient can reflect the stability differences of different land types.
[0051] As a further implementation of this embodiment, based on the actual land use planning needs of the study area and combined with dual carbon targets and ecological protection requirements, three scenario parameters are set as follows: Baseline development scenario: Following the current land use development trend in the study area, without adding additional constraints on ecological protection or urban expansion, the inertia coefficient and conversion cost of each land type are kept at default settings to simulate the natural evolution trend of future land use.
[0052] Urban expansion scenario: Focuses on the development of urban construction land, reduces the conversion cost of urban construction land (e.g., 0.2), increases the inertia coefficient of urban construction land (e.g., 1.3); at the same time, reduces the conversion constraint between cultivated land and urban construction land, simulating land use changes under the background of rapid urbanization.
[0053] Ecological protection scenario: Focusing on ecological environmental protection, increasing the inertia coefficient of forest land and water area (Inertia=1.4 for forest land and Inertia=1.6 for water area), increasing the conversion cost of forest land and water area to other land types (Cost=0.95), and strictly restricting the conversion of arable land to urban construction land (Cost=0.6), simulating land use change under the background of ecological priority.
[0054] For the three scenarios, an adaptive inertia mechanism is adopted to dynamically adjust the inertia coefficient of each land type according to the total demand of each land type each year, so as to ensure that the total conversion of each land type meets the target set by the scenario (e.g., in the ecological protection scenario, the area of forest land and water area shall not be less than 40% of the total area of the study area).
[0055] In this embodiment, the grid state change is achieved through a roulette wheel algorithm, and the specific steps are as follows: 1. For each raster cell p, convert it into a combined conversion probability for different target land types m. As the weights for roulette, the sum of the weights is normalized to 1; 2. Each target land type m corresponds to a region in the roulette wheel. The size of the region is related to the overall conversion probability of that land type. Proportional; 3. Generate a random number between 0 and 1 using Python, and determine the final land type of the raster cell p based on the area of the roulette wheel where the random number falls; Repeat steps 1 to 3 above for all grids in the study area to complete the land use state simulation at the current time t, and then proceed to the next time step.
[0056] In this embodiment, by leveraging the spatial evolution capability and multi-scenario settings of CA, the dynamic transition probability of DBN is implemented in a spatial grid, thereby realizing the simulation of the spatiotemporal pattern of land use change. Furthermore, the three scenarios can cover different development needs, providing diversified decision-making methods for regional land planning and the delineation of ecological protection red lines.
[0057] This embodiment integrates the time-delay effect of driving factors into the model building process for the first time, combining the nonlinear mapping capability of ANN, the time-series probabilistic reasoning capability of DBN, and the spatial evolution capability of CA. It breaks through the key limitations of traditional models that ignore the time-delay effect of driving factors, lack cumulative dynamics representation, and have weak uncertainty reasoning capability, and realizes the leap from static mapping to dynamic causality.
[0058] The above are merely preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A land use change scenario simulation method that takes into account the time lag of driving factors, characterized in that, Includes the following steps: S1: Standardize and perform time-delay feature engineering on the collected multi-source heterogeneous data, extract the driving factors from the multi-source heterogeneous data, and construct the time-delay sequence set of the driving factors as the basic input sequence to characterize the time accumulation effect. S2: Using an artificial neural network with the aforementioned basic input sequence as input, the high-dimensional nonlinear relationship between driving factors and land use conversion probability is mined to obtain the land use suitability probability and the relative contribution of each driving factor. S3: Using a dynamic Bayesian network, the land use suitability probability is used as the prior probability of land state. Combined with the significant lag period of each driving factor and the extracted driving factors, a causal relationship between driving factors and land state is constructed in the dynamic Bayesian topology. The time cumulative effect and causal lag logic of driving factors are quantitatively captured to obtain the dynamic transfer probability of considering the time lag effect. S4: Using the dynamic transition probability as the core input, the comprehensive transition probability is obtained by coupling calculation with the neighborhood effect of the adaptive cellular automaton, the adaptive inertia coefficient set by the land use type characteristics, and the land type conversion cost. Then, through multi-scenario iterative allocation, the spatiotemporal pattern simulation of land use change is completed.
2. The land use change scenario simulation method considering the time lag of driving factors according to claim 1, characterized in that, The multi-source heterogeneous data includes land use remote sensing data, natural geographic data, and socio-economic statistical data.
3. The land use change scenario simulation method considering the time lag of driving factors according to claim 2, characterized in that, The standardization process and time-delay feature engineering include the following steps: By using a GIS platform, land use remote sensing data, natural geographic data, and socio-economic statistical data are unified to the same coordinate system and preset raster spatial resolution to obtain standardized multi-source heterogeneous data. Then, natural driving factors and socio-economic driving factors are extracted from standardized multi-source heterogeneous data, with natural driving factors extracted from standardized natural geographic data and socio-economic driving factors extracted from standardized socio-economic statistical data. For each extracted driving factor, the mutual information between the sequence of each driving factor at different lag times and the current land state is calculated to determine the significant lag period of the impact of each driving factor on land use. The lag period corresponding to the maximum value of the mutual information is the significant lag period of the corresponding driving factor. Based on the determined significant lag period, a time lag sequence set of the corresponding driving factor is constructed, and the time lag sequence set is used as the basic input sequence.
4. The land use change scenario simulation method considering the time lag of driving factors according to claim 3, characterized in that, The natural geographic data includes at least one of topographic data, meteorological data, and soil data, and the corresponding natural driving factors include at least one of topographic factors, meteorological factors, and soil factors.
5. A land use change scenario simulation method considering the time lag of driving factors according to claim 3, characterized in that, The socioeconomic statistics include at least one of population data, GDP data, and road network data, and the corresponding socioeconomic driving factors include at least one of population factors, GDP factors, and road network factors.
6. The land use change scenario simulation method considering the time lag of driving factors according to claim 1, characterized in that, Step S2 includes the following steps: The basic input sequence is used as the input of the artificial neural network, and the artificial neural network processes the data to output the static suitability probability of the mutual conversion of various land uses. The Garson algorithm is used to analyze the relative contribution percentage of each driving factor to the probability of land use type conversion by calculating the absolute weight product from the input layer to the hidden layer and from the hidden layer to the output layer in the artificial neural network, thus obtaining the relative contribution of each driving factor.
7. A land use change scenario simulation method considering the time lag of driving factors according to claim 6, characterized in that, The relative contribution is a standardized value between 0 and 1. The relative contribution is sorted in descending order, and driving factors with relative contribution greater than the threshold are selected as core driving factors based on a preset threshold to adjust the dynamic Bayesian network structure.
8. A land use change scenario simulation method considering the time lag of driving factors according to claim 1, characterized in that, Step S3 includes the following steps: Dynamic Bayesian Network Construction: Define two types of nodes: land state nodes at each time step and driving factor observation nodes at each time step. In the dynamic Bayesian network, explicitly establish directed edges from the observation node corresponding to the significant lag period of each driving factor to the land state node at the current time step, representing the impact of each driving factor on the current land transformation before the significant lag period. Data from all nodes in the dynamic Bayesian network is collected, collinearity is tested and nodes with collinearity exceeding the threshold are removed. The remaining node data is discretized using the optimal binning and discrete classification standard methods and then organized into training and test sets according to time slices. Using the land use suitability probability as the prior probability of land state nodes, the expectation-maximization algorithm is used to handle missing data, and the conditional probability of each node under different combinations of parent node states is calculated to obtain the dynamic transition probability of time lag effect.
9. A land use change scenario simulation method considering the time lag of driving factors according to claim 1, characterized in that, Step S4 includes the following steps: The dynamic transfer probability of considering the time lag effect, the neighborhood interaction coefficient corresponding to the horizontal neighborhood factor of the adaptive cellular automaton, the adaptive inertia coefficient set based on the natural attributes of land use type, and the land type conversion cost set based on the difficulty of land use type conversion are coupled and calculated to obtain the comprehensive conversion probability of each raster cell from the initial land type to the target land type at each time. Based on actual land use planning needs, at least three scenario parameters are set, and an adaptive inertial mechanism is adopted to adjust the total demand balance of each land type. Using the roulette wheel algorithm, based on the obtained comprehensive conversion probability, the land state of each grid cell is changed, thereby simulating the spatiotemporal pattern of land use under different scenarios.
10. A land use change scenario simulation method considering the time lag of driving factors according to claim 9, characterized in that, The execution steps of the roulette wheel betting algorithm are as follows: The overall conversion probability of each raster cell to different target land types is used as the weight of the roulette wheel. Each target land type corresponds to a region in the roulette wheel, and the size of the region is proportional to the overall conversion probability of the corresponding land type. The final land type of the corresponding raster cell is determined by randomly selecting a region in the roulette wheel.