Land utilization space-time prediction method and system

By constructing a multi-source data set and using gradient enhancement decision tree and cellular automata model, combined with dynamic sampling strategies, the problem of data-driven single and insufficient prediction of key areas of land use prediction methods in the existing technology is solved, and more efficient land use spatio-temporal prediction is achieved.

CN120069218APending Publication Date: 2025-05-30GUANGDONG UNIV OF TECH

Patent Information

Application Number
CN202510218043.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, the land use prediction methods have problems such as single data-driven, insufficient prediction of key areas, and lack of constraints and guidance on land use planning goals, resulting in insufficient prediction accuracy.

Method used

By collecting multi-source data, classifying and integrating, building multi-source data sets, and using gradient enhancement decision tree model and cellular automata model, combined with dynamic sampling strategies, an intelligent goal-oriented spatial and temporal prediction system for land use is built.

Benefits of technology

The model's ability to portray land use change laws is improved, the prediction accuracy of high-development potential areas is significantly improved, the waste of computing resources for low-priority areas is reduced, and the allocation of model resources is more efficient. The prediction results are highly consistent with future planning needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069218A_ABST
    Figure CN120069218A_ABST
Patent Text Reader

Abstract

The invention discloses a land utilization space-time prediction method and system, and the method comprises the following steps: collecting multi-source data, and carrying out the classification and integration of the multi-source data, and obtaining a basic data set; constructing a multi-source data set by using the basic data set; setting a time interval of each round of evolution, and obtaining a development probability and a transition probability matrix of each land utilization type and a distribution target of each round of land utilization type by using the multi-source data set in combination with a preset method; the method comprises the following steps: constructing an intelligent target-oriented cellular automaton model, introducing a dynamic sampling strategy, carrying out land utilization continuous space-time prediction by using the cellular automaton model, and outputting a prediction result. The precision and the flexibility of land utilization space-time prediction and the matching degree with an actual planning target are remarkably improved, and meanwhile, the resource allocation and the calculation efficiency of a key area are optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of land use prediction, and more specifically, relates to a method and system for spatio-temporal prediction of land use. Background Art

[0002] In the context of social and economic development and urbanization, the conflict between land supply and demand is intensifying day by day, leading to challenges such as urban sprawl, traffic congestion, and negative environmental impacts. Therefore, the accuracy of land use prediction is crucial for policy-making, urban management, environmental protection, water system planning, and climate change adaptation strategies, which can help decision-makers make wise decisions that balance economic growth with sustainable land use and water system management. Currently, traditional land use prediction models have limitations such as input restrictions on land use data, the singularity of data mining algorithms, insufficient flexibility and fineness in parameter settings, and the inability to simulate continuous time and space.

[0003] As a commonly used method for predicting spatio-temporal changes in land use, cellular automata have the advantage of being able to comprehensively consider the impacts of various land use types, simulate the spatio-temporal evolution process of complex systems, and have high accuracy and reliability. This model has good openness and strong data compatibility, and is suitable for solving problems of simulating micro-spatial changes in complex systems. However, traditional models can only simulate discrete time and space, and it is difficult to handle continuous time and space changes. Moreover, challenges remain in effectively using land use data and driving data over multiple years, fully mining the macro factors behind land use changes, and improving the model simulation effect, which require further research and improvement.

[0004] The invention patent with the publication number CN117114176A in the prior art proposes a method and system for predicting land use changes based on data analysis and machine learning, including the following steps: collecting data to construct a data set; determining the classification criteria for land use types and performing data preprocessing; calculating the land use transfer matrix; constructing a land use change prediction model based on the CA-Markov model through machine learning algorithms according to land use transfer changes and land use change driving factors; performing model training, and predicting future land use changes through the trained model. This solution pays equal attention to all regions, ignores the priority of high-development-potential regions, and lacks constraint guidance for land use planning goals, resulting in insufficient prediction accuracy. Summary of the Invention

[0005] In order to overcome the problems in the prior art such as single data driving in the prediction method, insufficient prediction of key regions, and lack of constraint guidance for land use planning goals, the present invention provides a method and system for spatio-temporal prediction of land use.

[0006] The primary objective of the present invention is to solve the above technical problems, and the technical solution of the present invention is as follows:

[0007] In a first aspect of the present invention, a method for spatio-temporal prediction of land use is provided, including the following steps:

[0008] Collect multi-source data and classify and integrate it to obtain a basic data set;

[0009] Use the basic data set to construct a multi-source data set;

[0010] Set the time interval for each round of evolution, and use the multi-source data set in combination with a preset method to obtain the development probability, transfer probability matrix of each land use type, and the distribution target of each round of land use types;

[0011] Use the multi-source data set, the development probability of each land use type, the transfer probability matrix, and the distribution target of each round of land use types to construct a cellular automata model with intelligent goal orientation, introduce a dynamic sampling strategy, and use the cellular automata model to perform continuous spatio-temporal prediction of land use and output the prediction results.

[0012] Further, the multi-source data includes land use data; elevation data and slope data under the terrain data category; temperature data, potential evapotranspiration data, normalized difference vegetation index data, and rainfall data under the climate data category; population distribution data and night light data under the population data category; urban arterial road data, urban secondary road data, and railway data under the location condition data category; general public budget revenue data, number of students in ordinary middle schools, regional gross domestic product data, and number of beds in various health institutions under the urban development data category.

[0013] Further, the classification and integration of multi-source data includes the following steps:

[0014] Use the population distribution data and night light data in the multi-source data to analyze the population quantity, spatial distribution, and activity characteristics, and generate more accurate population information data for the basic data set;

[0015] In the processing of urban development data, particularly consider the number of students in ordinary middle schools and the number of beds in various health institutions, generate corresponding urban development index data, and add it to the basic data set;

[0016] Unify the time dimension and ensure that the basic data set covers a time span of n years.

[0017] Further, the method for constructing a multi-source data set using the basic data set includes the following steps:

[0018] Perform extended analysis on two adjacent periods of land use data in the basic dataset, extract the land use change areas to generate a land expansion map, and on this basis, extract the driving factor data at the corresponding spatial positions and supplement the driving factor data to the basic dataset;

[0019] Flatten, splice, and stack different data types in the basic dataset to generate a four-dimensional multi-source data structure, where the first dimension is the time dimension, the second dimension is the data dimension, and the third and fourth dimensions are the spatial dimensions, recording the coordinates of the pixels;

[0020] Clean the multi-source data, including identifying and processing missing values, outliers, and duplicate data: use bilinear interpolation to calculate the missing data based on the weighted average of the four nearest input pixels; replace the outlier data by removing or taking values from adjacent pixels; remove the duplicate data;

[0021] Normalize the data, and the calculation formula is as follows:

[0022]

[0023] where, x * is the data after normalization, x is the original data to be processed, x min is the minimum value in the original data, x max is the maximum value in the original data, and the value of the data after normalization is between 0 and 1;

[0024] Compress the normalized multi-source dataset and output the multi-source dataset.

[0025] Furthermore, the method for obtaining the development probability of each land use type using the multi-source dataset includes the following steps:

[0026] Flatten the time dimension and spatial dimension of the multi-source dataset, merge them into the data dimension, and form an input sample dataset;

[0027] Randomly shuffle the sample dataset and divide it into a training set and a validation set according to a preset ratio;

[0028] Use the gradient boosting decision tree model to iteratively train the sample training set dataset, gradually construct the decision tree for each round, and the expression of the gradient boosting decision tree model is as follows:

[0029]

[0030] where, x is the driving factor, F 0 (x) is the initial decision tree model, ω t is the parameter of the gradient boosting decision tree model, T is the number of decision trees; f t(x) is the residual prediction value of decision tree t for driving factor x, and F(x) is the output gradient boosting decision tree model.

[0031] After each round of decision tree training, the loss function value is calculated using the validation set. The expression of the loss function is as follows:

[0032]

[0033] Among them, L(y, f(x, ω)) is the loss function, and f(x, ω) is the predicted value of the gradient boosting decision tree model when the driving factor is x and the parameter is ω; (x i , y i ) is the observed data of the i-th sample in the m samples, x i is the driving factor of the sample, and y i is the actual observed value of the sample;

[0034] According to the loss change trend of the validation set, the number of decision trees and model parameters are selected to make the model perform best on the validation set;

[0035] Using the trained gradient boosting decision tree model, the relationship between the driving factors and land use types is mined to obtain the contribution degree of the driving factors for each land use type. The driving factors include land use, terrain, climate, population, location conditions, and urban development;

[0036] Combined with future driving factor data, the gradient boosting decision tree model is used to predict the growth probability of each future land use type, and the development probability of each land use type is output.

[0037] Furthermore, the method for obtaining the transition probability matrix using the multi-source dataset includes the following steps:

[0038] Using the adjacent two-phase land use data in the multi-source dataset, calculate the land use type change of each pixel in different periods to generate a land use transfer matrix;

[0039] Sum and average the land use transfer matrices for multiple time periods to form a basic transfer probability matrix;

[0040] According to the policy requirements or planning objectives, combined with the protection requirements or conversion restrictions of specific land use types, adjust the basic transfer probability matrix and output the final transfer probability matrix.

[0041] Furthermore, the method for obtaining the distribution target of each round of land use type using the multi-source dataset includes the following steps:

[0042] Using the multi-source dataset, statistically analyze the land use data for multiple time periods, and calculate the historical distribution quantity and change trend of each land use type;

[0043] Construct a linear model of multi-objective planning based on historical trend data, combined with the finiteness of land resources, the restrictive conditions of policies and regulations, and the requirements of environmental protection.

[0044] Set constraints for the linear model of the multi-objective planning, and the constraints include: the finiteness of land resources, the restrictions of policies and regulations, and the requirements of environmental protection.

[0045] Use the linear model of the multi-objective planning to calculate the distribution objectives of land use types for each round.

[0046] Furthermore, a method for constructing a cellular automata model with intelligent goal orientation and conducting continuous spatio-temporal prediction of land use includes the following steps:

[0047] Use the multi-source data set of the base period, combined with the transition probability matrix and the land use distribution objectives of this round, to initialize the cellular automata model, and set the initial land use type distribution and its conversion probability of the model sample pixels; initialize the dynamic sampling strategy weight and the spatial position screening rule. The dynamic sampling strategy weight is used to control the extraction priority of sample pixels, and the dynamic sampling strategy spatial position screening rule is used to guide the geographical range of dynamic sampling. The sampling of the initial sample pixels needs to keep the quantity proportion of each land use type consistent with the actual data.

[0048] Start the evolution of this round. Select the geographical range through the dynamic sampling strategy spatial position screening rule, and preferentially extract sample pixels in the high dynamic sampling weight area from the geographical range. Calculate the neighborhood effect of each sample pixel using the transition probability matrix, and the expression is as follows:

[0049]

[0050] where, Ω i,k is the neighborhood effect of the k-th land use type in the neighborhood of sample pixel i, and ∑ N×N con(c i =k) is the number of pixels belonging to the k-th land use type among the pixels with a radius of N in the neighborhood of sample pixel i. N is the neighborhood parameter, and TPM k is the transition probability obtained from the transition probability matrix for the conversion of the land use type of sample pixel i to the k-th land use type.

[0051] Adjust the sample pixels with a neighborhood effect of 0, introduce a multi-type random patch seeding mechanism based on the Monte Carlo method, gradually approximate its true transition probability through random sampling, and introduce a threshold that gradually decreases with the number of evolutions to suppress the seeding of patches.

[0052] After sowing, the neighborhood effect, the development probability of each land use type, and the adaptive coefficient are combined to calculate the overall conversion probability of the sample pixel, and the expression is as follows:

[0053]

[0054] Among them, TP i,k is the overall conversion probability that the sample pixel i is converted into the k-th land use type, and Ω i,k is the neighborhood effect of the k-th land use type within the neighborhood of the sample pixel i, and F i,k is the development probability that the sample pixel i develops into the k-th land use type, is the adaptive coefficient of the k-th land use type within the t-th round of evolution, and the expression is as follows:

[0055]

[0056] Among them, is the difference between the distribution target of the land use type within the (t - 1)-th round of evolution and the quantity of the land use type k at the beginning of the (t - 1)-th round, is the difference between the distribution target of the land use type within the (t - 2)-th round of evolution and the quantity of the land use type k at the beginning of the (t - 2)-th round, is the adaptive coefficient of the k-th land use type within the (t - 1)-th round of evolution;

[0057] The overall conversion probability is input into the roulette algorithm to output the candidate land use types of the sample pixels, and combined with the transition probability matrix to determine whether each sample pixel undergoes a land use type conversion;

[0058] Update the land use distribution of the current round as the input of the land use type distribution for the next round, calculate the sample demand difference of the next round and update the adaptive coefficient

[0059] According to the adaptive parameter Adjust the weight of the dynamic sampling strategy, and according to the development probability and the proportion of the sample pixels in the area that has not changed during this round of evolution, adjust the spatial position screening rule of the dynamic sampling strategy;

[0060] Repeat the above evolution steps until the difference between the distribution quantity of the land use type and the distribution target of the land use type in this round converges to the set threshold, and output the land use distribution of this round;

[0061] Reset the adaptive coefficient and the weight value of the dynamic sampling strategy, and perform the evolution of the next round until the evolution of all rounds is completed, and output the land use distribution result of continuous spatio-temporal prediction.

[0062] In a second aspect of the present invention, a land use spatio-temporal prediction system is provided, which includes a memory and a processor. The memory includes a land use spatio-temporal prediction method program, and when the land use spatio-temporal prediction method program is executed by the processor, the steps of a land use spatio-temporal prediction method are implemented.

[0063] In a third aspect of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium includes a land use spatio-temporal prediction method program, and when the land use spatio-temporal prediction method program is executed by a processor, the steps of a land use spatio-temporal prediction method are implemented.

[0064] Compared with the prior art, the beneficial effects of the technical solution of the present invention are as follows:

[0065] By classifying and integrating multi-source data, the present invention generates a multi-dimensional data structure, avoiding the problem of insufficient prediction caused by single data driving in traditional methods, and improving the model's ability to depict the law of land use change; by introducing a dynamic sampling strategy, the prediction accuracy of the model for high-development-potential regions is significantly improved, while reducing the waste of computing resources for low-priority regions, making the model resource allocation more efficient; during the evolution process, the model gradually approaches the planning goal, making the prediction results not only conform to the historical change law but also highly coincide with the future land use planning requirements, providing a scientific basis for urban planning, ecological protection, and policy making. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] In order to make the objectives and technical solutions of the present invention clearer, the following drawings are provided by the present invention and described:

[0067] Figure 1 It is a flowchart of the method provided by an embodiment of the present invention;

[0068] Figure 2 It is a schematic diagram comparing the simulation results of the present invention with the actual land use types provided by an embodiment of the present invention;

[0069] Figure 3 It is a schematic diagram of the precision and recall rate of the model of the present invention provided by an embodiment of the present invention;

[0070] Figure 4 It is a comparison diagram of the continuous spatio-temporal simulation results of the model of the present invention and the simulation results of the traditional discrete model provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0071] In order to more clearly understand the above objectives, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other.

[0072] In the following description, many specific details are set forth to provide a thorough understanding of the present invention. However, the present invention may be implemented in other ways different from those described herein. Therefore, the scope of protection of the present invention is not limited by the specific embodiments disclosed below.

[0073] Embodiment 1:

[0074] The present invention provides a method for spatio-temporal prediction of land use. As Figure 1 shown in the flowchart of a method for spatio-temporal prediction of land use, the specific steps are as follows:

[0075] S1: Collect multi-source data of the research area, classify and integrate it to obtain a basic data set.

[0076] More specifically, the multi-source data includes land use data; elevation data and slope data under the terrain data category; temperature data, potential evapotranspiration data, normalized difference vegetation index data, and rainfall data under the climate data category; population distribution data and night light data under the population data category; urban arterial road data, urban secondary road data, and railway data under the location condition data category; general public budget revenue data, number of students in ordinary middle schools, regional gross domestic product data, and number of beds in various health institutions data under the urban development data category.

[0077] The multi-source data of the research area is obtained from multiple authoritative data sources. For example, land use data is obtained from the Zenodo platform, terrain data is obtained from NASA, climate data is obtained from the National Tibetan Plateau Data Center of Sciences, population data is obtained from Worldpop, location condition data is obtained from Open Street Map, and urban development data is obtained from the statistical yearbook of the research area.

[0078] The classification and integration of the multi-source data include the following steps:

[0079] Utilize the population distribution data and night light data in the multi-source data to analyze the population quantity, spatial distribution, and activity characteristics, and generate more accurate population information data for the basic data set.

[0080] In the processing of urban development data, specifically consider the number of students in ordinary middle schools and the number of beds in various health institutions to reflect the investment in educational resources, educational popularization degree, and medical service capacity of the region, generate corresponding urban development index data, and add it to the basic data set.

[0081] Ensure that the basic data set covers a time span of 10 - 20 years, unify the time dimension, and maintain the multi-source nature and analysis depth of the data.

[0082] S2: Use the basic data set to construct a multi-source data set.

[0083] More specifically, a method for constructing a multi-source dataset using a basic dataset includes the following steps:

[0084] Perform extended analysis on two adjacent periods of land use data in the basic dataset, extract the land use change areas to generate a land expansion map, and on this basis, extract the driving factor data corresponding to their spatial positions, and supplement the driving factor data to the basic dataset.

[0085] Flatten, splice, and stack different data types in the basic dataset to generate a four-dimensional multi-source data structure Numpy array, where the first dimension is the time dimension, the second dimension is the data dimension, and the third and fourth dimensions are the spatial dimensions, recording the coordinates of the pixels.

[0086] Perform data cleaning on the multi-source data, including identifying and processing missing values, outliers, and duplicate data: use bilinear interpolation to calculate the missing data based on the weighted average of the four nearest input pixels; replace the outlier data by removing or taking values from adjacent pixels; remove the duplicate data.

[0087] Normalize the data, and the calculation formula is as follows:

[0088]

[0089] Where, x * is the data after normalization, x is the original data to be processed, x min is the minimum value in the original data, x max is the maximum value in the original data, and the value of the data after normalization is between 0 and 1;

[0090] Compress the normalized multi-source dataset for storage and subsequent calls, and output the multi-source dataset.

[0091] S3: Set the time interval for each round of evolution, and use the multi-source dataset to obtain the development probability, transfer probability matrix, and distribution target of each land use type in each round in combination with a preset method.

[0092] More specifically, based on the multi-source dataset, use the Gradient Boosting Decision Tree (GBDT) model to construct a rule mining model, and obtain the development probability of each land use type by iteratively optimizing the residuals.

[0093] The specific process is as follows:

[0094] Flatten the time dimension and spatial dimension of the multi-source dataset, merge them into the data dimension, and form an input sample dataset.

[0095] Randomly shuffle the sample data set and divide it into a training set and a validation set according to a preset ratio of 9:1. The learning rate, number of decision trees, and number of nodes for generating decision trees in GBDT are 0.001, 100, and 15 respectively.

[0096] Use the Gradient Boosting Decision Tree (GBDT) model to iteratively train the sample training set data, and gradually construct the decision tree for each round. The expression of the Gradient Boosting Decision Tree (GBDT) model is as follows:

[0097]

[0098] where x is the driving factor, F 0 (x) is the initial decision tree model, ω t is the Gradient Boosting Decision Tree (GBDT) model parameter, T is the number of decision trees, and the default value is 100. The more decision trees, the better the model can fit the training data and improve the prediction accuracy to a certain extent; f t (x) is the predicted value of the residual (the difference between the true value and the predicted value) of the decision tree t for the driving factor x, and F(x) is the output Gradient Boosting Decision Tree (GBDT) model.

[0099] After each round of decision tree training, calculate the loss function value using the validation set. The expression of the loss function is as follows:

[0100]

[0101] where L(y, f(x, ω)) is the loss function, and f(x, ω) is the predicted value of the corresponding Gradient Boosting Decision Tree (GBDT) model when the driving factor is x and the parameter is ω; (x i , y i ) is the observed data of the i-th sample in the sample with a quantity of m, x i is the driving factor of the sample, and y i is the actual observed value of the sample.

[0102] According to the loss change trend of the validation set, select the number of decision trees and model parameters to make the model perform best on the validation set.

[0103] Use the trained Gradient Boosting Decision Tree (GBDT) model to explore the relationship between the driving factor and the land use type, and obtain the contribution degree of the driving factor for each land use type. The driving factors include land use, terrain, climate, population, location conditions, and urban development.

[0104] Combine future driving factor data and use the Gradient Boosting Decision Tree (GBDT) model to predict the growth probability of each future land use type and output the development probability of each land use type.

[0105] GBDT is used to explore and mine the relationship between the growth of each land use type and multiple driving factors. It can not only return the contribution degree of each driving factor to the conversion of each land use type, but also predict the development probability of each land use type in the future based on the future driving factor data, providing support for subsequent continuous spatio-temporal prediction of land use.

[0106] More specifically, based on the multi-source dataset, a statistical method is used to construct a transition probability matrix to describe the conversion probability between land use types.

[0107] The specific process is as follows:

[0108] Using the land use data of two adjacent periods in the multi-source dataset, calculate the change of land use type of each pixel at different times to generate a land use transfer matrix.

[0109] Sum and average the land use transfer matrices of multiple time periods to form a basic transfer probability matrix.

[0110] According to policy requirements or planning goals, combined with the protection requirements or conversion restrictions of specific land use types, adjust the basic transfer probability matrix and output the final transfer probability matrix.

[0111] In this step, using the multi-source dataset, with the land use data of two adjacent periods as a unit, accurately match the land use type of each pixel at different time points to obtain a land use transfer matrix, which clearly shows the conversion probability between various land use types. In order to obtain a more representative and stable transfer probability matrix, further sum the land use transfer matrices of multiple time periods and calculate their average value. This step not only smooths the abnormal fluctuations that may occur in a single year, but also reflects the long-term trend and general law of land use type conversion. Because its data comes from historical data and the calculation process is automatically counted by the model, it can simplify the model parameters to a large extent.

[0112] More specifically, based on the multi-source dataset, a linear programming method is used to construct a linear model of multi-objective programming, with land use demand as the goal, and use the model to calculate the distribution target of each round of land use type.

[0113] The specific process is as follows:

[0114] Using the multi-source dataset, statistically analyze the land use data of multiple time periods, and calculate the historical distribution quantity and change trend of each land use type.

[0115] According to the historical trend data, combined with the finiteness of land resources, the constraint conditions of policies and regulations, and the requirements of environmental protection, construct a linear model of multi-objective programming.

[0116] Set constraints for the linear model of the multi-objective programming. The constraints include: the finiteness of land resources, the restrictions of policies and regulations, and the requirements of environmental protection.

[0117] Use the linear model of the multi-objective programming to calculate the distribution objectives of land use types in each round, ensuring that they meet the policy requirements and environmental protection objectives.

[0118] S4: Utilize multi-source datasets, the development probabilities of each land use type, the transition probability matrix, and the distribution objectives of land use types in each round to construct an intelligent goal-oriented cellular automata (CA) model. Introduce a dynamic sampling strategy to make the model more flexible in adapting to the goal changes in each round. Use the cellular automata (CA) model for continuous spatio-temporal prediction of land use and output the prediction results.

[0119] The specific process is as follows:

[0120] Utilize the base-period multi-source datasets, combine the transition probability matrix and the land use distribution objectives in this round to initialize the cellular automata model, and set the initial land use type distribution and its conversion probability of the model sample pixels; initialize the dynamic sampling strategy weight and the spatial location screening rule. The dynamic sampling strategy weight is used to control the extraction priority of the sample pixels, and the dynamic sampling strategy spatial location screening rule is used to guide the geographical scope of dynamic sampling. The sampling of the initial sample pixels needs to keep the quantity proportion of each land use type consistent with the actual data. In this step, the initial weight of the dynamic sampling strategy is zero. First, the model will extract the corresponding land use types from the base-period land use data, and the quantity proportion of each land use type remains consistent, and the total quantity accounts for about 5% of the total sample quantity. The spatial locations of the extracted samples are based on the development probabilities of each land use type obtained in the previous steps, where 80% of the samples with a development probability greater than 50% account for the total sample quantity, and 20% of the samples with a development probability greater than 50%.

[0121] Start the evolution in this round. Select the geographical scope through the dynamic sampling strategy spatial location screening rule, and preferentially extract the sample pixels in the high dynamic sampling weight area from the geographical scope. Use the transition probability matrix to calculate the neighborhood effect of each sample pixel, that is, calculate the states of its neighboring pixels, which is used to describe the potential of the sample pixel i to be converted into the k-th land use type. The expression is as follows:

[0122]

[0123] where, Ω i,k is the neighborhood effect of the k-th land use type within the neighborhood of the sample pixel i, ∑ N×N con(c iN_{i,k} is the number of pixels belonging to the k-th land use type among the pixels within the neighborhood of sample pixel i with a radius of N. N is the neighborhood parameter that defines the size of the neighborhood radius. Neighborhood land use changes do not occur in isolation but are affected by surrounding pixels. The neighborhood parameter defines the propagation distance of this influence. The value of the neighborhood range needs to be based on the characteristics of the study area. In this embodiment, the value is 7, TPM k TPM_{i,k} is the transition probability of the land use type corresponding to sample pixel i obtained from the transition probability matrix being converted to the k-th land use type.

[0124] In this step, using the transition probability matrix obtained in the previous step, calculate the neighborhood effect of each sample pixel's land use type conversion from the initial type to other land use types. There are significant neighborhood effects between different land use combinations, and the neighborhood effects have influences such as inertia, repulsion, and attraction. Traditional neighborhood effects only count the number of land use types surrounding each sample pixel without considering that different land types have different sensitivities and responses to neighborhood effects. The addition of the transition probability matrix can adjust the transition probabilities of each land use type to make them as close to the actual situation as possible.

[0125] Adjust the sample pixels with a neighborhood effect of 0, introduce a multi-type random patch seeding mechanism based on the Monte Carlo method, and gradually approximate its true transition probability through random sampling. Introduce a decreasing threshold to suppress patch seeding, which is determined by the number of evolutions and gradually decreases as the evolution progresses, thereby reducing the perturbation of patch seeding on the model results;

[0126] After seeding, combine the neighborhood effect, the development probability of each land use type, and the adaptive coefficient to calculate the overall conversion probability of the sample pixel, and the expression is as follows:

[0127]

[0128] where, TP i,k is the overall conversion probability of sample pixel i being converted to the k-th land use type, Ω i,k is the neighborhood effect of the k-th land use type within the neighborhood of sample pixel i, F i,k is the development probability of sample pixel i developing into the k-th land use type, is the adaptive coefficient of the k-th land use type within the t-th round of evolution, and its initial value is 1 for all, and the expression is as follows:

[0129]

[0130] where, is the difference between the distribution target of land use types within the (t - 1)-th round of evolution and the quantity of land use type k at the beginning of the (t - 1)-th round, is the difference between the distribution target of land use types within the (t - 2)-th round of evolution and the quantity of land use type k at the beginning of the (t - 2)-th round, is the adaptive coefficient of the k-th land use type within the (t - 1)-th round of evolution;

[0131] Input the overall conversion probability into the roulette algorithm, output the candidate land use types of sample pixels, and combine with the transfer probability matrix to determine whether each sample pixel undergoes land use type conversion;

[0132] Update the land use distribution of the current round as the input of the land use type distribution for the next round, and calculate the sample demand difference of the next round and update the adaptive coefficient

[0133] According to the adaptive parameter Adjust the weight of the dynamic sampling strategy, and according to the development probability and the proportion of sample pixels in the area that has not changed during this round of evolution, adjust the spatial position screening rule of the dynamic sampling strategy;

[0134] Repeat the above evolution steps until the difference between the distribution quantity of land use types and the distribution target of land use types in this round converges to the set threshold, and output the land use distribution of this round;

[0135] Reset the adaptive coefficient and the weight value of the dynamic sampling strategy, and conduct the evolution of the next round until the evolution of all rounds is completed, and output the land use distribution result of continuous spatio-temporal prediction.

[0136] Based on the actual land use data in 2000, the land use types in 2020 were simulated through experiments, and the simulation results are as Figure 2 shown. The simulation results of the present invention were compared with the actual land use types in 2020, and in-depth analysis was carried out for three specific regions. According to the comparison results, the accuracy rate of the model was calculated to be 91.01%, indicating that the simulation results are highly consistent with the actual data. In addition, from the perspective of visual quality, the prediction results of the model of the present invention are also consistent with the actual results, further verifying the effectiveness of the model.

[0137] Figure 3The precision rate and recall rate of the model's predictions for different land use types are shown. The best model has the best prediction effect on green land (the precision rate reaches 97%), and a relatively poor prediction effect on urban land (but the precision rate also reaches 83%), indicating that the model's predictions for various land use types are relatively accurate, with only a few prediction errors. The recall rate of the model's predictions for different land use types is consistent with the precision rate, indicating that the model comprehensively predicts patches where the land use type has changed before and after.

[0138] Figure 4 The continuous spatio-temporal evolution simulation results of the lake evolution and urban expansion by the model of the present invention are shown. It can be clearly seen from the figure that the lake shows an obvious shrinking trend, while the town shows a significant expanding trend. The final simulation result of the model highly coincides with the actual land use situation. Compared with the traditional time-discrete land use simulation method, the continuous land use simulation method proposed by the present invention can utilize data more efficiently, accurately capture the changes of settlements with significant evolution trends such as lakes and towns, and thus obtain more accurate and reliable simulation results.

[0139] Example 2:

[0140] This example provides a land use spatio-temporal prediction system, including a memory and a processor. The memory includes a land use spatio-temporal prediction method program. When the land use spatio-temporal prediction method program is executed by the processor, the steps of a land use spatio-temporal prediction method as described in Example 1 are implemented.

[0141] Example 3:

[0142] This example provides a computer-readable storage medium. The computer-readable storage medium includes a land use spatio-temporal prediction method program. When the land use spatio-temporal prediction method program based on it is executed by a processor, the steps of a land use spatio-temporal prediction method as described in Example 1 are implemented.

[0143] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limiting the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. A method for spatiotemporal prediction of land use, characterized in that: The steps include: Collect, classify and integrate multi-source data to obtain basic data sets; constructing a multi-source dataset using the basic dataset; Setting the time interval for each round of evolution, using the multi-source data set combined with a preset method to obtain the development probability of each land use type, the transition probability matrix, and the distribution target of each round of land use type; Using multi-source data sets, the development probability of each land use type, the transition probability matrix and the distribution target of each round of land use type, an intelligent goal-oriented cellular automaton model is constructed, a dynamic sampling strategy is introduced, and the cellular automaton model is used to perform continuous spatiotemporal prediction of land use and output the prediction results.

2. A land use spatiotemporal prediction method according to claim 1, characterized in that: The multi-source data include land use data; elevation data and slope data under the terrain data category; temperature data, potential evapotranspiration data, normalized difference vegetation index data and rainfall data under the climate data category; population distribution data and night light data under the population data category; Data on urban trunk roads, urban secondary roads, and railways under the location conditions data category; data on general public budget revenue, the number of students in ordinary middle schools, regional GDP data, and the number of beds in various health institutions under the urban development data category.

3. A method for spatiotemporal prediction of land use according to claim 2, characterized in that: Classifying and integrating multi-source data includes the following steps: Utilize population distribution data and night light data from multi-source data to analyze population size, spatial distribution and activity characteristics, and generate more accurate population information data for the basic data set; In the processing of urban development data, the number of students in ordinary middle schools and the number of beds in various health institutions are specially considered to generate corresponding urban development indicator data and add them to the basic data set; Unify the time dimension and ensure that the basic data set covers a time span of n years.

4. A method for spatiotemporal prediction of land use according to claim 1, characterized in that: The method for constructing a multi-source dataset using a basic dataset includes the following steps: The land use data of two consecutive periods in the basic data set are extended and analyzed, and the land use change area is extracted to generate a land expansion map. On this basis, the driving factor data of the corresponding spatial position is extracted and the driving factor data is supplemented into the basic data set. Flatten, concatenate and stack different data types in the basic data set to generate a four-dimensional multi-source data structure, where the first dimension is the time dimension, the second dimension is the data dimension, and the third and fourth dimensions are the space dimensions, recording the coordinates of the pixels; Data cleaning of multi-source data, including identification and processing of missing values, outliers, and duplicate data: using bilinear interpolation to calculate missing data based on the weighted average of the four nearest input pixels; replacing outlier data by removing or taking values ​​from neighboring pixels; removing duplicate data; The data is normalized and the calculation formula is as follows: Among them, x * is the normalized data, x is the original data to be processed, and x min is the minimum value in the original data, x max It is the maximum value in the original data, and the data value after normalization is between 0 and 1; The normalized multi-source data set is compressed and output as a multi-source data set.

5. A method for spatiotemporal prediction of land use according to claim 1, characterized in that: The method for obtaining the development probability of each land use type using multi-source datasets includes the following steps: Flatten the time dimension and space dimension of the multi-source data set and merge them into the data dimension to form the input sample data set; The sample data set is randomly shuffled and divided into a training set and a validation set according to a preset ratio; The sample training set data set is iteratively trained using the gradient boosting decision tree model, and each round of decision trees is gradually constructed. The expression of the gradient boosting decision tree model is as follows: Among them, x is the driving factor, F0(x) is the initial decision tree model, ω t is the gradient boosting decision tree model parameter, T is the number of decision trees; f t (x) is the residual prediction value of decision tree t for driving factor x, and F(x) is the output gradient boosting decision tree model; After each round of decision tree training, the validation set is used to calculate the loss function value. The expression of the loss function is as follows: Among them, L(y,f(x,ω)) is the loss function, f(x,ω) is the corresponding gradient boosting decision tree model prediction value when the driving factor is x and the parameter is ω; (x i ,y i ) is the observed data of the ith sample in the sample with a number of m, x i is the driving factor of the sample, y i is the actual observed value of the sample; According to the loss trend of the validation set, select the number of decision trees and model parameters to achieve the best performance of the model on the validation set; Using the trained gradient boosting decision tree model, the relationship between driving factors and land use types is explored to obtain the contribution of driving factors to each land use type. The driving factors include land use, topography, climate, population, location conditions and urban development. Combined with future driving factor data, the gradient boosting decision tree model is used to predict the growth probability of each land use type in the future and output the development probability of each land use type.

6. A method for spatiotemporal prediction of land use according to claim 1, characterized in that: The method for obtaining a transition probability matrix using a multi-source data set comprises the following steps: Using the land use data of two consecutive periods in the multi-source data set, the land use type changes of each pixel in different periods are calculated to generate a land use transfer matrix; The land use transfer matrices of multiple rounds of time are summed and averaged to form a basic transfer probability matrix; According to policy needs or planning objectives, combined with the protection requirements or conversion restrictions of specific land use types, the basic transfer probability matrix is ​​adjusted and the final transfer probability matrix is ​​output.

7. A method for spatiotemporal prediction of land use according to claim 1, characterized in that: The method for obtaining the distribution target of each round of land use type using multi-source datasets includes the following steps: Using multi-source data sets, we collect statistics on land use data from multiple time periods and calculate the historical distribution quantity and change trend of each land use type; Based on historical trend data, combined with the limited nature of land resources, the constraints of policies and regulations, and the requirements of environmental protection, a linear model of multi-objective programming is constructed; Setting constraints for the linear model of the multi-objective programming, wherein the constraints include: limited land resources, restrictions of policies and regulations, and requirements for environmental protection; The distribution target of each round of land use type is calculated using the linear model of the multi-objective programming.

8. A method for spatiotemporal prediction of land use according to claim 1, characterized in that: The method for constructing an intelligent goal-oriented cellular automation model and performing continuous spatiotemporal prediction of land use includes the following steps: Using the base period multi-source data set, combined with the transition probability matrix and the current round of land use distribution targets, the cellular automaton model is initialized, and the initial land use type distribution and conversion probability of the model sample pixels are set; the dynamic sampling strategy weights and spatial location screening rules are initialized. The dynamic sampling strategy weights are used to control the extraction priority of the sample pixels, and the dynamic sampling strategy spatial location screening rules are used to guide the geographical scope of dynamic sampling. The sampling of the initial sample pixels must keep the quantitative proportion of each land use type consistent with the actual data; At the beginning of this round of evolution, the geographical range is selected through the dynamic sampling strategy spatial position screening rule, and sample pixels in the area with high dynamic sampling weight are preferentially extracted from the geographical range. The neighborhood effect of each sample pixel is calculated using the transition probability matrix. The expression is as follows: Among them, Ω i,k is the neighborhood effect of the kth land use type in the neighborhood of sample pixel i, ∑ N×N con(c i =k) ​​is the number of pixels belonging to the kth land use type in the neighborhood of sample pixel i with a radius of N, N is the neighborhood parameter, TPM k is the transition probability of converting the land use type of the corresponding sample pixel i into the kth land use type obtained from the transition probability matrix; The sample pixels with zero neighborhood effect are adjusted, and a multi-type random patch seeding mechanism based on the Monte Carlo method is introduced to gradually approach the true transfer probability through random sampling, and a threshold that gradually decreases with the number of evolutions is introduced to suppress the seeding of patches. After seeding, the neighborhood effect, the development probability of each land use type and the adaptive coefficient are combined to calculate the overall conversion probability of the sample pixel. The expression is as follows: Among them, TP i,k is the overall conversion probability of sample pixel i to the kth land use type, Ω i,k is the neighborhood effect of the kth land use type in the neighborhood of sample pixel i, F i,k is the probability of sample pixel i developing into the kth land use type, is the adaptive coefficient of the k-th land use type in the t-th round of evolution, and its expression is as follows: in, is the difference between the distribution target of land use type in the t-1th round of evolution and the number of land use type k at the beginning of the t-1th round, is the difference between the distribution target of land use type in the t-2th round of evolution and the number of land use type k at the beginning of the t-2th round, is the adaptive coefficient of the kth land use type in the t-1th round of evolution; The overall conversion probability is input into the roulette algorithm, the candidate land use types of the sample pixels are output, and the transition probability matrix is ​​combined to determine whether each sample pixel should undergo land use type conversion; Update the current round of land use distribution as the input of the next round of land use type distribution, and calculate the difference in sample demand for the next round And update the adaptive coefficient According to the adaptive parameters Adjust the weight of the dynamic sampling strategy, and adjust the spatial location screening rules of the dynamic sampling strategy according to the development probability and the proportion of regional sample pixels that have not changed in this round of evolution; Repeat the above evolution steps until the difference between the distribution quantity of land use types and the distribution target of this round of land use types converges to the set threshold, and output the land use distribution of this round; Reset the adaptive coefficients and weight values ​​of the dynamic sampling strategy, and perform the next round of evolution until all rounds of evolution are completed, and output the land use distribution results of continuous spatiotemporal prediction.

9. A land use spatiotemporal prediction system, characterized in that: The system includes: a memory and a processor, wherein the memory includes a land use spatiotemporal prediction method program, and when the land use spatiotemporal prediction method program is executed by the processor, the steps of a land use spatiotemporal prediction method as described in any one of claims 1 to 8 are implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a land use spatiotemporal prediction method program, and when the land use spatiotemporal prediction method program is executed by a processor, the steps of a land use spatiotemporal prediction method as described in any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Land utilization change prediction method and system based on data analysis and machine learning

    CN117114176A

Cited By

  • Soil drought monitoring method based on multi-source data dynamic optimization

    CN121682246A