A land analysis method and system based on a deep forest model and a FLUS model
By combining the deep forest model and the FLUS model, the problems of low land change simulation accuracy and model complexity in the existing technology are solved, and more accurate land use change simulation and driving mechanism analysis are achieved, and the model is easy to use and stable.
Patent Information
- Application Number
- CN202211045930.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-29
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-08-29
AI Technical Summary
When the prior art simulates the complex nonlinear relationship between land changes and natural environment and socio-economic factors, the expression ability is weak, resulting in low simulation accuracy, and the deep neural network model is complex, the calculation efficiency is low, and it is susceptible to superparameters.
The land analysis method based on the deep forest model and FLUS model is adopted to model the nonlinear mapping relationship between driver factors and land use changes through the deep forest model, and the land use change simulation is carried out in combination with the FLUS model to obtain more accurate land use distribution data and driver factor contribution weights.
The accuracy of the land use change simulation of the FLUS model is improved, and more accurate analysis of the land use driving mechanism is provided. The model is easier to use than the deep neural network, less affected by randomness, and more stable.
Smart Images

Figure CN115374709B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of geographic information science, and particularly to a land analysis method and system based on a deep forest model and a FLUS model. Background Art
[0002] Land is an important place for the interaction between social systems and natural systems. Human activities have a profound impact on land use patterns, and land changes, in turn, affect the natural environment and human society. Studying historical and future land changes is an important basis for assessing the impact of land changes and formulating effective mitigation measures, and thus realizing the sustainable development of human society. The CA model can be used to explore the interaction between land changes and natural environment and socioeconomic factors, simulate the complex spatio-temporal dynamics of land, and explore the driving mechanisms of land changes. Currently, it has been widely applied to the simulation study of land changes. The FLUS model introduces an adaptive inertia mechanism on the basis of CA and is one of the most widely used CA models today. It has been successfully applied to the simulation study of land use changes at the regional scale and the global scale.
[0003] Modeling the complex non-linear relationships between various spatial and non-spatial driving factors and land changes and mining accurate land change rules are important components of the CA model. In early research, traditional data mining methods, including machine learning, bionics, and ensemble learning methods, were used to mine CA land change rules. However, the expression ability of these methods is weak, resulting in low simulation accuracy of the model.
[0004] Recent research has attempted to apply deep neural network methods to mine CA land change rules. Although deep neural networks have sufficient expression ability, the models are complex, the computational efficiency is low, and they are easily affected by hyperparameters, with low usability and difficulty in being used for the driving mechanisms of land changes.
[0005] Therefore, the existing technology still needs to be further improved. Summary of the Invention
[0006] The object of the present invention is to provide a land analysis method and system based on a deep forest model and a FLUS model, to overcome the problems of poor expression ability of traditional machine learning models and excessive complexity of deep neural networks, improve the accuracy of FLUS model land use change simulation, provide the possibility for the importance analysis of land use change driving factors, and endow the FLUS model with the ability to accurately analyze land use driving mechanisms.
[0007] To achieve the above object, the present invention provides a land analysis method and system based on a deep forest model and a FLUS model.
[0008] In a first aspect, the present invention provides a land analysis method based on a deep forest model and a FLUS model. The method includes: obtaining driving factor data and land use distribution data that drive land use change;
[0009] According to the land expansion analysis strategy, sampling the multi-period land use change data and aggregating the driving factor data to obtain a training sample set;
[0010] Inputting the training sample set into a deep forest model for training to obtain land change development suitability data for each land;
[0011] Inputting the land change development suitability data, neighborhood effect, conversion cost, and inertia coefficient into a FLUS model for simulation to obtain simulated land use distribution data;
[0012] Calculating the driving factor data set through the deep forest model to obtain contribution degree weight data of each driving factor to each land change.
[0013] In a further implementation, the step of sampling the multi-period land use distribution data according to the land expansion analysis strategy and aggregating the driving factor data to obtain a training sample set further includes:
[0014] Using a superposition analysis method to extract expansion areas from multi-period historical land use distribution data;
[0015] Performing stratified random sampling within the expansion areas to obtain sample positions;
[0016] Constructing the training sample set according to the driving factor data and the land use data corresponding to the sample positions.
[0017] In a further implementation, the step of inputting the training sample set into a deep forest model for training to obtain land change development suitability data for each land further includes:
[0018] The training sample set is represented in the following manner:
[0019] {x 1 ,x 1 ,…,x i ,…,x N ,y j}, where x i represents the value of the i-th driving factor, N is the number of driving factors, y j is the land use type corresponding to the land change, and M is the number of land use types;
[0020] Model the non - linear mapping relationship between driving factors and land changes through a deep forest model, and take the average value of the classification probability vectors output by the base forests in the last layer of the deep forest model as the suitability value for the development of land use changes;
[0021] The classification probability vector is represented in the following way:
[0022] {P 1 ,P 2 ,…,P j ,…,P M}
[0023] The average value of the classification probability vector is calculated by the following formula:
[0024]
[0025] where P j is the classification probability of land use type y j , H t (X) is the classification result of a single decision tree, X is the feature vector corresponding to the driving factor x i in the training sample set, T is the number of decision trees in the ordinary forest, and I is the indicator function.
[0026] In a further implementation, calculating the contribution degree weight data of each driving factor to each land change through the deep forest model for the driving factor dataset further includes:
[0027] The deep forest model includes a multi - level cascade structure and each level has several ordinary forest models, and the ordinary forest is composed of decision trees;
[0028] Select a splitting rule, and calculate the Gini coefficient of the node where the driving factor dataset is located according to the splitting rule; where the splitting rule is defined as θ, θ=(i, t), i represents a driving factor variable, t represents the splitting threshold, the node where the driving factor dataset is located is defined as Node, and the Gini coefficient is defined as G, specifically as follows:
[0029]
[0030]
[0031] Divide the driving factor dataset into two new nodes according to the following splitting rule:
[0032] Node left ={(X,y)|x i ≤t}
[0033] Node right =Node / Nodeleft
[0034] Among them, Node left and Node right are two new nodes after division, and the number of samples contained in the two data sets are m left and m right ;
[0035] The Gini coefficient sum of the two new nodes is calculated using the following formula:
[0036]
[0037] The impurity reduction value of the division rule θ is calculated using the following formula:
[0038] D(Node|θ) = G(Node) - G(Node|θ)
[0039] The nodes are divided by maximizing the impurity reduction value to construct a decision tree: Among them, dividing the nodes by maximizing the impurity reduction value is denoted as θ * , θ * = argmax θ D(Node|θ).
[0040] In a further embodiment, the maximized impurity reduction value is used as the contribution weight of the driving factor to evaluate the contribution degree of the driving factor to each land change.
[0041] In a further embodiment, using the maximized impurity reduction value as the contribution weight of the driving factor further includes:
[0042] Training each base forest model of each layer of forest model in the deep forest model to obtain the contribution weight of each driving factor in each base forest;
[0043] Calculating the average value of the contribution weights of each driving factor in each base forest and the average value of the contribution weights of each driving factor in each layer of forest model;
[0044] Taking the average of the average values of the contribution weights of each driving factor in each layer of forest model to obtain the integrated contribution degree, and forming the contribution degree weight data of each driving factor to each land change from the integrated contribution degree.
[0045] In a further embodiment, the simulated land use distribution data is evaluated, and the real land use distribution data is compared and evaluated with the simulated land use distribution data of the DF-FLUS model to obtain the accuracy of the simulated land use distribution data of the DF-FLUS model.
[0046] Second aspect, the present invention provides a land analysis system based on a deep forest model and a FLUS model, wherein the system includes:
[0047] A data collection module for obtaining driving factor data and multi-period land use distribution data that drive land use changes;
[0048] A data sampling module for sampling the multi-period land use distribution data according to a land expansion analysis strategy and aggregating the driving factor data to obtain a training sample set;
[0049] A model training module for training the training sample set by inputting it into a deep forest model to obtain land change development suitability data;
[0050] A data processing module for inputting the land change development suitability data, neighborhood effect, conversion cost, and inertia coefficient into a FLUS model for simulation to obtain simulated land use distribution data;
[0051] A data evaluation module for calculating the driving factor data set through the deep forest model to obtain contribution degree weight data of each driving factor to each land change.
[0052] Third aspect, the present invention further provides a computer device, including a processor and a memory, the processor is connected to the memory, the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory so that the computer device executes the steps of implementing the above method.
[0053] Fourth aspect, the present invention further provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of implementing the above method are realized.
[0054] The present invention provides a land analysis method and system based on a deep forest model and a FLUS model. Compared with the prior art, its beneficial effects are as follows:
[0055] By modeling the relationship between driving factors and land use changes based on a deep forest model, more accurate land use change rules can be obtained compared with traditional machine learning methods. At the same time, the model is easier to use than existing deep neural networks, can effectively improve the accuracy of land use change simulation of the FLUS model, is less affected by randomness, is more stable, and endows the FLUS model with the ability to accurately analyze the land use driving mechanism. Description of the Drawings
[0056] Figure 1 It is a schematic flowchart of a land analysis method based on a deep forest model and a FLUS model provided by an embodiment of the present invention;
[0057] Figure 2 It is a schematic diagram of the method for simulating land use change and analyzing driving mechanisms based on the deep forest model and the FLUS model in the embodiments of the present invention;
[0058] Figure 3 It is a framework diagram of the deep forest model provided by the embodiments of the present invention;
[0059] Figure 4 It is a comparison diagram of the simulation accuracy of the DF-FLUS of the present invention and the simulation accuracy of the ordinary method;
[0060] Figure 5 It is a comparison diagram of the distribution details of the actual land use pattern in the study area in 2010 and the land use patterns simulated by the present invention and the ordinary method provided by the embodiments of the present invention;
[0061] Figure 6 It is the contribution degree of the driving factors of the land use change data in the study area from 2000 to 2010 calculated by the present invention;
[0062] Figure 7 It is a comparison diagram of the random volatility of the analysis results of the contribution degree of the driving factors between the present invention and the ordinary method;
[0063] Figure 8 A block diagram of a land analysis system based on the deep forest model and the FLUS model provided by the embodiments of the present invention;
[0064] Figure 9 It is a schematic diagram of the structure of the computer device provided by the embodiments of the present invention. Detailed implementation manners
[0065] The following specifically clarifies the implementation manners of the present invention in conjunction with the accompanying drawings. The presentation of the embodiments is only for illustrative purposes and should not be construed as a limitation of the present invention. The accompanying drawings are for reference and illustration only and do not constitute a limitation on the scope of patent protection of the present invention, because many changes can be made to the present invention without departing from the spirit and scope of the present invention.
[0066] In one embodiment, as Figure 1 shown, the present invention provides a land analysis method based on the deep forest model and the FLUS model, and the method includes:
[0067] S1. Obtain the driving factor data and multi-period land use distribution data that drive land use change;
[0068] Among them, the driving factor data includes natural environmental factors and socio-economic factors;
[0069] Natural environmental factors include elevation, slope, distance to rivers, annual average temperature, annual average precipitation, temperature seasonality, and precipitation seasonality;
[0070] Socio - economic factors include distance to provincial centers, distance to city centers, distance to county centers, distance to airports, distance to highways and railways, and distance to general roads;
[0071] Land - use distribution data includes arable land, forest land, water bodies, and construction land.
[0072] S2. According to the land - expansion analysis strategy, sample the multi - period land - use distribution data and aggregate the driving - factor data to obtain a training sample set;
[0073] Preferably, the step of sampling the multi - period land - use distribution data according to the land - expansion analysis strategy and aggregating the driving - factor data to obtain a training sample set further includes:
[0074] Adopt an overlay analysis method to extract the expansion areas from the multi - period historical land - use distribution data;
[0075] Conduct stratified random sampling within the expansion areas to obtain sample locations;
[0076] Construct the training sample set according to the driving - factor data and the land - use data corresponding to the sample locations.
[0077] The land - expansion analysis strategy extracts the "expansion" areas from the multi - period historical land - use distribution data through overlay analysis, conducts stratified random sampling within the expansion areas to obtain sample locations, and then constructs training samples based on the driving - factor data and land - use data. This sampling method only focuses on new land types and ignores old land types. Therefore, while capturing the characteristics of land - use changes, it effectively reduces the complexity of sampling and model training.
[0078] S3. Input the training sample set into a deep - forest model for training to obtain the suitability data for the development of each land change;
[0079] Preferably, the step of inputting the training sample set into a deep - forest model for training to obtain the suitability data for the development of each land change further includes:
[0080] The training sample set is represented in the following way:
[0081] {x 1 ,x 1 ,…,x i ,…,x N ,y j}, where x irepresents the value of the i-th driving factor, N is the number of driving factors, and y j is the land use type corresponding to the land change, and M is the number of land use types;
[0082] The non-linear mapping relationship between the driving factors and the land change is modeled by a deep forest model, and the average value of the classification probability vectors output by the base forest in the last layer of the deep forest model is used as the development suitability value of the land use change;
[0083] The classification probability vector is represented in the following way:
[0084] {P 1 , P 2 , …, P j , …, P M}
[0085] The average value of the classification probability vector is calculated by the following formula:
[0086]
[0087] where P j is the classification probability of the land use type y j , H t (X) is the classification result of a single decision tree, X is the feature vector corresponding to the driving factor x i in the training sample set, T is the number of decision trees in the ordinary forest, and I is the indicator function.
[0088] The deep forest model is a deep learning method based on the ensemble learning strategy, with the characteristics of few hyperparameters, in-model feature transformation, hierarchical feature processing, and adaptive model complexity. Using the deep forest model to model the non-linear mapping relationship between the driving factors and the land change, the development suitability values of various land changes are calculated. The deep forest is an ensemble forest model with a cascaded structure, which consists of K layers, and each layer consists of F ordinary forest models (the models can be random forest, extreme forest). Each ordinary forest in each layer of the cascaded structure receives all the features output by the previous layer, and then outputs M high-level features. Then, the high-level features output by F ordinary forests (a total of F×M) and the original input features (a total of N) are connected to form a new feature vector and input into the next layer of the model.
[0089] Using the deep forest model with high expressiveness and availability to mine the CA land change rules can significantly improve the land use change simulation accuracy of FLUS.
[0090] S4. Input the land change development suitability data, neighborhood effect, conversion cost, and inertia coefficient into the FLUS model for simulation to obtain the simulated land use distribution data;
[0091] Preferably, the simulated land use distribution data is evaluated, and the real land use distribution data is compared and evaluated with the simulated land use distribution data of the DF-FLUS model to obtain the accuracy of the simulated land use distribution data of the DF-FLUS model.
[0092] The deep forest model can adaptively determine the complexity of the model. When using the default parameters, it can determine the model size adapted to the problem scale, avoiding the reduction of fitting accuracy caused by too large or too small networks in the CNN model. Through comparison and analysis, it is found that the land use pattern simulated by the present invention is closer to the actual situation than the ordinary methods, highlighting the high accuracy characteristics of the DF-FLUS model, that is, using the deep forest model can effectively improve the simulation accuracy of land use change of FLUS.
[0093] S5. Calculate the contribution degree weight data of each driving factor to each land change through the deep forest model for the driving factor data set.
[0094] Preferably, the calculation of the contribution degree weight data of each driving factor to each land change through the deep forest model for the driving factor data set further includes:
[0095] The deep forest model includes a multi-level cascade structure and each layer has several ordinary forest models, and the ordinary forest is composed of decision trees.
[0096] Select a partitioning rule, and calculate the Gini coefficient of the node where the driving factor data set is located according to the partitioning rule; wherein, the partitioning rule is defined as θ, θ = (i, t), i represents a driving factor variable, t represents the partitioning threshold, the node where the driving factor data set is located is defined as Node, and the Gini coefficient is defined as G, as follows:
[0097]
[0098]
[0099] Partition the driving factor data set into two new nodes according to the following partitioning rule:
[0100] Node left = {(X, y)|x i ≤ t}
[0101] Node right = Node / Node left
[0102] where, Node left and Node rightFor the two new nodes after division, the number of samples contained in the two data sets is m left and m right ;
[0103] The Gini coefficient sum of the two new nodes is calculated using the following formula:
[0104]
[0105] The impurity reduction value of the division rule θ is calculated using the following formula:
[0106] D(Node|θ) = G(Node) - G(Node|θ)
[0107] The node is divided by maximizing the impurity reduction value to construct a decision tree: Among them, dividing the node by maximizing the impurity reduction value is represented as θ * θ * = argmax θ D(Node|θ);
[0108] The contribution weight of the driving factor is taken as the maximum impurity reduction value, and the contribution degree of the driving factor to each land change is evaluated;
[0109] The above taking the maximum impurity reduction value as the contribution weight of the driving factor further includes:
[0110] Each basic forest model in each layer of the deep forest model is trained to obtain the contribution weight of each driving factor in each basic forest;
[0111] The average value of the contribution weights of each driving factor in each basic forest and the average value of the contribution weights of each driving factor in each layer of the forest model are calculated;
[0112] The average values of the contribution weights of each driving factor in each layer of the forest model are averaged to obtain the integrated contribution degree, and the contribution degree weights data of each driving factor to each land change are constituted by the integrated contribution degree;
[0113] In one embodiment, the variable importance analysis method is adopted, that is, important variables can divide samples into more certain categories, so they have lower impurity values. Therefore, the mean decrease in impurity (MDI) in the forest model can be used for variable importance evaluation. The weighted average of the impurity reduction values of a variable at all nodes of all decision trees in the forest is the MDI of the variable, and the MDI of all variables will be normalized. Therefore, MDI can also evaluate the contribution weight of the factor.
[0114] Since each layer of the deep forest model is composed of forest models, the variable importance calculated by the internal basic forest can be used for the contribution weight analysis of the driving factor.
[0115] For each basic forest n in layer 1, perform k - times of training to obtain k internal models {Fn.1, Fn.2, …, Fn.k}. Each model will calculate a corresponding contribution degree during training, and finally a total of k variable contribution degrees {Cn.1, Cn.2, …, Cn.k} are obtained. Average the k contribution degrees output by a forest to obtain Cn. Repeat the above operations for all random forests and extreme forests in layer 1 to obtain {C1, C2, …, Cn}, and finally take their average as the final integrated contribution degree. This method makes full use of all MDI information calculated during the construction process of the deep forest model, effectively reducing the impact of randomness on the Gini - coefficient importance method. At the same time, the calculation of variable importance is synchronized with the training process of the deep forest model, without additional time cost;
[0116] It makes full use of all MDI information calculated during the construction process of the deep forest model, can effectively reduce the impact of randomness on importance assessment, and more accurately analyze the driving mechanism of land - use change.
[0117] Based on the deep forest model to model the relationship between driving factors and land - use change, more accurate land - use change rules (development suitability) can be obtained compared with traditional machine - learning methods. At the same time, the model is easier to use than existing deep neural networks, and can effectively improve the accuracy of land - use change simulation of the FLUS model. The variable - importance analysis method of the present invention is less affected by randomness, more stable, and endows the FLUS model with the ability to accurately analyze the driving mechanism of land - use change.
[0118] In one embodiment, the research object in the present invention is the Pearl River Delta region of China. The data used in the research area are: CLUD land - use distribution data in 2000 and 2010, and the land - use data are classified into 4 categories (cultivated land, forest land, water body, construction land). Natural - environment factors (elevation, slope, distance to rivers, annual average temperature, annual average precipitation, temperature seasonality, precipitation seasonality) and socio - economic factors (distance to provincial center, distance to city center, distance to county center, distance to airport, distance to highway and railway, distance to general road). All factors are normalized to [0, 1] to eliminate the influence of scale effect. The spatial resolution of the data is uniformly resampled to 100m.
[0119] As Figure 2 shown, the method of the present invention includes the following steps:
[0120] Step 1: Collect the CLUD land use distribution data for 2000 and 2010, collect the spatial vector database, use GIS software to calculate the distance from each pixel in the area to the feature, and generate a distance driving factor layer with a resolution of 100m. Collect elevation data with a resolution of 100m, use GIS software to calculate the slope of each pixel in the area, and generate a slope driving factor layer. Collect climate raster data, project and resample it to a resolution of 100m, and finally obtain 13 driving factors;
[0121] Step 2: Use the expansion analysis strategy from the land use data of 2000 and 2010 to randomly sample 4 types of sample locations in layers, namely "0: remain unchanged", "1: change to cultivated land", "2: change to forest land", and "3: change to construction land". In this example, it is assumed that the water body remains unchanged. For types 1, 2, and 3, 5% of the samples are sampled respectively. For type 0, the number of its samples is the same as the total number of samples of the first three types. Based on the obtained sample locations, obtain sample data from the driving factor data and the land use distribution data;
[0122] Step 3: Input the samples into the deep forest model for training to fit the relationship between the driving factors and the land use change types. The parameter of the deep forest, the number of base forests F, is set to 2, and the number of decision trees in each base forest is 100, both of which are default parameters. Other comparison models are also trained using default parameters. After the model training is completed, input the complete driving factor data set to obtain the land use change development suitability surface of the study area;
[0123] Step 4: As Figure 3 shown, use the obtained land use change development suitability surface, integrate the neighborhood effect, conversion cost, and inertia coefficient to calculate the total conversion probability of land use change, and then use the roulette wheel to determine the new land use type of each pixel. Repeat this step until the number of land uses reaches the requirement and then stop the simulation. In this example, the land demand in 2010 is used as the termination condition of the simulation respectively;
[0124] Preferably, as Figure 4 shown, the present invention has a great improvement in the overall accuracy compared with the ordinary method. Especially when compared with the traditional deep learning model CNN, CNN uses default parameters without parameter tuning and only has a low simulation accuracy. While the deep forest model can adaptively determine the complexity of the model. In the case of using default parameters, it can determine the model size suitable for the problem scale, avoiding the reduction of fitting accuracy caused by too large or too small networks in the CNN model;
[0125] Among them, as Figure 5 and Figure 6As shown, the actual land use pattern in the Pearl River Delta in 2010 is compared with the land use patterns simulated by the present invention and the conventional method. The land use pattern simulated by the present invention is closer to the actual one than the conventional method.
[0126] Step 5: Calculate the contribution degree of the driving factors to land change using the MDI information saved during the training process of the deep forest model. The calculation method is to average the normalized MDI values output by the k-fold cross-validation training of all basic forests in the first layer of the deep forest model.
[0127] Preferably, as Figure 7 shown, the contribution degrees of the respective driving factors calculated by applying the present invention to the land use change in the study area from 2000 to 2010;
[0128] The volatility of the ranking of the contribution degrees of the driving factors calculated by the present invention and the conventional method for multiple times is compared. The method of the present invention makes full use of all the MDI information calculated during the construction of the deep forest model, can effectively reduce the influence of randomness on importance assessment, and more accurately analyze the driving mechanism of land use change.
[0129] Based on the above land analysis method based on the deep forest model and the FLUS model, as Figure 8 shown, an embodiment of the present invention provides a land analysis system based on the deep forest model and the FLUS model, and the system includes:
[0130] A data collection module 101, configured to obtain driving factor data for driving land use distribution and multi-period land use distribution data;
[0131] A data sampling module 102, configured to sample the multi-period land use distribution data according to the land expansion analysis strategy, and combine the driving factor data to obtain a training sample set;
[0132] A model training module 103, configured to input the training sample set into a deep forest model for training to obtain land change development suitability data;
[0133] A data processing module 104, configured to input the land change development suitability data, neighborhood effect, conversion cost, and inertia coefficient into the FLUS model for simulation to obtain simulated land use distribution data;
[0134] A data evaluation module 105, configured to calculate the contribution degree weight data of each driving factor to each land change by calculating the driving factor data set through the deep forest model.
[0135] Based on the above land analysis method and system based on the deep forest model and the FLUS model, as Figure 9As shown in the figure, a computer device provided by an embodiment of the present invention includes a memory, a processor, and a transceiver, which are connected through a bus; the memory is used to store a set of computer program instructions and data, and can transmit the stored data to the processor, and the processor can execute the program instructions stored in the memory to perform the steps of the above method.
[0136] Among them, the memory may include a volatile memory or a non-volatile memory, or may include both a volatile and a non-volatile memory; the processor may be a central processing unit, a microprocessor, an application specific integrated circuit, a programmable logic device, or a combination thereof. By way of example but not limitation, the above programmable logic device may be a complex programmable logic device, a field programmable gate array, a generic array logic, or any combination thereof.
[0137] In addition, the memory may be a physically independent unit or may be integrated with the processor.
[0138] Those of ordinary skill in the art can understand that Figure 9 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have the same component arrangement.
[0139] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0140] In summary, the embodiments of the present invention provide a land analysis method and system based on a deep forest model and a FLUS model. By modeling the relationship between driving factors and land use change based on the deep forest model, more accurate land use change rules (development suitability) can be obtained compared with traditional machine learning methods. At the same time, the model is easier to use than existing deep neural networks, can effectively improve the accuracy of land use change simulation of the FLUS model, is less affected by randomness, is more stable, and endows the FLUS model with the ability to accurately analyze the land use driving mechanism.
[0141] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the counting principle of the present invention, several improvements and replacements can be made, and these improvements and replacements should also be regarded as the protection scope of the present invention.
Claims
1. A land analysis method based on a deep forest model and a FLUS model, characterized in that: The method includes: Obtaining driving factor data for driving land use change and multi-period land use distribution data; Sampling the multi-period land use distribution data according to a land expansion analysis strategy, and aggregating the driving factor data to obtain a training sample set; Inputting the training sample set into a deep forest model for training to obtain suitability data for the development of each land change; Inputting the suitability data for the development of the land change, neighborhood effect, conversion cost, and inertia coefficient into a FLUS model for simulation to obtain simulated land use distribution data; Calculating, through the deep forest model, contribution degree weight data of each driving factor to each land change for a driving factor data set; The step of inputting the training sample set into a deep forest model for training to obtain suitability data for the development of each land change further includes: The training sample set is represented in the following manner: {x 1 , x 2 , …, x i , …, x N , y j}, where x i represents the value of the i-th driving factor, N is the number of driving factors, and y j is the land use type corresponding to the land change, and M is the number of land use types; Modeling the non-linear mapping relationship between driving factors and land change through a deep forest model, and taking the average value of the classification probability vectors output by the basic forests in the last layer of the deep forest model as the suitability value for the development of land use change; The classification probability vector is represented in the following manner: {P 1 ,P 2 ,…,P j ,…,P M} The average value of the classification probability vectors is calculated using the following formula: Among them, P j is the classification probability of land use type y j , H t (X) is the classification result of a single decision tree, where X is the feature vector corresponding to the driving factor x in the training sample set i , T is the number of decision trees in the ordinary forest, and I is the indicator function 2. The land analysis method based on a deep forest model and a FLUS model according to claim 1, characterized in that: The step of sampling the multi-period land use distribution data according to a land expansion analysis strategy and aggregating the driving factor data to obtain a training sample set further includes: Using an overlay analysis method to extract expansion areas from multi-period historical land use distribution data; Performing stratified random sampling within the expansion areas to obtain sample locations; Constructing the training sample set according to the driving factor data and the land use distribution corresponding to the sample locations.
3. The land analysis method based on a deep forest model and a FLUS model according to claim 1, characterized in that The step of calculating, through the deep forest model, contribution degree weight data of each driving factor to each land change for a driving factor data set further includes: The deep forest model includes a multi-level cascade structure and each layer has several ordinary forest models, and the ordinary forest is composed of decision trees; Selecting a splitting rule, and calculating the Gini coefficient of the node where the driving factor data set is located according to the splitting rule; wherein, the splitting rule is defined as θ, θ = (i, t), i represents a driving factor variable, t represents the splitting threshold driving factor, the node where the data set is located is defined as Node, and the Gini coefficient is defined as G, specifically as follows: Dividing the driving factor data set into two new nodes according to the following splitting rule: Node left = {(X,y)|x i ≤t} Node right = Node / Node left Among them, Node left and Node right are two new nodes after division, and the number of samples contained in the two data sets are m left and m right ; Calculating the sum of the Gini coefficients of the two new nodes using the following formula: Calculating the impurity reduction value of the splitting rule θ using the following formula: D(Node|θ) = G(Node) - G(Node|θ) Construct a decision tree by partitioning nodes using the maximum impurity reduction value: Among them, partitioning nodes using the maximum impurity reduction value is denoted as θ * , θ * = argmax θ D(Node|θ).
4. The land analysis method based on a deep forest model and a FLUS model according to claim 3, characterized in that, Taking the maximum impurity reduction value as the contribution weight of the driving factor, evaluate the contribution degree of the driving factor to each land change.
5. A land analysis method based on a deep forest model and a FLUS model as described in claim 4, characterized in that the step of taking the maximum impurity reduction value as the contribution weight of the driving factor further includes: Training each basic forest model in each layer of the forest model in the deep forest model to obtain the contribution weights of each driving factor in each basic forest; Calculating the average value of the contribution weights of each driving factor in each basic forest and the average value of the contribution weights of each driving factor in each layer of the forest model; Taking the average of the average values of the contribution weights of each driving factor in each layer of the forest model to obtain an integrated contribution degree, and forming contribution degree weight data of each driving factor to each land change from the integrated contribution degree.
6. A land analysis method based on a deep forest model and a FLUS model as described in claim 1, characterized in that Comparing and evaluating the real land use distribution data with the simulated land use distribution data simulated by the DF-FLUS model to obtain the accuracy of the simulated land use distribution data of the DF-FLUS model.
7. A land analysis system based on a deep forest model and a FLUS model, characterized in that the system includes: A data collection module for obtaining driving factor data and multi-period land use distribution data that drive land use changes; A data sampling module for sampling the multi-period land use distribution data according to a land expansion analysis strategy and combining the driving factor data to obtain a training sample set; A model training module for inputting the training sample set into a deep forest model for training to obtain suitability data for the development of each land change; A data processing module for inputting the suitability data for the development of the land change, neighborhood effect, conversion cost, and inertia coefficient into a FLUS model for simulation to obtain simulated land use distribution data; A data evaluation module for calculating a driving factor data set through the deep forest model to obtain contribution degree weight data of each driving factor to each land change; The step of inputting the training sample set into a deep forest model for training to obtain suitability data for the development of each land change further includes: The training sample set is represented in the following manner: {x 1 , x 2 , …, x i , …, x N , y j}, where x i represents the value of the i-th driving factor, N is the number of driving factors, and y j is the land use type corresponding to the land change, and M is the number of land use types; Modeling the non-linear mapping relationship between the driving factor and the land change through a deep forest model, and taking the average value of the classification probability vectors output by the basic forest in the last layer of the deep forest model as the suitability value for the development of land use change; The classification probability vector is represented in the following manner: {P 1 ,P 2 ,…,P j ,…,P M} The average value of the classification probability vector is calculated using the following formula: Among them, P j is the classification probability of land use type y j , H t (X) is the classification result of a single decision tree, X is the driving factor x in the training sample set i corresponding feature vector, T is the number of decision trees in the ordinary forest, and I is the indicator function.
8. A computer device, characterized in that: It includes a processor and a memory. The processor is connected to the memory. The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory so that the computer device executes the method described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: A computer program is stored in the computer-readable storage medium, and when the computer program is run, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Tensor-based land utilization simulation method, system and device and storage medium
CN111814368A
Change scene simulation method based on FLUS model and biodiversity model
CN113222316A