Land utilization driving mode excavation method and related equipment

By combining the method of minimum spanning tree and Bayesian network, the land use-driven mode is partitioned and optimized, which solves the problems of spatial heterogeneity, uncertainty quantification and partitioning process fragmentation of land use driving forces in the prior art, and realizes the precise excavation of land use-driven mode and captures complex variations.

CN120217062AActive Publication Date: 2025-06-27深圳市规划和自然资源数据管理中心(深圳市空间地理信息中心) +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510695211.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-06-27
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

The prior art has problems such as overly rigid partitioning, lack of uncertainty quantification mechanisms, and separation of driving force analysis and partitioning process in the spatial heterogeneity of land use driving force, making it difficult to accurately capture the spatial variation and uncertainty of complex driving mechanisms.

Method used

By obtaining land use classification data, socio-economic data, natural data and transportation data, processing and calculations are performed to obtain land use change characteristics, neighborhood land use characteristics and land use scenario characteristics. These features are input into the minimum spanning tree for partitioning, and the partitioning results are iteratively optimized using the Bayesian network until the iteration termination condition is met, and the optimal partition structure and Bayesian network are output.

Benefits of technology

Accurate excavation of the land use-driven model has been achieved, which can better capture the spatial variation and uncertainty of complex driving mechanisms, and provide more accurate land use change prediction and policy formulation support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217062A_ABST
    Figure CN120217062A_ABST
Patent Text Reader

Abstract

The invention provides a land utilization driving mode mining method and related equipment, and the method comprises the steps: carrying out the processing of obtained original raster data of a research region, and obtaining the processed raster data; calculating the processed raster data to obtain land utilization change features, neighborhood land utilization features and land utilization scene features of the research area; inputting the processed raster data, the land utilization change features, the neighborhood land utilization features and the land utilization scene features into a generated minimum spanning tree for partitioning to obtain a partitioning result, and performing iteration on the initial Bayesian network by using the partitioning result to obtain an iterated Bayesian network; optimizing a partitioning result according to the iterated Bayesian network until an iteration termination condition is met, outputting an optimal partitioning structure of the research area and taking the Bayesian network of each partition in the optimal partitioning structure as a land utilization driving mode mining result of the research area; accurate excavation of a land utilization driving mode can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of spatial data mining, and particularly to a method for mining land use driving patterns and related devices. Background Art

[0002] With the progress of urbanization and the development of social economy, the accurate identification of land use driving mechanisms has become a key and difficult issue in the research of land change science. Land use change is a complex spatial process, affected by multiple factors such as natural environment, social economy, and policy systems. Its driving mechanisms vary significantly in different regions and different periods. Accurately revealing the driving mechanisms of land use change has important theoretical and practical significance for formulating scientific national territorial space planning and promoting regional sustainable development.

[0003] However, the spatial heterogeneity problem of land use driving forces makes it difficult for global modeling methods to accurately capture local variation characteristics, affecting the accuracy of land use change prediction and the effectiveness of policy formulation. To address the spatial heterogeneity problem of land use driving forces, there are currently three main types of spatial zoning methods: zoning methods based on spatial clustering, zoning methods based on spatial optimization, and network-based zoning methods.

[0004] Zoning methods based on spatial clustering mainly include the K-Means clustering method, hierarchical clustering algorithms, and the Density-Based Spatial Clustering of Applications with Noise (DBSCAN), etc., which aggregate spatial units with similar attribute characteristics into the same region; such methods have high computational efficiency, but usually ignore spatial continuity constraints, resulting in the dispersion of regions of the same category in space and being difficult to directly apply to regional planning and management practices.

[0005] Zoning methods based on spatial optimization mainly include the Automated Zoning Procedure (AZP) and the Regionalization with dynamically constrained agglomerative clustering and partitioning (REDCAP), which optimize regional division by setting spatial continuity constraints to maximize within-region homogeneity; such methods can maintain spatial continuity, but have low computational efficiency when dealing with large-scale data and are prone to falling into local optimal solutions. At the same time, the optimization function with fixed parameters is difficult to capture non-linear change characteristics, resulting in a decline in boundary recognition accuracy.

[0006] The network-based partitioning method represents spatial relationships using network structures in graph theory and achieves spatial partitioning through methods such as network pruning or community discovery. Among them, widely used network partitioning methods include the Minimum Spanning Tree (MST) and the Spatial ’K’ luster Analysis by Tree Edge Removal (SKATER). Although spatial partitioning is generated by constructing an MST and pruning it, achieving a balance between spatial continuity and attribute similarity; however, the solution space of SKATER is limited by the initial MST structure, resulting in a lack of flexibility in the shape and scale of the partition. Therefore, in areas with significant heterogeneity such as the urban-rural transition zone, SKATER may generate suboptimal partitions due to the fixed tree structure and fail to capture the spatial variation of complex driving mechanisms.

[0007] In summary, the defects of the existing technologies are as follows: (1) Ignoring spatial continuity and the partition being too rigid; traditional clustering methods such as K-Means ignore the spatial proximity constraint, resulting in the spatial dispersion of areas in the same category, which is not conducive to regional management and policy implementation. Although MST-based methods such as SKATER maintain spatial continuity, the partition structure is too rigid and limited by the initial tree structure, making it difficult to adapt to the spatial variation of driving mechanisms in complex geographical environments, especially performing poorly in areas such as the urban-rural transition zone; (2) Lack of an uncertainty quantification mechanism; uncertainty in land use driving force analysis is a key consideration factor in formulating policies. Existing partitioning methods mostly focus on extracting deterministic patterns and lack an effective uncertainty quantification mechanism. Regression models assume linear relationships and independence among variables and are difficult to capture complex non-linear interactions. Association rule mining mainly focuses on the co-occurrence frequency of events and cannot express causal relationships and uncertainties; (3) The separation between the driving force analysis and the partitioning process; existing methods usually regard spatial partitioning and driving force analysis as two independent processes. First, spatial partitioning is performed based on certain preset criteria, and then driving force analysis is carried out separately within each partition. This separation may result in the partition results not being able to optimally serve the mining of driving mechanisms and affecting the accuracy of subsequent analysis. Summary of the Invention

[0008] The present invention provides a method for mining land use driving patterns and related devices, and its purpose is to achieve accurate mining of land use driving patterns.

[0009] To achieve the above purpose, the present invention provides a method for mining land use driving patterns, including: Step 1, obtaining land use classification data, socio-economic data, natural data, and traffic data of the research area as original raster data; Step 2: Process the original raster data to obtain the processed raster data; Step 3: Calculate the processed raster data to obtain the land use change characteristics, neighborhood land use characteristics, and land use scenario characteristics of the study area; Step 4: Input the processed raster data, land use change characteristics, neighborhood land use characteristics, and land use scenario characteristics into the generated minimum spanning tree for partitioning to obtain the partitioning result, and use the partitioning result to iterate the initial Bayesian network to obtain the iterated Bayesian network. The partitioning result is optimized according to the iterated Bayesian network until the iteration termination condition is met, and the optimal partitioning structure of the study area and the Bayesian network of each partition in the optimal partitioning structure are output as the mining result of the land use driving pattern in the study area.

[0010] Furthermore, Step 2 includes: Resample the land use classification data in the original raster data to obtain multiple grid cells; Align the socioeconomic data, natural data, and traffic data in the original raster data to each grid cell according to the spatio-temporal position to obtain the aligned grid cells; Discretize the socioeconomic data, natural data, and traffic data in each grid cell to obtain the discretized socioeconomic data, natural data, and traffic data; Use the aligned grid cells and the discretized socioeconomic data, natural data, and traffic data as the processed raster data.

[0011] Furthermore, Step 3 includes: Count the mode of land use classification in each aligned grid cell as the land use characteristic of the study area; Calculate the land use characteristic change vector for multiple years using the land use characteristics to obtain the land use change characteristics of the study area; Calculate the neighborhood of each aligned grid cell through the nearest neighbor algorithm and count the mode of land use characteristics in each neighborhood to obtain the neighborhood land use characteristics of each grid cell in the study area; Convert the land use classification data in each aligned grid cell into vector data, generate a buffer, and remap the buffer and assign it to each aligned grid cell to obtain the land use scenario characteristics of the study area.

[0012] Furthermore, before inputting the processed raster data, land use change characteristics, neighborhood land use characteristics, and land use scenario characteristics into the generated minimum spanning tree for partitioning, it also includes: Based on each aligned grid cell, an initial spatial adjacency graph is established through a triangulation network, and the edges in the initial spatial adjacency graph are randomly assigned values to generate a weight matrix for representing an undirected graph. Based on the weight matrix, a spatial adjacency graph is obtained. The Kruskal algorithm is used to generate an initial minimum spanning tree on the spatial adjacency graph. The initial minimum spanning tree includes edges, indicating the number of aligned grid cells.

[0013] Furthermore, based on an update operator, the initial Bayesian network is iterated using the partitioning result to obtain an iterated Bayesian network, and the partitioning result is optimized according to the iterated Bayesian network.

[0014] Furthermore, the update operator includes: The first update operator is used to randomly select a partition in the partitioning result and split it into two sub - partitions; The second update operator is used to randomly select two adjacent partitions in the partitioning result and merge them into one partition; The third update operator is used to randomly select a partition in the partitioning result, split it into two sub - partitions, and then select two adjacent partitions and merge them into one partition; The fourth update operator is used to randomly assign values between 0 and 1 to the edge weights of the edges located in the same partition and regenerate the minimum spanning tree.

[0015] Furthermore, the iteration termination conditions include: The number of iterations reaches a preset total number of iterations; There is no partitioning result and Bayesian network in several consecutive iterations that are better than the previous iteration.

[0016] The present invention also provides a device for mining land - use driving patterns, including: An acquisition module, configured to acquire land - use classification data, socio - economic data, natural data, and traffic data of a research area as original raster data; A processing module, configured to process the original raster data to obtain processed raster data; A calculation module, configured to calculate the processed raster data to obtain land - use change characteristics, neighborhood land - use characteristics, and land - use scenario characteristics of the research area; An iterative module is used to partition the processed raster data, land use change features, neighborhood land use features, and land use scenario features into the generated minimum spanning tree to obtain a partitioning result, and use the partitioning result to iterate the initial Bayesian network to obtain an iterated Bayesian network. The partitioning result is optimized according to the iterated Bayesian network until the iteration termination condition is met, and the best partitioning structure of the research area and the Bayesian network of each partition in the best partitioning structure are output as the land use driving pattern mining result of the research area.

[0017] The present invention also provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the land use driving pattern mining method is implemented.

[0018] The present invention also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the land use driving pattern mining method is implemented.

[0019] The above solution of the present invention has the following beneficial effects: The present invention processes the obtained land use classification data, social economic data, natural data, and traffic data of the research area as the original raster data to obtain the processed raster data; calculates the processed raster data to obtain the land use change features, neighborhood land use features, and land use scenario features of the research area; inputs the processed raster data, land use change features, neighborhood land use features, and land use scenario features into the generated minimum spanning tree for partitioning to obtain a partitioning result, and uses the partitioning result to iterate the initial Bayesian network to obtain an iterated Bayesian network. The partitioning result is optimized according to the iterated Bayesian network until the iteration termination condition is met, and the best partitioning structure of the research area and the Bayesian network of each partition in the best partitioning structure are output as the land use driving pattern mining result of the research area. Compared with the prior art, the present invention combines the minimum spanning tree and the Bayesian network. After each partition of the minimum spanning tree is completed, the partition quality is immediately evaluated by the Bayesian network, and the minimum spanning tree is guided to optimize the previous partition. This tightly coupled design can achieve precise mining of the land use driving pattern.

[0020] Other beneficial effects of the present invention will be described in detail in the subsequent specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is a schematic flowchart of an embodiment of the present invention; Figure 2 is a schematic structural diagram of the Bayesian network in an embodiment of the present invention; Figure 3 This is a schematic structural diagram of the land use driving mode mining device in the embodiment of the present invention; Figure 4 This is a schematic structural diagram of the terminal device in the embodiment of the present invention. Detailed implementation manners

[0022] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0023] In the description of the present invention, it should be noted that the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.

[0024] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0025] The present invention provides a land use driving mode mining method and related equipment for existing problems.

[0026] As Figure 1 shown, the embodiment of the present invention provides a land use driving mode mining method, including: Step 1, obtaining land use classification data, socio-economic data, natural data, and traffic data of the research area as original raster data; Step 2, processing the original raster data to obtain processed raster data; Step 3, calculating the processed raster data to obtain the land use change characteristics, neighborhood land use characteristics, and land use scenario characteristics of the research area; Step 4, inputting the processed raster data, land use change characteristics, neighborhood land use characteristics, and land use scenario characteristics into the generated minimum spanning tree for partitioning to obtain a partitioning result, and using the partitioning result to iterate the initial Bayesian network to obtain an iterated Bayesian network. The partitioning result is optimized according to the iterated Bayesian network until the iteration termination condition is met, and the best partitioning structure of the research area and the Bayesian network of each partition in the best partitioning structure are output as the land use driving mode mining result of the research area; in the embodiment of the present invention, the best partitioning structure may include urban areas, agricultural areas, and mountainous areas, and the Bayesian network of each partition is used to show the driving mechanism of land use change within the partition.

[0027] In the embodiments of the present invention, the land use classification data uses existing Landsat image interpretation products. Landsat image interpretation products refer to the surface information and features extracted by processing and analyzing remote sensing image data obtained by Landsat satellites, usually including the current situation of land use, vegetation cover, water area distribution, urban expansion data, etc.; the social and economic data is characterized by three indicators: gross domestic product, population density data, and night light data. The gross domestic product uses existing, rasterized, and research area-corresponding gross domestic product data. The night light data selects NPP-VIIRS image data. NPP-VIIRS image data is a kind of night light remote sensing image data with higher spatial resolution and more accurate night light radiation intensity records; the natural data is mainly terrain data, represented by digital elevation models and their derived slope data. These natural data are based on the publicly available SRTM DEM products. The SRTM DEM products are digital elevation models (Digital Elevation Model) made from data obtained by the Shuttle Radar Topography Mission and uniformly resampled to a resolution of 300 meters; the traffic data uses the road network data of OpenStreetMap and the point of interest (POI) data of Baidu Maps to calculate the distance from the target area to the city center (citydis), the distance to the railway (CZT_D2R), the distance to the main road, and the distance to key facilities, etc. All data finally generate original raster data with a resolution of 300 meters.

[0028] Specifically, step 2 includes: Resample the land use classification data in the original raster data to obtain multiple grid cells; Align the social and economic data, natural data, and traffic data in the original raster data to each grid cell according to the spatio-temporal position to obtain the aligned grid cells; Discretize the social and economic data, natural data, and traffic data in each grid cell to obtain the discretized social and economic data, natural data, and traffic data; Use the aligned grid cells and the discretized social and economic data, natural data, and traffic data as the processed raster data.

[0029] In the embodiments of the present invention, the land use classification data in the original raster data is resampled into grid cells of 300 meters. The resampling is based on the mode of land use classification in the 10×10 grid cells. The resampling calculation formula is: ; Wherein, represents the resampling result, Represents land use classification data, Represents the mode operation, Represents the row and column of the grid cell, Represents the offset in the row direction of the raster data, Represents the offset in the column direction of the raster data.

[0030] Specifically, align the socioeconomic data, natural data, and traffic data in the original raster data to each grid cell according to the spatio-temporal position to obtain the aligned grid cell, including: Use arcgis to perform spatial calculations on the socioeconomic data, natural data, and traffic data in the original raster data to obtain continuous values; Then discretize based on the natural break method to obtain eigenvalue; Then assign the eigenvalue to the grid cell at the same spatio-temporal position according to the spatial position, and delete the grid cells whose quantiles of the eigenvalue are above the top 5%, to obtain the aligned grid cell. The representation form of the aligned grid cell is: ; Among them, Represents the th aligned grid cell, Represents the central coordinate of the th aligned grid cell, Represents the land use classification data in the th aligned grid cell, Represents the th aligned grid cell, and the th eigenvalue.

[0031] Specifically, in the embodiment of the present invention, the decision tree is used to discretize the socioeconomic data, natural data, and traffic data in each grid cell. The decision tree uses the information gain as the splitting index, and the discretization process is as follows: Given the continuous feature and the target variable (such as land use change type) , the expression for calculating the information gain is: ; Among them, Represents the information gain, is the feature discretized by the threshold , is the entropy of the target variable, is the conditional entropy; Based on the optimal splitting point , select the threshold that maximizes the information gain: ; When selecting the best split point, the effect of splitting is evaluated based on the reduction in entropy. At each step of splitting, an attempt is made to maximize the information gain to automatically find the "best" threshold for converting continuous data into discrete intervals, thus completing the discretization process.

[0032] Specifically, step 3 includes: Count the mode of land use classification in each aligned grid cell as the land use characteristic of the study area; Use the land use characteristics to calculate the land use characteristic change vectors for multiple years to obtain the land use change characteristics of the study area. The calculation expression is: ; where, represents the land use change characteristic in the th aligned grid cell, represents the land use characteristic of the th aligned grid cell in the th year; Calculate the neighborhood of each aligned grid cell through the nearest neighbor algorithm, and count the mode of land use characteristics in each neighborhood to obtain the neighborhood land use characteristics of each grid cell in the study area. The expression is: ; where, represents the neighborhood land use characteristic of the th aligned grid cell; Convert the land use classification data in each aligned grid cell into vector data, generate a buffer, and remap the buffer and assign it to each aligned grid cell to obtain the land use scenario characteristics of the study area.

[0033] In the embodiments of the present invention, remapping the buffer is for features 1 to 3, which respectively represent the cultivated land core area, the urban core area, the forest land core area, and the land use edge area.

[0034] Specifically, before inputting the processed raster data, land use change characteristics, neighborhood land use characteristics, and land use scenario characteristics into the generated minimum spanning tree for partitioning, it also includes: Based on each aligned grid cell, establish an initial spatial adjacency graph through a triangulation network, and randomly assign values to the edges in the initial spatial adjacency graph to generate a weight matrix for representing an undirected graph. Based on the weight matrix, obtain the spatial adjacency graph; Use the Kruskal algorithm to generate an initial minimum spanning tree on the spatial adjacency graph. The initial minimum spanning tree includes edges, Indicates the number of aligned grid cells.

[0035] In an embodiment of the present invention, based on each aligned grid cell, an initial spatial adjacency graph is established through a Delaunay triangulation network, expressed as: ; Wherein, represents the initial spatial adjacency graph, represents the set of grid cells, represents the set of edges; Then, for in the edges, randomly assign values from 0 to 1 to generate a weight matrix for characterizing an undirected graph : ; Finally, use the Kruskal algorithm to generate an initial minimum spanning tree with the number of edges being _, where is the number of grid cells ; The generated contains edges and is a spanning tree subgraph of , expressed as: ; Wherein, represents the edge connecting nodes and and satisfies is the minimum.

[0036] In an embodiment of the present invention, a Bayesian network is a probabilistic graphical model that represents conditional independence between variables through a directed acyclic graph, formally expressed as , where is the graph structure, is the parameter set, and the calculation method of the Bayesian network score is as follows: ; Wherein, is the number of variables, is the number of parent node configurations of node , is the number of values of node , represents the number of samples in the dataset where the value of node is and the parent node configuration is , , is the prior hyperparameter, , , is the gamma function.

[0037] To improve the computational efficiency, in the embodiments of the present invention, when performing Bayesian network training, the decomposability of the scoring function is utilized, and the network structure search process is optimized through a dynamic programming algorithm. Specifically, instead of directly searching in the huge space of directed acyclic graphs, the algorithm decomposes the problem of learning the globally optimal structure into a series of local optimization sub-problems: that is, for each variable (node) (where ), find the optimal parent node configuration and its corresponding optimal local score when its parent node set is limited to a certain subset of variables ; by systematically calculating, storing, and reusing these local optimal solutions for different variable subsets, dynamic programming can avoid repeated scoring calculations for a large number of possible parent node combinations, thereby significantly reducing the complexity of finding the globally optimal structure from super-exponential to single-exponential (with respect to the number of variables ), effectively improving the computational efficiency of structure learning.

[0038] The embodiments of the present invention also introduce a complexity penalty term to limit the network complexity: ; wherein, represents the complexity penalty term, is the penalty coefficient, is a network complexity metric, which is used to represent the ratio of the number of occurrences of the parent-child transaction categories in the dataset to the number of occurrences of the parent transaction category in the dataset, multiplied by the product of the parent transaction and the child transaction categories. In the optimization process, the closer the network score is to zero, the more significant the learning effect of the Bayesian network.

[0039] Specifically, based on the update operator, the initial Bayesian network is iterated using the partitioning result to obtain the iterated Bayesian network, and the partitioning result is optimized according to the iterated Bayesian network.

[0040] Specifically, the update operator includes: The first update operator is used to randomly select a partition in the partitioning result and split it into two sub-partitions. The acceptance probability of this update operator is equal to the operation probability ratio × the Bayesian network fitness ratio × (1 - penalty term c); The second update operator is used to randomly select two adjacent partitions in the partitioning result and merge them into one partition. The acceptance probability of this update operator is equal to the operation probability ratio × the Bayesian network fitness ratio / (1 - penalty term c); The third update operator is used to randomly select a partition in the partition result, split it into two sub - partitions, and then select two adjacent partitions to merge into one partition. The acceptance probability of this update operator is equal to the operation probability ratio × the Bayesian network fitness ratio × (1 - penalty term c); The fourth update operator is used to randomly assign values between 0 and 1 to the edge weights of edges located in the same partition, and regenerate the minimum spanning tree.

[0041] During the iteration process, based on the current partition result, use the above four update operators to generate candidate partitions, and then calculate the acceptance probability And compare it with a random number between 0 and 1. If the random number is greater than or equal to the acceptance probability, then use the candidate partition as the new partition result.

[0042] It should be noted that in order to avoid the optimization process falling into a local optimal solution, the embodiment of the present invention adopts an optimization mechanism similar to simulated annealing, and its core lies in adopting a non - deterministic state acceptance criterion. This criterion allows the system, during the iteration process, not only to accept candidate partition states that can improve the fitness score (usually meaning superior to or equal to in terms of scoring ), but also to accept, according to a specific probability states that may lead to a temporary decrease in fitness (that is inferior to ). This ability to randomly accept "inferior" solutions gives the optimization algorithm the possibility to get out of the local optimal attraction basin and explore a wider solution space, thereby increasing the chance of finally finding the global optimal or high - quality partition scheme.

[0043] The specific acceptance probability is calculated as follows: ; This formula shows that: when the fitness score of the candidate partition is not inferior to the current score , the acceptance probability is 1 (definitely accept); when is inferior to , the acceptance probability is the power of the ratio of the absolute value of the current score to the absolute value of the new score ( ). The absolute value is used here to handle the case where the score may be negative (such as the BIC score), and what is compared is the "good - bad degree" of the score rather than the simple numerical size.

[0044] Among them, the key parameter affecting the calculation of the acceptance probability (its role is similar to the temperature in the simulated annealing algorithm), and its reference value (or other relevant parameters directly affecting calculation) will change with the number of iterations Gradually decrease and decay according to the following formula: ; where, is the initial reference value, is the decay rate (usually set to 0.0005), is the current iteration number; As the iteration progresses, (or related parameters) decreases, resulting in a overall decrease in the probability of accepting inferior solutions, making the algorithm tend to explore widely (accept more inferior solutions) in the initial stage of the search and focus more on exploitation (concentrate near high-quality solutions) in the later stage, and finally converge.

[0045] The operation selection probability corresponding to the operator is dynamically adjusted according to the current number of partitions: when the number of partitions is 1, it is set to , and normalized; when the number of partitions is , it is set to , and normalized; in other cases, it is set to , and normalized. If a non-normalized probability vector is given, the formula for its normalization is: ; That is, each element in the vector is divided by the sum of all elements in .

[0046] It should be noted that in the embodiments of the present invention, for the dataset of each partition , a Bayesian network is used to train the model of a single partition. Here, the "model of a single partition" refers to constructing a local Bayesian network model that can best fit the data subset within the partition by using the exact structure learning method based on BIC scoring and dynamic programming and the corresponding parameter learning. Then, compare all partition sets after the operation with the partition set before the operation to obtain the overall fitness score .

[0047] The core part of the fitness calculation method is the contribution corresponding to a single partition , which is calculated based on the following formula: ; where, represents that the value of node in partition is and the parent node is configured to the number of samples, . The formula calculates the log-likelihood part of the data within the partition , which reflects the learned local Bayesian network model for the data in this partition degree of fit. In actual evaluation, usually the contributions of this item for all partitions (or more precisely, the complete local score including the model complexity penalty, such as ) are accumulated, and other global penalty terms (such as penalty for partition size) may be combined to obtain the overall fitness score , which essentially measures the degree to which the entire dataset can be explained by a series of local Bayesian network models under the current partition scheme, or the certainty of the rules. The better the score (for example, the smaller the absolute value of the BIC score), the higher the overall quality of the partition scheme and the corresponding network structure.

[0048] Specifically, the iteration termination conditions include: the number of iterations reaches the preset total number of iterations; there is no partition result and Bayesian network in several consecutive iterations that is better than the previous iteration.

[0049] The embodiments of the present invention use the land use classification data of a certain urban area from 2015 to 2020 to illustrate the provided method: First, select the original raster data within the scope of a certain urban area. The original raster data includes land use classification data, socio-economic data, natural data, and traffic data. The resolution of the land use classification data in the original raster data is 30 meters; Resample the land use classification data with a resolution of 30 meters into 300-meter grid cells according to a 10×10 window, align the socio-economic data, natural data, and traffic data in the original raster data into each grid cell, and delete abnormal grid cells; Statistically analyze the mode of the land use classification data in the aligned raster cells as the land use feature, calculate the annual change sequence of the land use feature from 2015 to 2020 as the land use change feature, calculate the neighboring cells of each aligned raster cell through the K8 nearest neighbor algorithm, and statistically analyze the mode of the land use feature in the neighborhood as the neighborhood land use feature; Convert all aligned raster cells into vector data, generate a 200m internal buffer zone, identify the cultivated land core area, urban core area, forest land core area, and land use edge area, and remap them into land use scenario features from 0 to 3; For continuous natural data and traffic data (such as altitude, slope, distance, etc.), apply the discretization method based on decision trees, and use information gain as the splitting index; for example, discretize altitude values into low altitude areas (0 - 200m), medium altitude areas (200 - 500m), and high altitude areas (>500m). For the already discretized categorical factors (such as land use types), retain their original category codes; Based on the aligned grid cells, establish a spatial adjacency graph through Delaunay triangulation Randomly assign weights from 0 to 1 to the edges in the graph to generate a weight matrix , and use Kruskal's algorithm to generate an initial minimum spanning tree on The initial minimum spanning tree contains , a total of edges ( is the number of grid cells). For the 15,241 grid cells in the study area, the initial minimum spanning tree contains 15,240 edges; Based on the initial partition data, train a Bayesian network model and calculate the BDeu score. Use the dynamic programming algorithm to optimize the network structure search. The initially trained Bayesian network shows that the land use change in the study area is mainly affected by the gross domestic product, traffic data, and neighborhood land use data. The BDeu score of the initial Bayesian network is -368,742.5; Partition iteration optimization; set the initial benchmark value , the decay rate , and the preset number of iterations is 15,000 times. During the optimization process, dynamically adjust the selection probabilities of the four update operators (Birth, Death, Change, Hyper) according to the current number of partitions. After 5,217 iterations, the optimal partition structure is obtained, which contains 5 spatially continuous partitions. The BDeu score of the Bayesian network is improved to -256,389.2, an improvement of 30.5% compared to the initial Bayesian network; Result evaluation; the optimal partition structure contains multiple spatially continuous partitions, namely the urban use area where land use change is mainly affected by population density data and gross domestic product, the plain agricultural area where land use change is mainly affected by the water area distribution in natural data, and the mountainous area where land use change is significantly restricted by slope and altitude. The Bayesian network structures of each partition are as shown in Figure 2 , clearly showing the differences in the driving mechanisms of land use change in different regions; To further optimize the connection method of the partition structure, according to the existing partition boundary constraints, the edge weights are resampled. The edge weights within the same partition are randomly assigned within the range of 0 - 0.5, and the edge weights between different partitions are randomly assigned within the range of 0.5 - 1. The Kruskal algorithm is used to generate a new minimum spanning tree. The newly generated tree structure maintains the original partition boundaries unchanged, while optimizing the connection relationships within the partitions, improving the interpretability and visualization effect of the model.

[0050] In the embodiment of the present invention, the obtained land use classification data, socioeconomic data, natural data, and traffic data of the research area are processed as the original raster data to obtain the processed raster data; the processed raster data is calculated to obtain the land use change characteristics, neighborhood land use characteristics, and land use scenario characteristics of the research area; the processed raster data, land use change characteristics, neighborhood land use characteristics, and land use scenario characteristics are input into the generated minimum spanning tree for partitioning to obtain the partitioning result, and the initial Bayesian network is iterated using the partitioning result to obtain the iterated Bayesian network. The partitioning result is optimized according to the iterated Bayesian network until the iteration termination condition is met, and the best partition structure of the research area and the Bayesian network of each partition in the best partition structure are output as the mining result of the land use driving pattern in the research area; compared with the prior art, in the embodiment of the present invention, the minimum spanning tree and the Bayesian network are combined. After each partition is completed by the minimum spanning tree, the partition quality is immediately evaluated by the Bayesian network, and the minimum spanning tree is guided to optimize the previous partition. This tightly coupled design can achieve accurate mining of the land use driving pattern.

[0051] Corresponding to the land use driving pattern mining method described in the above embodiment, as Figure 3 shown, the embodiment of the present invention further provides a land use driving pattern mining device 100, and the land use driving pattern mining device 100 includes: An acquisition module 101, configured to acquire the land use classification data, socioeconomic data, natural data, and traffic data of the research area as the original raster data; A processing module 102, configured to process the original raster data to obtain the processed raster data; A calculation module 103, configured to calculate the processed raster data to obtain the land use change characteristics, neighborhood land use characteristics, and land use scenario characteristics of the research area; The iterative module 104 is used to partition the processed raster data, land use change features, neighborhood land use features, and land use scenario features into the generated minimum spanning tree to obtain a partitioning result, and use the partitioning result to iterate the initial Bayesian network to obtain an iterated Bayesian network. The partitioning result is optimized according to the iterated Bayesian network until the iteration termination condition is met, and the best partitioning structure of the study area and the Bayesian network of each partition in the best partitioning structure are output as the mining result of the land use driving pattern in the study area.

[0052] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units, due to being based on the same concept as the method embodiment of the present application, for their specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details will not be repeated here.

[0053] Those skilled in the art can clearly understand that for the convenience and conciseness of description, only the above-mentioned division of each functional unit and module is used as an example. In practical applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be repeated here.

[0054] The embodiment of the present invention also provides a terminal device, such as Figure 4 shown. The terminal device D10 in this embodiment includes: at least one processor D100 ( Figure 4 only one processor is shown in the figure), a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100. When the processor D100 executes the computer program D102, the above-mentioned land use driving pattern mining method is implemented.

[0055] The terminal device D10 can be a computing device such as a desktop computer, a notebook, a palm computer, a server, a server cluster, and a cloud server. The terminal device may include, but is not limited to, a processor D100 and a memory D101. Those skilled in the art can understand, Figure 4This is only an example of the terminal device D10, which does not constitute a limitation on the terminal device D10. It may include more or fewer components than those shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0056] The so-called processor D100 may be a central processing unit (CPU, Central Processing Unit). This processor D100 may also be other general-purpose processors, digital signal processors (DSP, Digital Signal Processor), application specific integrated circuits (ASIC, Application Specific Integrated Circuit), field-programmable gate arrays (FPGA, Field-Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0057] In some embodiments, the memory D101 may be an internal storage unit of the terminal device D10, such as the hard disk or memory of the terminal device D10. In other embodiments, the memory D101 may also be an external storage device of the terminal device D10, such as a plug-in hard disk, smart media card (SMC, SmartMedia Card), secure digital (SD, Secure Digital) card, flash card (Flash Card), etc. equipped on the terminal device D10. Further, the memory D101 may also include both the internal storage unit and the external storage device of the terminal device D10. The memory D101 is used to store the operating system, application programs, boot loader (BootLoader), data, and other programs, such as the program code of the computer program, etc. The memory D101 may also be used to temporarily store the data that has been output or will be output.

[0058] It should be noted that for the information interaction, execution process, etc. between the above-mentioned devices / units, since they are based on the same concept as the method embodiments of this application, for their specific functions and the technical effects brought, please refer to the method embodiment part specifically, and details will not be elaborated here.

[0059] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.

[0060] The present invention also provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the method for mining the land use driving mode.

[0061] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of this application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the construction device / terminal device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc.

[0062] The above is the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for mining land use driving patterns, characterized in that Including: Step 1: Obtain the land use classification data, socioeconomic data, natural data, and traffic data of the research area as the original raster data; Step 2: Process the original raster data to obtain the processed raster data; Step 3: Calculate the processed raster data to obtain the land use change characteristics, neighborhood land use characteristics, and land use scenario characteristics of the research area; Step 4: Input the processed raster data, the land use change characteristics, the neighborhood land use characteristics, and the land use scenario characteristics into the generated minimum spanning tree for partitioning to obtain the partitioning result, and use the partitioning result to iterate the initial Bayesian network to obtain the iterated Bayesian network. The partitioning result is optimized according to the iterated Bayesian network until the iteration termination condition is met, and the best partitioning structure of the research area and the Bayesian network of each partition in the best partitioning structure are output as the mining result of the land use driving pattern in the research area.

2. The method for mining the land use driving mode according to claim 1, characterized in that The Step 2 includes: Resample the land use classification data in the original raster data to obtain multiple grid cells; Align the socioeconomic data, natural data, and traffic data in the original raster data to each grid cell according to the spatio-temporal position to obtain the aligned grid cells; Discretize the socioeconomic data, natural data, and traffic data in each grid cell to obtain the discretized socioeconomic data, natural data, and traffic data; Use the aligned grid cells and the discretized socioeconomic data, natural data, and traffic data as the processed raster data.

3. The method for mining land use driving patterns according to claim 2, wherein The Step 3 includes: Count the land use classification mode in each aligned grid cell as the land use characteristic of the research area; Calculate the land use characteristic change vector of multiple years using the land use characteristic to obtain the land use change characteristics of the research area; Calculate the neighborhood of each aligned grid cell through the nearest neighbor algorithm and count the mode of the land use characteristics in each neighborhood to obtain the neighborhood land use characteristics of each grid cell in the research area; Convert the land use classification data in each aligned grid cell into vector data, generate a buffer, and remap the buffer and assign it to each aligned grid cell to obtain the land use scenario characteristics of the research area.

4. The land use driving pattern mining method according to claim 3, characterized in that Before inputting the processed raster data, the land use change characteristics, the neighborhood land use characteristics, and the land use scenario characteristics into the generated minimum spanning tree for partitioning, it also includes: Based on each aligned grid cell, establish an initial spatial adjacency graph through a triangulation network, randomly assign values to the edges in the initial spatial adjacency graph to generate a weight matrix for representing an undirected graph, and obtain the spatial adjacency graph based on the weight matrix; Generate an initial minimum spanning tree on the spatial adjacency graph using the Kruskal algorithm, where the initial minimum spanning tree includes edges, indicating the number of aligned grid cells.

5. The land use driving mode mining method according to claim 4, characterized in that, Based on the update operator, use the partitioning result to iterate the initial Bayesian network to obtain the iterated Bayesian network. The partitioning result is optimized according to the iterated Bayesian network.

6. The method for mining land use driving patterns according to claim 5, characterized in that, The update operator includes: The first update operator is used to randomly select a partition in the partition result and split it into two sub - partitions; The second update operator is used to randomly select two adjacent partitions in the partition result and merge them into one partition; The third update operator is used to randomly select a partition in the partition result, split it into two sub - partitions, and then select two adjacent partitions and merge them into one partition; The fourth update operator is used to randomly assign values between 0 and 1 to the edge weights of the edges located in the same partition and regenerate the minimum spanning tree.

7. The method for mining land use driving patterns according to claim 1, wherein The iteration termination conditions include: The number of iterations reaches the preset total number of iterations; There is no partition result and Bayesian network in several consecutive iterations that are better than the previous iteration.

8. An apparatus for mining a land use driving mode, characterized in that, It includes: An acquisition module is used to acquire the land use classification data, socioeconomic data, natural data, and traffic data of the research area as the original raster data; A processing module is used to process the original raster data to obtain the processed raster data; A calculation module is used to calculate the processed raster data to obtain the land use change characteristics, neighborhood land use characteristics, and land use scenario characteristics of the research area; An iteration module is used to input the processed raster data, the land use change characteristics, the neighborhood land use characteristics, and the land use scenario characteristics into the generated minimum spanning tree for partitioning to obtain a partition result, and use the partition result to iterate the initial Bayesian network to obtain the iterated Bayesian network. The partition result is optimized according to the iterated Bayesian network until the iteration termination conditions are met, and the optimal partition structure of the research area and the Bayesian network of each partition in the optimal partition structure are output as the mining result of the land use driving pattern of the research area.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the land use driving pattern mining method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the land use driving pattern mining method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Land use change driving force identification method, system and apparatus

    CN108428007A

  • Land use change simulation method and system

    CN109543278A

  • Land utilization change simulation method based on deep learning

    CN113297174A

  • Land utilization change driving factor mining method based on random forest and public source geographic information

    CN114398951A

  • Urban function facility cluster effect mining method based on spatial causal discovery

    CN118365006A