A method for causal mining of geographic features considering topological neighborhood
By constructing topological neighborhood relationships and causal identification criteria, the problem of identifying the causal relationships of geographical elements was solved, and quantitative analysis of the direction and intensity of the causal effects of geographical elements was achieved, thereby improving the accuracy of urban planning and resource allocation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANHUI AGRICULTURAL UNIVERSITY
- Filing Date
- 2026-04-22
- Publication Date
- 2026-07-21
AI Technical Summary
Existing technologies struggle to identify causal relationships between geographic elements, as well as their direction and intensity, and fail to adequately consider topological neighborhood characteristics, resulting in insufficient accuracy and reliability of the analysis results.
By acquiring multi-type geographic element data, constructing topological neighborhood relationships, screening candidate causal relationship pairs, and controlling confounding variables in conjunction with causal identification criteria, a directed causal network of geographic elements is constructed to achieve quantitative identification of the direction and intensity of causal effects.
It improves the accuracy and interpretability of geographic element relationship analysis, effectively characterizes the structural connections between geographic elements, reduces the impact of spurious correlations, and enhances the accuracy and stability of causal relationship identification.
Smart Images

Figure CN122432231A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of geographic information processing and spatial data mining technology, and in particular relates to a method for causal mining of geographic elements that takes into account topological neighborhoods. Background Technology
[0002] With the development of urban informatization and geographic information technology, geographic feature analysis based on Points of Interest (POI) data has become an important tool for urban spatial research and application. Existing technologies typically acquire multi-type geographic feature data and analyze their spatial distribution characteristics. For example, they utilize methods such as cluster analysis, kernel density estimation, hotspot analysis, and spatial association rule mining to identify the spatial clustering characteristics and interrelationships of different categories of geographic features. These methods can reveal, to some extent, the spatial distribution patterns among urban functional facilities and are widely used in urban planning, business site selection, and public resource allocation.
[0003] However, most of the existing technologies mentioned above focus on analyses based on spatial distance or statistical correlation, which can only reflect the co-occurrence or proximity relationships between different categories of geographic elements, and are difficult to further identify their causal relationships and their direction and intensity. At the same time, existing methods typically do not fully consider the structural connectivity relationships between geographic elements, lacking modeling of topological neighborhood features such as spatial adjacency and path connectivity. Furthermore, in real urban environments, different categories of geographic elements are often influenced by multiple factors such as population distribution, land use, and transportation conditions. Existing technologies lack effective identification and control of confounding variables, easily generating spurious correlations, thereby reducing the accuracy and reliability of the analysis results.
[0004] Therefore, how to introduce topological neighborhood relationships that can characterize the structural connections between geographic elements based on multi-type geographic element data, and on this basis, screen candidate causal relationship pairs, and further combine causal identification criteria to control confounding variables, thereby achieving accurate identification of the direction and intensity of causal interactions between geographic elements, has become an urgent technical problem to be solved. This invention addresses the above problem by proposing a geographic element causal mining method that takes topological neighborhood into account, in order to improve the accuracy and interpretability of geographic element relationship analysis. Summary of the Invention
[0005] The purpose of this invention is to provide a method for causal mining of geographic elements that takes into account topological neighborhoods, in order to solve the problems mentioned in the background art.
[0006] This invention is implemented as follows: a method for causal mining of geographic elements that takes into account topological neighborhoods, comprising the following steps:
[0007] Acquire multi-type geographic feature data within the target area and preprocess it to obtain a set of geographic features containing spatial location coordinates and category attributes;
[0008] The target area is divided into research units based on the set of geographic elements to obtain several basic research units. Based on the category attributes of the geographic elements, the spatial distribution of different categories of geographic elements within each basic research unit is statistically analyzed to obtain distribution characteristic values.
[0009] A topological neighborhood relationship is constructed based on the spatial adjacency or connectivity relationship between basic research units, and a topological adjacency matrix is generated. The connectivity relationship includes path-based connectivity relationships.
[0010] Candidate causal relationship pairs are determined based on the topological adjacency matrix, and treatment variables, outcome variables, and confounding variables are determined for each candidate causal relationship pair for causal analysis.
[0011] The confounding variables are screened based on a preset causal identification criterion to determine the set of adjustment variables used for causal effect estimation;
[0012] Based on the set of adjustment variables, the causal effect of the candidate causal relationship pairs is estimated to obtain the direction and intensity of causal interaction between geographical elements.
[0013] Construct a directed causal network of geographic elements based on the direction and intensity of the causal effects.
[0014] As a further limitation of the technical solution of the present invention, the geographic element data is point of interest (POI) data, and the POI data contains at least two different categories of geographic elements.
[0015] As a further limitation of the technical solution of this embodiment of the invention, the step of dividing the target area into research units includes: using a clustering algorithm to cluster geographical elements to obtain multiple clusters, merging the cluster centers, and constructing a Voronoi diagram based on the merged cluster centers to form the basic research unit.
[0016] As a further limitation of the technical solution of the present invention, the distribution characteristic value includes at least one of the following: number of geographical elements, density, proportion, kernel density value or clustering index.
[0017] As a further limitation of the technical solution of the embodiment of the present invention, the topological adjacency matrix is a weighted matrix, and the values of the matrix elements in the weighted matrix are determined based on parameters characterizing the topological neighborhood relationship, including at least one of shared boundary length, connected path length, or number of connected paths.
[0018] As a further limitation of the technical solution of the present invention, the method for determining candidate causal relationship pairs is as follows: only geographical element pairs located within the same basic research unit, adjacent basic research units, or located in a preset order topological neighborhood are retained as candidate causal relationship pairs.
[0019] As a further limitation of the technical solution of this embodiment of the invention, the step of screening the confounding variables based on a preset causal identification criterion to determine the set of adjustment variables for causal effect estimation includes:
[0020] Based on the candidate causal relationships, an initial causal relationship graph is constructed for the corresponding processing variables, outcome variables, and confounding variables.
[0021] Identify the backdoor path between the processing variable and the result variable based on the preset causal identification criteria;
[0022] Filter out the target variables that can block the backdoor path from the mixed variables;
[0023] The target variable is defined as the set of adjustment variables used for causal effect estimation.
[0024] As a further limitation of the technical solution of this invention, the step of estimating the causal effect of the candidate causal relationship pairs based on the set of adjustment variables to obtain the direction and intensity of causal interaction between geographical elements includes:
[0025] Based on the candidate causal relationships, a causal effect estimation model is constructed for the corresponding set of treatment variables, outcome variables, and adjustment variables.
[0026] The causal effect value of the treatment variable on the outcome variable is calculated using the causal effect estimation model.
[0027] The direction of causal interaction between geographic elements is determined based on the correspondence between the processing variables and the outcome variables in the candidate causal relationship pair.
[0028] The strength of causal interaction between geographical elements is determined based on the magnitude of the causal effect value.
[0029] As a further limitation of the technical solution of the present invention, the causal effect estimation adopts at least one of the structural causal model, the potential outcome model, or the convergent cross-mapping method.
[0030] Compared with existing technologies, this invention, based on multi-type geographic element data, introduces topological neighborhood relationships to constrain candidate causal relationship pairs and combines causal identification criteria to screen confounding variables, thereby achieving quantitative identification of the direction and intensity of causal interactions between geographic elements. Compared with existing methods based solely on spatial distance or statistical correlation, this invention can effectively characterize the structural connections between geographic elements, reduce the impact of spurious correlations, and improve the accuracy and stability of causal relationship identification. Furthermore, by constructing a directed causal network of geographic elements, it achieves a holistic expression of the interaction mechanisms of multiple categories of geographic elements, exhibiting good interpretability. This method can be widely applied in fields such as urban functional layout optimization, commercial site selection analysis, public facility configuration, and spatial structure evolution research, demonstrating strong application value and promising prospects for wider application. Attached Figure Description
[0031] Figure 1 A flowchart of the method provided in the embodiments of the present invention;
[0032] Figure 2 This is a flowchart illustrating the process of determining the set of adjustment variables in the method provided in this embodiment of the invention;
[0033] Figure 3 This is a flowchart illustrating the determination of the intensity and direction of causal interaction in the method provided in this embodiment of the invention. Detailed Implementation
[0034] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0035] Figure 1 A flowchart of the method provided by an embodiment of the present invention is shown.
[0036] Specifically, a method for causal mining of geographic elements that takes into account topological neighborhoods includes the following steps:
[0037] Step S100: Obtain multi-type geographic feature data within the target area and perform preprocessing to obtain a set of geographic features containing spatial location coordinates and category attributes.
[0038] The geographic feature data is Point of Interest (POI) data, which contains at least two different categories of geographic features.
[0039] In this embodiment of the invention, the invention is applied to the field of geographic information processing and urban spatial analysis, and in particular relates to a method for causal mining of multiple types of geographic elements based on point of interest (POI) data, which can be used in scenarios such as urban functional facility layout analysis, spatial structure identification, and spatial decision support.
[0040] Specifically, the geographic feature data is preferably Points of Interest (POI) data, which originates from map platforms or geographic information databases and is used to characterize geographic entities within the target area that have clear spatial locations and category attributes. These geographic entities can include various types such as commercial facilities, transportation facilities, and public service facilities, with each geographic feature corresponding to a spatial location coordinate and its category attribute. By acquiring and preprocessing multi-type geographic feature data, a unified format set of geographic features can be obtained, providing a data foundation for subsequent analysis.
[0041] Current research on POI data mainly focuses on spatial distribution analysis, cluster feature identification, and spatial correlation mining. It typically only considers co-occurrence relationships or spatial proximity relationships between different categories of geographic elements, making it difficult to identify causal relationships between geographic elements and their direction and intensity. Furthermore, existing methods are mostly based on distance or statistical correlation analysis, failing to fully consider the structural connections between geographic elements and the influence of potential confounding factors, which can easily lead to spurious correlations and affect the accuracy of the analysis results.
[0042] To address the aforementioned issues, the core of this invention lies in unifying spatial location coordinates and category attributes during the data preprocessing stage, based on multi-type geographic element data, to form a set of geographic elements suitable for causal analysis. This provides data support for subsequent construction of candidate causal relationship pairs based on topological neighborhood relationships and estimation of causal effects. Simultaneously, by ensuring that the POI data contains at least two different categories of geographic elements, potential interaction relationships can be formed between different categories of geographic elements. This provides the foundation for subsequently determining processing variables, outcome variables, and confounding variables, thereby laying the groundwork for the identification and quantification of causal relationships between geographic elements.
[0043] Furthermore, the method for causal mining of geographic elements that takes into account topological neighborhoods also includes the following steps:
[0044] Step S200: Based on the set of geographic elements, the target area is divided into research units to obtain several basic research units. Based on the category attributes of the geographic elements, the spatial distribution of different categories of geographic elements within each basic research unit is statistically analyzed to obtain distribution characteristic values.
[0045] The steps for dividing the target area into research units include: using a clustering algorithm to cluster geographical elements to obtain multiple clusters, merging the cluster centers, and constructing a Voronoi diagram based on the merged cluster centers to form the basic research unit.
[0046] In this embodiment of the invention, after obtaining a set of geographic elements including spatial location coordinates and category attributes in step S100, the target area is divided into research units, and the spatial distribution of different categories of geographic elements is statistically analyzed at the research unit scale to form the distribution characteristic values required for subsequent causal analysis.
[0047] Specifically, the research unit division process is constructed based on the spatial distribution of geographic elements. First, based on the spatial coordinates of the geographic element set, a clustering algorithm is used to cluster the geographic elements, dividing the relatively densely distributed geographic elements into several clusters. This clustering algorithm can identify areas of similar density based on the spatial distance relationship between geographic elements, thus avoiding the spatial heterogeneity problem caused by simply using regular grid division. After obtaining multiple clusters, the center position of each cluster is further extracted as the cluster center.
[0048] Based on this, the cluster centers are merged. Specifically, cluster centers with a distance less than a preset threshold are merged according to the distance relationship between them, reducing the number of overly dense or fragmented research units and thus making the subsequently constructed research units more spatially balanced. After merging the cluster centers, a Voronoi diagram is constructed based on the merged cluster centers, dividing the target area into several irregular polygonal regions, each corresponding to a basic research unit. The basic research units formed in this way can better reflect the actual spatial distribution characteristics of geographic elements, ensuring that the geographic elements within each research unit have relatively consistent spatial density and functional attributes.
[0049] The distribution characteristic values include at least one of the following: quantity, density, proportion, kernel density value, or clustering index of geographic elements, used to characterize the distribution of different categories of geographic elements within the study unit. For example, the quantity of a certain category of geographic elements within each basic study unit can be counted, and the density value can be calculated by combining the area of the study unit; the proportion of the quantity of that category of geographic elements to the total quantity of all geographic elements within the study unit can also be calculated to reflect its proportion characteristic; the spatial distribution of geographic elements can be smoothed using kernel density estimation methods to obtain kernel density values; or the degree of clustering of geographic elements within the study unit can be quantitatively described using the clustering index.
[0050] Furthermore, the method for causal mining of geographic elements that takes into account topological neighborhoods also includes the following steps:
[0051] Step S300: Construct a topological neighborhood relationship based on the spatial adjacency or connectivity relationship between basic research units, and generate a topological adjacency matrix. The connectivity relationship includes path-based connectivity relationships.
[0052] The topological adjacency matrix is a weighted matrix, and the values of the matrix elements in the weighted matrix are determined based on parameters characterizing the topological neighborhood relationships. The parameters include at least one of the following: shared boundary length, connected path length, or number of connected paths.
[0053] In this embodiment of the invention, the topological neighborhood relationship is preferably determined based on the spatial adjacency or connectivity relationship between basic research units. Spatial adjacency refers to the relationship where two basic research units share a boundary in space, meaning the polygonal regions corresponding to the two research units are in contact at the boundary. Connectivity refers to the relationship where two basic research units have reachable paths, including path-based connectivity, such as connectivity formed through road networks, walking paths, or other spatial networks. When either of the above relationships exists between two basic research units, they can be considered to be in the same topological neighborhood.
[0054] After determining the topological neighborhood relationships, a topological adjacency matrix is constructed to formally express these relationships. The topological adjacency matrix is used to characterize whether topological neighborhood relationships exist between the basic research units and the strength of these relationships. Preferably, the topological adjacency matrix is a weighted matrix, where each element corresponds to a pair of basic research units, and its value reflects the strength of the topological neighborhood relationship between the corresponding research units.
[0055] The values of the matrix elements in the weighted matrix are determined based on parameters characterizing the topological neighborhood relationships. Specifically, they can be assigned values based on at least one of the shared boundary length, connected path length, or number of connected paths. For example, when two basic research units share a boundary, the shared boundary length can be used as the value of the matrix element; the longer the boundary, the stronger the adjacency relationship. When two basic research units are connected by a path, the reciprocal of the connected path length or the number of paths can be used as the weight; the shorter the path or the more paths, the higher the degree of connectivity.
[0056] For example, in one specific embodiment, a topological neighborhood relationship can be constructed for a certain urban area based on the basic research units obtained in step S200. For any two adjacent basic research units, if their corresponding Voronoi polygons share a boundary, the length of the shared boundary is calculated and used as the value of the corresponding matrix element; for basic research units that are not directly adjacent but reachable through a road network, the value of the corresponding matrix element is determined according to the shortest path length or the number of optional paths in the road network. In this way, a complete topological adjacency matrix can be obtained, which is used for the subsequent determination of candidate causal relationship pairs.
[0057] Furthermore, the method for causal mining of geographic elements that takes into account topological neighborhoods also includes the following steps:
[0058] Step S400: Based on the topological adjacency matrix, candidate causal relationship pairs are determined, and for each candidate causal relationship pair, processing variables, outcome variables, and confounding variables for causal analysis are determined.
[0059] The method for determining candidate causal relationship pairs is as follows: only geographical element pairs located within the same basic research unit, adjacent basic research units, or within a preset order topological neighborhood are retained as candidate causal relationship pairs.
[0060] In this embodiment of the invention, the determination of candidate causal relationship pairs is constrained based on the topological neighborhood relationships reflected by the topological adjacency matrix. Only geographical feature pairs located within the same basic research unit, adjacent basic research units, or within a topological neighborhood of a preset order are retained as candidate causal relationship pairs. Specifically, geographical feature pairs located within the same basic research unit reflect potential interactions within a local spatial unit; geographical feature pairs located within adjacent basic research units reflect spatial influence relationships between adjacent regions; and geographical feature pairs located within a topological neighborhood of a preset order characterize indirect interactions formed over a larger area through multi-level adjacency or path connectivity. Through these constraints, the search space for candidate causal relationships can be effectively reduced, avoiding the computational complexity and noise interference caused by performing full connectivity analysis among all geographical features.
[0061] After identifying candidate causal pairs, for each candidate causal pair, treatment variables, outcome variables, and confounding variables are determined for causal analysis. Preferably, for two different categories of geographic elements in any candidate causal pair, one category of geographic elements can be used as the treatment variable, and the other category as the outcome variable, thereby constructing a directed analysis relationship. The treatment and outcome variables can be quantitatively represented based on the distribution characteristic values obtained in step S200, for example, using the quantity, density, or proportion of the corresponding category of geographic elements within the basic research unit as variable values. The confounding variables are used to characterize other factors that may simultaneously affect the treatment and outcome variables. They can originate from the distribution characteristic values of other categories of geographic elements within the same basic research unit, or from external spatial attribute data related to the basic research unit.
[0062] For example, in one specific embodiment, multiple types of Points of Interest (POI) data within a city can be selected as a set of geographic features, and divided into multiple basic research units according to step S200. For a certain basic research unit and its adjacent research units, restaurant-type geographic features and public transport stop-type geographic features can be used to form candidate causal relationship pairs, where public transport stop-type geographic features are used as processing variables, restaurant-type geographic features are used as outcome variables, and the number or density of restaurants within the corresponding research unit is used as the value of the outcome variable. At the same time, the distribution characteristic values of other categories of geographic features such as shopping malls and entertainment venues can be used as confounding variables for interference control in subsequent causal analysis.
[0063] Furthermore, the method for causal mining of geographic elements that takes into account topological neighborhoods also includes the following steps:
[0064] Step S500: Based on the preset causal identification criteria, the confounding variables are screened to determine the set of adjustment variables used for causal effect estimation.
[0065] Specifically, Figure 2 A flowchart showing the set of adjustment variables is provided.
[0066] The process of selecting the set of adjustment variables for causal effect estimation based on a preset causal identification criterion includes the following steps:
[0067] Step S501: Construct an initial causal relationship graph based on the candidate causal relationships for the corresponding processing variables, outcome variables, and confounding variables;
[0068] Step S502: Identify the backdoor path between the processing variable and the result variable according to the preset causal identification criteria;
[0069] Step S503: Select target variables that can block the backdoor path from the mixed variables;
[0070] Step S504: The target variable is determined as a set of adjustment variables for causal effect estimation.
[0071] In this embodiment of the invention, the initial causal relationship graph is used to describe the potential dependencies between processing variables, outcome variables, and various confounding variables. It can be represented in the form of a directed graph, where nodes correspond to different variables and edges are used to represent the possible influence relationships between variables.
[0072] The preset causal identification criterion is preferably a backdoor criterion, which identifies backdoor paths formed by connecting mixed variables instead of directly acting on the processing variables in the initial causal relationship graph by analyzing the paths from the processing variables to the result variables.
[0073] In step S503, key confounding variables located on the backdoor path can be selected according to the backdoor criterion, so that after controlling these variables, the non-causal relationship between the processing variables and the result variables is eliminated.
[0074] The set of adjustment variables is used to control the subsequent causal effect estimation process to ensure that the causal effect of the treatment variable on the outcome variable can be estimated under the condition of excluding the influence of confounding factors.
[0075] For example, in one specific embodiment, based on the candidate causal relationship pairs determined in step S400, a causal relationship graph can be constructed by using bus stop-type geographic elements as processing variables, restaurant-type geographic elements as outcome variables, and the distribution characteristic values of supermarket-type and entertainment venue-type geographic elements as confounding variables. In this causal relationship graph, paths such as "bus stop → supermarket → restaurant" or "bus stop ← supermarket → restaurant" may exist. Based on the backdoor criterion, backdoor paths formed through supermarket-type geographic elements can be identified, and the distribution characteristic values corresponding to supermarket-type geographic elements can be used as variables to be controlled, thus incorporating them into the set of adjustment variables.
[0076] Furthermore, the method for causal mining of geographic elements that takes into account topological neighborhoods also includes the following steps:
[0077] Step S600: Based on the set of adjustment variables, estimate the causal effect of the candidate causal relationship pairs to obtain the direction and intensity of causal interaction between geographical elements.
[0078] Specifically, Figure 3 A flowchart illustrating the strength and direction of causal interactions is provided.
[0079] Specifically, estimating the causal effects of the candidate causal relationships based on the set of adjustment variables to obtain the direction and intensity of causal interactions between geographical elements includes the following steps:
[0080] Step S601: Based on the candidate causal relationship, construct a causal effect estimation model for the corresponding set of processing variables, outcome variables and adjustment variables. The causal effect estimation adopts at least one of the following: structural causal model, potential outcome model or convergent cross-mapping method.
[0081] Step S602: Calculate the causal effect value of the treatment variable on the outcome variable using the causal effect estimation model;
[0082] Step S603: Determine the causal direction between geographic elements based on the correspondence between the processing variables and the outcome variables in the candidate causal relationship pair;
[0083] Step S604: Determine the intensity of causal interaction between geographical elements based on the magnitude of the causal effect value.
[0084] In this embodiment of the invention, when constructing a causal effect estimation model, the treatment variable is used as the input variable of the causal effect, the outcome variable is used as the output variable of the causal effect, and each variable in the set of adjustment variables is introduced into the model as a control variable, thereby establishing a causal relationship model between the treatment variable and the outcome variable under the condition of controlling for confounding factors.
[0085] The causal effect value is used to quantify the degree of influence of changes in the processing variable on the outcome variable under the condition of controlling the set of adjustment variables. It can be obtained directly through model parameter estimation or model output results.
[0086] In step S603, if a certain category of geographic features is set as a processing variable and another category of geographic features is set as a result variable, then the corresponding causal direction is from the geographic feature corresponding to the processing variable to the geographic feature corresponding to the result variable.
[0087] The larger the causal effect value, the more significant the influence of the treatment variable on the outcome variable, and the stronger the corresponding causal effect.
[0088] For example, in one specific embodiment, multi-type Points of Interest (POI) data within a city can be selected as a set of geographic features. Based on step S400, restaurant-type geographic features and bus stop-type geographic features are determined to constitute candidate causal relationship pairs. Bus stop-type geographic features are used as processing variables, restaurant-type geographic features as outcome variables, and the distribution characteristic values of supermarket-type geographic features are included as part of the set of adjustment variables. When constructing the causal effect estimation model, a structural causal model or a latent outcome model can be used to model the impact of changes in bus stop distribution on restaurant distribution, and the causal effect value is calculated under the condition of controlling for confounding factors such as supermarket distribution.
[0089] Furthermore, the method for causal mining of geographic elements that takes into account topological neighborhoods also includes the following steps:
[0090] Step S700: Construct a directed causal network of geographic elements based on the causal action direction and causal action intensity.
[0091] In this embodiment of the invention, the directed causal network of geographic elements can be represented by a directed graph structure, wherein nodes represent different categories of geographic elements, and edges represent the causal relationships between geographic elements. For each candidate causal relationship pair, based on the correspondence between the processing variable and the result variable determined in step S603, the geographic element corresponding to the processing variable is taken as the starting node, and the geographic element corresponding to the result variable is taken as the ending node, thereby determining the direction of the directed edge; simultaneously, based on the causal effect value obtained in step S604, it is used as the weight or strength identifier of the directed edge to characterize the strength of the corresponding causal effect.
[0092] Using the above method, multiple candidate causal pairs can be integrated into a unified directed causal network of geographic elements, thereby comprehensively characterizing the interaction structure between different categories of geographic elements. During the construction process, directed edges can be filtered based on causal effect values, for example, only retaining relationships with causal effect values higher than a preset threshold, to reduce noise impact and improve the stability and interpretability of the network structure.
[0093] For example, in one specific embodiment, a directed causal network containing multiple types of geographical elements such as restaurants, bus stops, and supermarkets can be constructed based on multi-type Points of Interest (POI) data within a city. If step S600 determines that bus stops have a significant causal effect on restaurants, then directed edges pointing from bus stops to restaurants are established in the network, with the corresponding causal effect value used as the edge weight. Simultaneously, if supermarkets also have a causal effect on restaurants or bus stops, then directed edges in the corresponding directions are established. In this way, a network structure reflecting the interaction relationships between urban functional facilities can be formed.
[0094] In another implementation, directed causal networks of geographic elements are constructed in different target regions to verify the applicability of this method in different spatial scenarios. For example, multi-type Points of Interest (POI) data are acquired within multiple city areas, and corresponding directed causal networks of geographic elements are constructed using the same method. In each city, the directed relationships and their strengths among various categories of geographic elements are obtained based on the division of research units, the construction of topological neighborhood relationships, and the estimation of causal effects. By comparing the directed causal networks of geographic elements constructed in different cities, the differences in causal interaction patterns among geographic elements in different regions can be analyzed. For example, in one city, catering geographic elements are mainly affected by transportation geographic elements, while in another city, they may be more affected by commercial geographic elements. In this way, the interaction mechanism between multiple types of geographic elements can be reflected from the perspective of network structure, providing a basis for subsequent spatial structure analysis.
[0095] It should be understood that although the steps in the flowcharts of the various embodiments of the present invention are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the various embodiments may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.
[0096] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0097] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0098] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.
[0099] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for causal mining of geographic elements that takes into account topological neighborhood, characterized in that, Includes the following steps: Acquire multi-type geographic feature data within the target area and preprocess it to obtain a set of geographic features containing spatial location coordinates and category attributes; The target area is divided into research units based on the set of geographic elements to obtain several basic research units. Based on the category attributes of the geographic elements, the spatial distribution of different categories of geographic elements within each basic research unit is statistically analyzed to obtain distribution characteristic values. A topological neighborhood relationship is constructed based on the spatial adjacency or connectivity relationship between basic research units, and a topological adjacency matrix is generated. Candidate causal relationship pairs are determined based on the topological adjacency matrix, and treatment variables, outcome variables, and confounding variables are determined for each candidate causal relationship pair for causal analysis. The confounding variables are screened based on a preset causal identification criterion to determine the set of adjustment variables used for causal effect estimation; Based on the set of adjustment variables, the causal effect of the candidate causal relationship pairs is estimated to obtain the direction and intensity of causal interaction between geographical elements. Construct a directed causal network of geographic elements based on the direction and intensity of the causal effects.
2. The method for causal mining of geographic elements considering topological neighborhood as described in claim 1, characterized in that, The geographic feature data is Point of Interest (POI) data, which contains at least two different categories of geographic features.
3. The method for causal mining of geographic elements considering topological neighborhood as described in claim 1, characterized in that, The steps for dividing the target area into research units include: Clustering algorithms are used to cluster geographic features to obtain multiple clusters. The cluster centers are merged, and a Voronoi diagram is constructed based on the merged cluster centers to form the basic research unit.
4. The method for causal mining of geographic elements considering topological neighborhood as described in claim 1, characterized in that, The distribution characteristic values include at least one of the following: number of geographical elements, density, proportion, kernel density value, or clustering index.
5. A method for causal mining of geographic elements considering topological neighborhood as described in claim 1, characterized in that, The topological adjacency matrix is a weighted matrix, and the values of the matrix elements in the weighted matrix are determined based on parameters characterizing the topological neighborhood relationships. The parameters include at least one of the following: shared boundary length, connected path length, or number of connected paths.
6. The method for causal mining of geographic elements considering topological neighborhood as described in claim 1, characterized in that, The method for determining candidate causal relationship pairs is as follows: only geographical element pairs located within the same basic research unit, adjacent basic research units, or within a preset order topological neighborhood are retained as candidate causal relationship pairs.
7. A method for causal mining of geographic elements considering topological neighborhood as described in claim 1, characterized in that, The steps of screening the confounding variables based on preset causal identification criteria to determine the set of adjustment variables for causal effect estimation include: Based on the candidate causal relationships, an initial causal relationship graph is constructed for the corresponding processing variables, outcome variables, and confounding variables. Identify the backdoor path between the processing variable and the result variable based on the preset causal identification criteria; Filter out the target variables that can block the backdoor path from the mixed variables; The target variable is defined as the set of adjustment variables used for causal effect estimation.
8. A method for causal mining of geographic elements considering topological neighborhood as described in claim 1, characterized in that, The steps of estimating the causal effects of the candidate causal relationship pairs based on the set of adjustment variables, and obtaining the direction and intensity of causal interactions between geographic elements, include: Based on the candidate causal relationships, a causal effect estimation model is constructed for the corresponding set of treatment variables, outcome variables, and adjustment variables. The causal effect value of the treatment variable on the outcome variable is calculated using the causal effect estimation model. The direction of causal interaction between geographic elements is determined based on the correspondence between the processing variables and the outcome variables in the candidate causal relationship pair. The strength of causal interaction between geographical elements is determined based on the magnitude of the causal effect value.
9. A method for causal mining of geographic elements considering topological neighborhood as described in claim 8, characterized in that, The causal effect estimation employs at least one of the following: a structural causal model, a potential outcome model, or a convergent cross-mapping method.