Spatial co-location pattern mining method based on time weighting and kernel estimation
By introducing time weighting and kernel estimation methods in spatial homogeneous mode mining, the problem of traditional methods ignoring time factors and failing to effectively deal with mode weights is solved, and more efficient and accurate mode mining is achieved, which is suitable for dynamically changing application scenarios.
Patent Information
- Application Number
- CN202510293473.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-27
AI Technical Summary
The traditional spatial isometric mode mining method ignores time factors, which leads to misjudgment in dynamically changing application scenarios, and fails to effectively deal with the weight problem of different instances' contribution to the mode, resulting in the mining results containing a large amount of irrelevant or useless content.
A spatial homogeneous mode mining method based on time weighting and kernel estimation is proposed. By obtaining the time series data of the instance, preprocessing and filtering, calculating the existence time span of the spatial instance pair, filtering the weighted spatial instances, calculating the spatial feature distance weight, using the step-by-step merge method to form a higher-order candidate mode, and performing pattern pruning and verification, ensuring that the excavated mode has high accuracy and comprehensiveness.
It improves the efficiency and accuracy of pattern mining, can accurately identify frequent patterns that meet the needs of actual application, and is suitable for large-scale spatial data and scenarios with significant changes in time and space, reduces the generation of useless patterns, provides more accurate spatial and temporal relationships, and supports more accurate management and planning.
Smart Images

Figure CN120217310A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to technical fields such as spatial data mining, co-location pattern mining, and time series pattern mining, and particularly relates to a method for mining spatial co-location patterns based on time weighting and kernel estimation. Background Art
[0002] As an important direction in spatial data mining, spatial co-location pattern mining aims to reveal the implicit relationships between spatial data and discover valuable information from it. This information can be widely applied in various fields and bring far-reaching impacts. For example, in the field of disease prevention, malaria is common in areas with serious mosquito breeding and water pollution; in botany, it is found that the growth probability of "orchid plants" is as high as 80% in the area of "semi-humid evergreen broad-leaved forest". These discoveries rely on the application of spatial co-location pattern mining technology.
[0003] Traditional methods for mining spatial co-location patterns mainly rely on Euclidean distance to measure the proximity relationship between spatial instances. However, this method ignores the influence of time factors, resulting in easy misjudgment in some dynamically changing application scenarios. In addition, existing methods usually fail to effectively handle the weight problem of the contribution of different instances to the pattern, often leading to mining results containing a large amount of irrelevant or useless content. Therefore, this paper proposes a method for mining spatial co-location patterns by combining the existence time of instances and the kernel estimation model, aiming to improve the accuracy and practicality of mining.
[0004] In the field of spatial co-location pattern mining, many improved methods are based on the classical Apriori method. Although these methods can effectively process deterministic spatial data, when faced with large-scale data, the computational complexity is too high to be applied to actual scenarios. To improve the computational efficiency, researchers have successively proposed optimization schemes including the partial-join method and the join-less method based on the star neighbor materialization model. However, these methods do not fully consider time constraints in their applications, resulting in their mining results often not meeting the actual needs.
[0005] In contrast, the proposed method for mining spatial co-location patterns based on time weighting and kernel estimation can not only improve the computational efficiency but also accurately identify frequent patterns that meet the actual application requirements, especially suitable for large-scale spatial data and scenarios with significant spatio-temporal changes. In practical applications, especially when dealing with large-scale data, traditional methods often face the problems of huge computational volume and slow processing speed. The method based on time weighting and kernel estimation can effectively reduce the time for traversing table instances by screening spatial instances, improve the efficiency of pattern mining, and capture more accurate spatio-temporal relationships. These patterns can reveal the internal laws in fields such as commercial site selection, urban and rural planning, and environmental monitoring, providing strong support for decision-makers in related fields and thus promoting more precise management and planning. Summary of the Invention
[0006] The purpose of the present invention is to provide a method for mining spatial co-location patterns based on time weighting and kernel estimation, aiming to change the problems such as slow mining efficiency, inaccurate mining results, and generation of a large number of useless or meaningless patterns.
[0007] To achieve the above object, the present invention provides a method for mining spatial co-location patterns based on time weighting and kernel estimation, including the following steps:
[0008] Obtain the time series of instances and summarize them into time series data;
[0009] Preprocess the time series data and transform it into transactional data;
[0010] Traverse the data set, calculate the existence time span of spatial instance pairs, and remove the spatial instance pairs with a time span greater than the maximum time span from the data set;
[0011] Traverse the remaining data set and screen weighted spatial instances;
[0012] Calculate the spatial feature distance weight;
[0013] Use the step-by-step merging method to form high-order candidate patterns and save the instance pairs corresponding to the patterns;
[0014] Prune the candidate patterns;
[0015] Judge the spatial frequency threshold of the spatial co-location pattern;
[0016] Delete redundant spatial co-location patterns.
[0017] In the process of obtaining the time series of instances and summarizing them into time series data, first obtain the data set from a public data set website or other data sources, and combine the location and timestamp of the object to obtain the complete time series data within a certain time period, providing the original data for pattern mining.
[0018] The data annotation process includes the following steps:
[0019] Normalize the data to remove duplicate data and missing data;
[0020] Then, perform data cleaning and processing on the time data of the instances to ensure the accuracy and consistency of the data, and convert the data into transactional data;
[0021] Traverse the transactional data set, and at the same time calculate the existence time span of all spatial instance pairs in the set. Remove the spatial instance pairs with a time span greater than the user-defined maximum time span max_time from the initial set, and retain the remaining spatial instance pairs;
[0022] Traverse the remaining spatial data set, merge the spatial instances belonging to the same spatial feature into a spatial feature set, calculate the existence time span of all spatial instance pairs under each spatial feature, and filter out the spatial instance pairs with an existence time span greater than the user-defined window interval threshold for subsequent mining;
[0023] By merging frequent second-order linear patterns, perform a linear pattern pruning operation on the generated higher-order linear patterns. According to the anti-monotonicity property, if a linear pattern is not a frequent linear pattern, then the higher-order pattern composed of it must also not be frequent;
[0024] Merge the patterns, and the row instances under the patterns also need to be merged accordingly to obtain the row instances of each pattern. At the same time, we introduce a verification pattern to determine whether a linear pattern can form a spatio-temporal frequent pattern;
[0025] To mine spatial frequent co-location patterns, we compare the frequency of the mined co-location patterns with the user-defined frequency threshold. If the frequency of the co-location pattern is greater than the user-defined frequency threshold, then the pattern is a spatio-temporal frequent pattern; otherwise, it does not belong to a frequent pattern.
[0026] Among them, we propose a strategy for pruning candidate patterns, which can mine all frequent co-location patterns faster, namely, second-order linear pattern pruning, candidate linear pattern pruning, and candidate loop pattern pruning strategies. According to the anti-monotonicity property, if a pattern is frequent, then all second-order patterns that make up the pattern must be frequent. During the pattern merging process, if it is found that the linear pattern that makes up the pattern is not frequent, then the linear pattern is directly deleted because any subsequent loop pattern containing the linear pattern must not be frequent and will not be merged. When the frequency of the table instances in the formed pattern is less than the user-defined threshold, then the pattern is also not frequent and should be pruned.
[0027] Among them, in order to fully ensure that the patterns in the result set are correct and there are no meaningless and redundant patterns, we screen the results and remove duplicate patterns.
[0028] The present invention provides a method for mining spatial co-location patterns based on time weighting and kernel estimation. By obtaining data from public data sets or other data sources, combining the timestamps and locations of objects, complete time series data is constructed. Through data cleaning and preprocessing, the data is converted into transactional data, and frequent second-order patterns are screened out. Then, a step-by-step merging method is used to generate higher-order candidate patterns, and pattern pruning is performed to ensure that only patterns that meet the frequency requirements are retained. Finally, through a verification mechanism and screening operations, it is ensured that the mined patterns have high accuracy and comprehensiveness, thereby obtaining the final frequent patterns. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0030] Figure 1 is a schematic flow chart of a method for mining spatial co-location patterns based on time weighting and kernel estimation of the present invention.
[0031] Figure 2 is an introduction to the pseudocode of the present invention.
[0032] Figure 3 is a schematic diagram of the join operation of the method for mining spatial co-location patterns based on time weighting and kernel estimation of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] The following details the embodiments of the present invention. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.
[0034] Please refer to Figure 1 , the present invention proposes a method for mining spatial co-location patterns based on time weighting and kernel estimation, including the following steps:
[0035] S1: Obtain the time series of the instances and summarize it into time series data;
[0036] S2: Preprocess the time series data and convert it into transactional data;
[0037] S3: Traverse the dataset, calculate the existence time span of spatial instance pairs, and remove the spatial instance pairs with a time span greater than the maximum time span from the dataset;
[0038] S4: Traverse the remaining dataset and filter weighted spatial instances;
[0039] S5: Calculate the spatial feature distance weights;
[0040] S6: Prune the candidate patterns;
[0041] S7: Judge the spatial frequency threshold of spatial co-location patterns;
[0042] S8: Delete redundant spatial co-location patterns;
[0043] In the process of obtaining the time series of instances and generalizing them into time series data, a dataset is obtained from a public dataset website or other data sources. Combining the location and timestamp of the object, complete time series data within a certain time period is obtained, providing raw data for pattern mining.
[0044] Clean the raw data by deleting missing data and abnormal data.
[0045] The input data required by this method is a transactional spatial dataset, the spatial features in the dataset, and the frequency threshold, distance threshold, window interval threshold, maximum time span threshold, and time weight defined by the user.
[0046] Traverse the instances in the dataset and simultaneously identify and filter out consecutive second-order linear co-location patterns. For each pattern, we check its frequency and retain those second-order linear co-location patterns with a frequency greater than or equal to the minimum threshold defined by the user.
[0047] By merging frequent second-order linear patterns, higher-order linear patterns are generated. In this process, we use linear pattern pruning operations to improve the method efficiency. According to the principle of anti-monotonicity, if a certain linear pattern itself is not frequent, then the pattern it forms must not be frequent either. Therefore, during the merging process, we will delete those linear patterns that do not meet the frequency requirements, thus avoiding unnecessary calculations for unimportant patterns.
[0048] We also perform corresponding merging operations on the row instances in the merged patterns to obtain the row instances of each pattern. In addition, we introduce a verification pattern to judge whether a certain linear pattern can form a frequent pattern. The construction of the verification pattern is based on the combination of the last feature and the first feature of the frequent linear pattern, which helps to ensure the rationality and effectiveness of pattern merging.
[0049] Traverse all spatial instances, and at the same time screen the frequency of consecutive second-order linear co-location patterns, retaining second-order linear co-location patterns that are greater than or equal to the user-defined minimum threshold.
[0050] By merging frequent second-order linear patterns, perform a linear pattern pruning operation on the generated higher-order linear patterns. According to the anti-monotonicity characteristic, if a linear pattern is not a frequent linear pattern, then the patterns composed of it must also not be frequent.
[0051] To more comprehensively mine frequent co-location patterns, we compare the frequency of the mined co-location patterns with the user-defined minimum threshold. If the frequency is greater than or equal to the threshold, retain the pattern; otherwise, remove it to ensure the accuracy and practicality of the final result.
[0052] Finally, we perform a further screening operation on the patterns in the result set to remove the redundancy of the results.
[0053] Furthermore, the present invention is further described in detail in conjunction with specific embodiments and attached pseudo-code:
[0054] Please refer to Figure 2 , the specific steps of the specific process of the algorithm are as follows:
[0055] Algorithm 1 shows the overall pseudo-code for mining all spatial co-location patterns. Among them, S represents the spatial data set, F represents the spatial feature set, d represents the distance threshold, min_pre represents the frequent threshold, min_window represents the window interval threshold, max_time represents the maximum time span threshold, and W represents the time weight.
[0056] (Step 1) Algorithm 1 shows the pseudo-code for mining spatial co-location patterns. First, take the spatial data set as the initial set, calculate the existence time span of all spatial instance pairs in the set, and remove the spatial instance pairs with an existence time span greater than the maximum time span max_time from the initial set.
[0057] (Steps 2 - 11) After the spatial data set is initialized, merge the spatial instances belonging to the same spatial feature into a spatial feature set, calculate the existence time span of all spatial instance pairs under each spatial feature, and screen out the spatial instance pairs with an existence time span greater than the window interval threshold min_window fi . And add the spatial instances in the spatial instance pairs to the empty set CL, and repeat step 2 until no instance is added to CL.
[0058] (Steps 12 - 15) Calculate the spatio-temporal proximity relationship of the spatial instance pairs in the spatial data set. If the instance pair contains the spatial instances in the set CL, then according to the following formula min_window fi <|tA -t B |≤max_time to determine whether the instance pair satisfies the spatio-temporal proximity relationship; if the instance pair does not contain the spatial instance in the set CL, then if the Euclidean distance of the instance pair is less than or equal to the distance threshold d, the instance pair satisfies the spatio-temporal proximity relationship.
[0059] (Steps 16 - 18) Next, calculate the spatial feature distance weights.
[0060] (Steps 19 - 23) Generate spatial co-location patterns based on the join method.
[0061] Furthermore,
[0062] An example of the join method is as Figure 3 shown. Combining with the schematic diagram, illustrate the specific steps of the join operation and how the pattern {A, B, C, D, E} is generated:
[0063] Step 1: Generate 1 - order patterns. All single objects are 1 - order patterns, namely {A}, {B}, {C}, {D}, {E}.
[0064] Step 2: Candidate generation. Join all 1 - order patterns, namely {A, B}, {A, C}, {A, D}, {A, E}, {B, C}, {B, D}, {B, E}, {C, D}, {C, E}, {D, E}, and verify whether there are spatially proximate instances for each pair of objects. Only retain the 2 - order frequent patterns that satisfy the proximity relationship. According to the following formula and judge the pattern frequency and compare it with the user - set frequency threshold to determine the frequent patterns.
[0065] Step 3: Generate 3 - order patterns. First, join the 2 - order patterns that share 1 common item, and then check whether all sub - patterns are frequent patterns. For example, {A, B, C} is generated by joining {A, B} and {B, C}.
[0066] Step 4: Repeat the above join operation. According to the anti - monotonicity characteristic, if any pattern is not a frequent linear pattern, then the higher - order pattern {A, B, C, D, E} composed of it must also not be frequent.
[0067] Step 5: Join the 4 - order patterns that share 3 common items, and verify the frequency of the pattern {A, B, C, D, E}. If its frequency is greater than or equal to the threshold, then the pattern {A, B, C, D, E} is a frequent co - location pattern.
[0068] The above-disclosed is only a preferred embodiment of the present invention. Of course, the scope of the rights of the present invention cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.
Claims
1. A spatial co-location pattern mining method based on time weighting and kernel estimation, characterized in that: The following steps are involved: Get the time series of the instance and summarize it into time series data; Preprocess time series data and transform it into transactional data; Traverse the trajectory data set and filter out all frequent second-order linear co-location patterns; Traverse the data set, calculate the existence time span of the space instance pairs, and remove the space instance pairs that are greater than the maximum time span from the data set; Traverse the remaining data sets and filter weighted space instances; Calculate spatial feature distance weights; Use the step-by-step merging method to form high-order candidate patterns and save instance pairs under corresponding patterns; Prune candidate patterns.
2. The spatial co-location pattern mining method based on time weighting and kernel estimation according to claim 1, characterized in that: In the process of obtaining the time series of instances and summarizing them into time series data, we first obtain the data set from the public data set website or other data sources, and combine the location and timestamp of the object to obtain the complete time series data within a certain period of time, providing raw data for pattern mining.
3. The spatial co-location pattern mining method based on time weighting and kernel estimation according to claim 1, characterized in that: Normalize the data, remove duplicate data, remove missing data, clean and process the time data of the instance to ensure the accuracy and consistency of the data, and convert the data into transactional data.
4. The spatial co-location pattern mining method based on time weighting and kernel estimation according to claim 1, characterized in that: The transactional dataset is traversed, and the existence time span of all spatial instance pairs in the collection is calculated. The spatial instance pairs that are greater than the user-defined maximum time span max_time are removed from the initial collection, and the remaining spatial instance pairs are retained.
5. The spatial co-location pattern mining method based on time weighting and kernel estimation according to claim 1, characterized in that: Traverse the remaining spatial data sets, merge the spatial instances belonging to the same spatial feature into a spatial feature set, calculate the existence time span of all spatial instance pairs under each spatial feature, and filter the spatial instance pairs whose existence time span is greater than the user-defined window interval threshold for subsequent mining.
6. The spatial co-location pattern mining method based on time weighting and kernel estimation according to claim 1, characterized in that: In the process of generating high-order linear patterns, higher-order patterns are gradually constructed by merging frequent second-order linear patterns. In order to improve efficiency, the generated high-order patterns are pruned. According to the anti-monotonicity principle, if a second-order linear pattern is not a frequent pattern, then any high-order pattern built based on this pattern must not be frequent. Therefore, in the process of pattern merging, once a linear pattern is found to not meet the frequent condition, it is directly eliminated to avoid further meaningless calculation and analysis of the generated high-order pattern. Effectively reduce the amount of calculation and ensure the accuracy of the final result.
7. The spatial co-location pattern mining method based on time weighting and kernel estimation according to claim 1, characterized in that: In the process of pattern merging, for each pattern, not only the pattern itself is merged, but also the row instances in the pattern are merged accordingly, so as to obtain the row instances corresponding to each merged pattern. In order to ensure that the merged pattern is valid, a mechanism of pattern verification is introduced to determine whether each linear pattern can constitute a spatiotemporal frequent pattern. The verification pattern helps confirm whether the pattern meets the requirements of the frequent pattern by checking the frequency of the linear pattern and its spatiotemporal proximity, thereby ensuring that the final mined pattern has high reliability and accuracy.
8. The spatial co-location pattern mining method based on time weighting and kernel estimation according to claim 1, characterized in that: In order to mine spatial frequent co-location patterns, we compare the frequency of the mined co-location patterns with the user-defined frequent threshold. If the frequency of the co-location pattern is greater than the user-defined frequent threshold, then the pattern is a spatiotemporal frequent pattern, otherwise it is not a frequent pattern.
9. The spatial co-location pattern mining method based on time weighting and kernel estimation according to claim 1, characterized in that: In order to ensure that the patterns in the result set are accurate and meaningful, the mining results are screened and meaningless or redundant patterns are removed, thus ensuring the simplicity and effectiveness of the final results.