Business Decision-making Method Based on Region Frequent Co-location Patterns

Through the two-stage mining method based on the huge instance cluster and multi-density grid clustering, the problem of lack of semantic information and low efficiency in regional spatial homogeneous mode mining is solved, and more accurate business model recognition and site selection decision support is achieved.

CN116091095BActive Publication Date: 2025-07-18YUNNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211596295.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-12
Publication Date
2025-07-18
Estimated Expiration
2042-12-12

AI Technical Summary

Technical Problem

The existing regional spatial homogeneous mode mining methods lack semantic information in regional division, which easily produces sub-regions that are not geographically adjacent but semantic similar, and have low mining efficiency, making it difficult to provide effective guidance in commercial site selection.

Method used

The two-stage mining method based on a huge instance cluster is used to combine multi-density grid clustering, and the participation of commercial homologous mode is calculated through the Bron–Kerbosch algorithm and CL-hash structure. The study area is subdivided into multiple subregions with multi-density grid clustering, and subregions with the same semantic information are merged to provide business decision support.

Benefits of technology

It can quickly identify frequent business isotopic patterns in areas, and the divided sub-regions are closer to the distribution of business models in reality, providing scientific business model selection and site selection guidance, and reducing entrepreneurial risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116091095B_ABST
    Figure CN116091095B_ABST
Patent Text Reader

Abstract

The present invention discloses a business decision-making method based on regional frequent co-location patterns, which mainly includes: determining business categories to obtain a spatial business category set and a business instance set; calculating the proximity relationship between business instances to obtain a spatial instance neighbor table, and calculating to obtain a maximum instance clique; calculating the participation degrees of all business co-location patterns, and dividing them into global frequent business co-location patterns and regional candidate business co-location patterns; dividing the entire region into multiple sub-regions; calculating the participation degrees of the regional candidate business co-location patterns, and those that meet the minimum participation degree threshold are regional frequent business co-location patterns, and the sub-regions where they are located are frequent regions; finding out the regional frequent business co-location patterns of the merged sub-regions; and giving corresponding decision results according to requirements. It solves the problem that the results of existing regional frequent co-location pattern mining lack semantic information and practical application value, and assists users in selecting business models and site selection, reducing the entrepreneurial risk and increasing the success probability of entrepreneurship.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of regional spatial data mining and decision support, and relates to a business decision-making method based on regional frequent co-location patterns. Background Art

[0002] Different business entities have extremely complex interrelationships in the economic distribution system of an entire city or township. Some business categories often have great connections and derivative relationships. For example, flower shops and pharmacies often appear together with hospitals, and bars and KTVs also usually appear within the same area. Such two or more mutually related business entities are called a group of business co-location patterns. In addition to this mutual connection between business co-location patterns, the influence of different regions on business co-location patterns is also very obvious. For example, bars and KTVs often appear near the city center or luxurious business districts, while they are not common in other areas. Reasonably considering these factors and through scientific analysis and actual research have crucial guiding significance for the choice of business models and business locations when users start their own businesses. As the saying goes, "Seven points depend on location, and three points depend on operation." The location selected by merchants directly affects the actual operating efficiency of enterprises and is related to the success or failure of enterprises. Especially in the current epidemic situation, many physical industries are facing the risk of bankruptcy and closure. At this time, choosing the correct business model and the correct business operating location is undoubtedly an important decision related to the survival of enterprises. With the increasing improvement of computer infrastructure construction and the support of massive data, according to the objective laws within the location selection, spatial data mining technology has been widely used in user business location selection, providing strong support for this strategic decision of business location selection. Using spatial mining technology can quickly help merchants find valuable information, contribute to more accurately discovering business opportunities, formulating development plans, and thus taking the lead in seizing the market.

[0003] Spatial co-location pattern mining is a very important research direction in the field of spatial data mining. A spatial co-location pattern refers to a set of spatial features, and the features in the set frequently appear together in space. For example, many communities around schools, pharmacies around hospitals, parking lots around shopping malls, etc. Spatial co-location pattern mining is to identify and discover these patterns from a large amount of spatial datasets. The discovery of spatial co-location patterns has also been widely applied in fields such as environmental protection, public safety, urban road planning, user behavior analysis, etc. However, due to the heterogeneity of spatial data distribution, some spatial co-location patterns may only be distributed in some specific regions. For example, in the field of public health, the overall prevalence rate of malignant tumors in a certain area does not seem to be high, but in areas with heavy industrial pollution, the prevalence rate of malignant tumors is particularly high. In the field of urban safety, the high crime rate in the city is mainly concentrated near the bars in the city center, and there is no such phenomenon in other regions. This implicit regional correlation is called the regional frequent spatial co-location pattern. The regional frequent spatial co-location pattern represents a subset of spatial features that frequently occur together in some regions (i.e., sub-regions) of the study area.

[0004] In recent years, many scholars have conducted a series of in-depth studies on the theories and methods of spatial co-location pattern mining and achieved fruitful results. However, most of the research focuses on mining global frequent co-location patterns, and only a very small number of studies are on regional frequent co-location pattern mining. Moreover, the existing research on regional frequent co-location patterns mostly aims to find objective laws and lacks the interpretation of the semantic information and practical implications of the mining results. Huang Y, Shekhar S, Xiong H. Discovering colocation patterns from spatial data sets: a general approach[J]. IEEE Transactions on Knowledge & Data Engineering, 2004, 16(12): 1472-1485 first proposed a general mining method for co-location patterns and gave related definitions. Due to the heterogeneity of spatial data distribution, some spatial co-location patterns may be distributed in some sub-regions. Zeng L, Wang L, Zeng Y, et al. Discovering Spatial Co-location Patterns with Dominant Influencing Features in Anomalous Regions[C] / / International Conference on Database Systems for Advanced Applications. Springer, Cham, 2021: 267-282 found that in the field of public health, the overall prevalence of malignant tumors in a certain area does not seem to be high, but in areas with heavy regional industrial pollution sources, the prevalence of malignant tumors is particularly high. P Mohan, S Shekhar, J A Shine, J P Rogers, N Wayant. A neighborhood graph based approach to regional co-location pattern discovery: a summary of results. ACM, https: / / doi.org / 10.1145 / 2093973.2093991 found that in the field of urban security, the high crime rate in the city is mainly concentrated near the bars in the city center, and this phenomenon does not exist in other areas.

[0005] Existing methods for discovering regional spatial co-location patterns usually first identify globally frequent co-location patterns and use globally infrequent co-location patterns as candidate regional co-location patterns. Then, all possible sub-regions within the study area are determined, and the frequency of the co-location patterns in each sub-region is evaluated. For the division of regions, existing research mainly falls into two aspects: hard-dividing the global study area into sub-regions based on pre-specified rules, and clustering-based methods that adaptively aggregate sub-regions according to the distribution of spatial instances.

[0006] Eick C F, Parmar R, Ding W, et al. Finding regional co-location patterns for sets of continuous variables in spatial datasets. Proceedings of the 16th ACM SIGSPATIAL international conference on Advances in geographic information proposed a quadtree-based method to divide the global region into four equal parts and then used a recursive method to further subdivide the divided regions until the desired effect was achieved. To reduce the subjectivity in determining the partitioning scheme, F. Qian, K. Chiew, Q. He and H. Huang, Mining regional co-location patterns with kNNG. J. Intell. Inf. Syst., doi:10.1007 / s10844-013-0280-5 developed a K-nearest neighbor-based partitioning method to divide the study area into several homogeneous sub-regions, in each of which the weights of the edges in the K-nearest neighbor graph were slightly different. However, this way of partitioning sub-regions based on fixed rules also has some obvious drawbacks. P. Mohan, S. Shekhar, J. A. Shine, J. P. Rogers, Z. Jiang and N. Wayant, A neighborhood graph-based approach to regional co-location pattern discovery: a summary of results, Proceedings of the 19th ACM SIGSPATIAL International Conference on Advances in Geographic Information Systems, New York, NY, USA, doi:10.1145 / 2093973.2093991 found that the actual location and shape of the region where the regional co-location pattern is located cannot be accurately determined by the specified spatial partitioning scheme, and the connection between co-location patterns may be wrongly cut off.

[0007] Traditional clustering methods require prior knowledge, such as K-Means which requires the number of clusters, and DBSCAN which requires determining the density. At the same time, the complexity of these clustering algorithms is relatively high. For large spatial datasets, they require a large amount of resources. Ding W, Eick C F, Yuan X, et al. A framework for regional association rule mining and scoping in spatial datasets [J]. Geoinformatica, 2011, 15(1): 1-28 proposed a grid-based clustering method for finding interesting feature hotspots, which can greatly accelerate the clustering speed and reduce spatial consumption. M. Deng, Q. Liu, T. Cheng, Y. Shi. An adaptive spatial clustering algorithm based on delaunay triangulation. Comput. Environ. Urban Syst., doi:10.1016 / j.compenvurbsys.2011.02.003 proposed an adaptive spatial clustering algorithm based on Delaunay triangulation, which can adaptively generate different sub-regions. And their team proposed a multi-level co-location pattern mining method in 2017: M. Deng, J. Cai, Q. Liu, Z. He, J. Tang. Multi-level method for discovery of regional co-location patterns. International Journal of Geographical Information Science first proposed using globally infrequent commercial co-location patterns as regional candidate co-location patterns. To improve the generality of the algorithm and reduce the algorithm's dependence on users' prior knowledge, Cai J, Liu Q, Deng M, et al. Adaptive detection of statistically significant regional spatial co-location patterns [J]. Computers, Environment and Urban Systems, 2018, 68: 53-63 improved the multi-level method and proposed an adaptive pattern clustering method that does not require users to specify parameters.

[0008] Although many scholars have conducted extensive research on the mining methods of regional co-location patterns over the years, there are still some problems to be solved: (1) The regions divided by existing regional division methods lack certain semantic information and it is difficult to combine with real-world scenarios to obtain some guiding opinions. (2) The finally divided sub-regions by existing methods may be too many, and there may be sub-regions with similar semantics but not geographically adjacent. (3) Whether it is global co-location mining or regional co-location pattern mining, if the candidate-generation and mining method is used step by step for co-location pattern mining, the efficiency will be very low. (4) The setting of the pattern participation threshold plays a crucial role in the mining results, but it is difficult for users lacking relevant prior knowledge to select a reasonable threshold, and once the participation threshold algorithm is modified in existing methods, the algorithm needs to be restarted, and the algorithm efficiency is not high. Summary of the Invention

[0009] To achieve the above object, the present invention provides a business decision-making method based on regional frequent co-location patterns, which solves the problems of the regions divided by the prior art lacking semantic information, being difficult to apply to real-world scenarios, being prone to having sub-regions with similar semantics but not geographically adjacent, and low mining efficiency, provides scientific and reasonable decision-making assistance for business users, reduces the entrepreneurial risk, and improves the success rate.

[0010] The technical solution adopted by the present invention is a business decision-making method based on regional frequent co-location patterns, including the following steps:

[0011] Step S1, determine the global mining region and all business categories of interest within the mining range, and obtain the spatial business category set F = {f1, f2,..., f n}, where n represents the number of business categories; each business category consists of multiple business instances, and the business instance set of a business category is represented as S = {S i}, 1 ≤ i ≤ n;

[0012] Step S2, calculate the proximity relationship between business instances based on the given distance threshold d, obtain the spatial instance neighbor table of each business instance in the business instance set S, and then use the partitioned Bron–Kerbosch algorithm to calculate all maximal instance cliques;

[0013] Step S3, according to all the maximal instance cliques obtained in S2, use the two-stage mining method based on maximal instance cliques to calculate the participation degrees of all business co-location patterns, classify the business co-location patterns that meet the minimum participation threshold as globally frequent business co-location patterns, and regard the patterns that do not meet the threshold as regional candidate business co-location patterns;

[0014] Step S4: Convert all the maximum instance cliques obtained in Step S2 into corresponding geometric center points, and then project the geometric center points onto an N×N grid; set multiple density thresholds, and perform multi-density grid clustering on the geometric center points to divide the entire region into multiple sub-regions;

[0015] Step S5: Treat each sub-region as an independent research area. In each sub-region, use a two-stage mining method to calculate the participation of all patterns in the set of candidate commercial co-location patterns in the region. The commercial co-location patterns that meet the minimum participation threshold are the region-frequent commercial co-location patterns; the sub-region where the region-frequent commercial co-location pattern is located is the corresponding frequent region; then merge the sub-regions with the same or similar semantic information into one sub-region and find the region-frequent commercial co-location patterns of the merged region;

[0016] Step S6: Give corresponding decision results according to different requirements.

[0017] Further, in Step S1, the global mining region is a spatial coordinate system, and the spatial business categories are represented in the form of a business category list {A, B,..., Z}, where each letter represents a business category; the commercial instance O i is represented in the form of a triple <belonging business category, instance number, spatial position>, where the instance number represents the serial number of the commercial instance among all instances of the belonging business category, and the spatial position is composed of a two-dimensional array [x, y], where x and y respectively represent the abscissa and ordinate of the commercial instance in the global spatial coordinate system.

[0018] Further, the specific implementation process in Step S2 is as follows:

[0019] Step S21: Given a proximity relation spatial proximity distance threshold d, when the Euclidean distance between two commercial instances is less than the threshold d, these two commercial instances have a neighbor relationship; sort the business categories in alphabetical order, and sort the commercial instances included in each business category by instance number; calculate the Euclidean distance between any two commercial instances O i and O j in the commercial instance set S according to the spatial positions of the commercial instances, where O i and O j belong to different commercial instance sets S; two commercial instances that meet the condition of being less than the distance threshold d are neighbors to each other, and a commercial instance neighbor table corresponding to the commercial instance set S = {s1, s2,..., s n} is obtained;

[0020] Step S22: Perform partition operation on all business instances according to the business instance neighbor table. The partition steps are as follows: Perform partition operation on all business instances. First, select an unvisited business instance P from the instance set S, put the business instance P and its neighbor business instances into the set N, then take the set N and the business instances having a proximity relationship with N as a partition, and finally update the instance set S;

[0021] Step S23: Pass the partitioned business instance sets into the Bron–Kerbosch algorithm in sequence to obtain all maximal instance cliques.

[0022] Furthermore, the specific steps of the Bron–Kerbosch algorithm in the step S23 are as follows: First, initialize three disjoint sets Q, R, and X. Among them, Q represents the set of generated maximal instance cliques, R represents the set of instances to be processed, and X represents the set of instances that have been processed. The initialized Q and X are empty sets, and R is obtained from the partitioned business instance set; First, select the vertex u with the largest number of adjacent vertices from P as the pivot element, put each non-neighbor vertex v of u in R into the set Q, and remove the vertices that are not adjacent to v from R and X. Then, recursively set the sets Q, R, and X, and move the selected vertex u from R to X; Repeat the above steps until the union of R and X is an empty set. At this time, Q is the finally obtained maximal instance clique.

[0023] Furthermore, the step S3 is specifically as follows:

[0024] Step S31: Store all maximal instance cliques through a Hash table structure and find the participating examples of business co-location patterns;

[0025] Step S32: Based on the CL-hash structure, calculate the participation degree PI(c) of all business co-location patterns through a two-stage mining method based on maximal instance cliques, and first find the globally frequent business co-location patterns according to the given minimum frequency threshold min_prev. The globally frequent business co-location patterns will no longer participate in the mining of regionally frequent business co-location patterns. The patterns that do not meet the threshold are used as regionally candidate business co-location patterns. The regionally candidate business co-location patterns will participate in the subsequent mining of regionally frequent business co-location patterns, and finally obtain the frequent regional business co-location patterns;

[0026] In the step S31, the storage method of the Hash table structure is a double Hash structure <business co-location pattern, <business category, instance number>>, called CL-hash, where the business co-location pattern is represented as k is the number of business categories included in the business co-location pattern c, called the order of the business co-location pattern c; A business category f i All non-repeated business instances in the table instances of the business co-location pattern c constitute the business category fi An example of participation in the commercial co-location pattern c, where 1 ≤ i ≤ k;

[0027] The table instance is all row instances of the commercial co-location pattern c;

[0028] The row instance means that when a spatial set cl contains all commercial categories in the commercial co-location pattern c, and no subset of cl contains all commercial categories in c, then the spatial set cl is called a row instance of the commercial co-location pattern c.

[0029] Furthermore, in step S32, the participation degree PI(c) of the commercial co-location pattern c is the minimum value of the participation rates of all commercial categories f i (1 ≤ i ≤ k), that is:

[0030]

[0031] Where,

[0032]

[0033] Where, PR(c, f i ) is the participation rate of the commercial category f i ; |∏ fi (T(c))| represents the number of participation examples of the commercial category f i in the co-location pattern c; |T{f i}| represents the total number of commercial instances of the commercial category f i ; Π is the projection operation of the relationship, indicating that the same commercial instance can only be counted once to remove redundancy; T(c) represents the table instance of the pattern c;

[0034] The specific process of the two-stage mining method based on the maximum instance clique is as follows:

[0035] In the first stage, the PI(c) values of all clique candidate business models are directly calculated by CL-hash, and the commercial co-location patterns with PI(c) values greater than or equal to min_pre are used as global frequent commercial co-location patterns, and those less than min_pre are used as regional candidate business models;

[0036] In the second stage, for the table instances of all remaining business models that are not in the clique candidate business models, they can be obtained by collecting all supersets containing the remaining business models, and the PI(c) values of all remaining business models are calculated according to the participation examples in the table instances, and those with PI(c) values greater than or equal to min_pre are used as global frequent commercial co-location patterns, and those less than min_pre are used as regional candidate business models.

[0037] Furthermore, the specific implementation process of step S4 is as follows:

[0038] Step S41: Convert all the maximum instance cliques obtained in Step S2 into corresponding geometric center points. After the conversion is completed, a new spatial data set is formed. In the new spatial data set, construct an N×N grid, where the value of N is not fixed and is determined considering the scale of the data set. The larger the data set scale, the larger the value of N. Project the converted spatial data set onto the grid plane;

[0039] Step S42: Set multiple density thresholds, and perform multi-density grid clustering on the converted spatial data set from large to small density thresholds to subdivide the entire region into multiple sub-regions.

[0040] Furthermore, in the above-mentioned Step S42, the specific process of multi-density grid clustering is as follows:

[0041] (1) Set a density threshold list density, where the value of density is defaulted to the value obtained by evenly mapping the scale of the spatial data set to each grid;

[0042] (2) Sort according to the density thresholds from large to small, and perform grid density clustering on the data set. If the clusters obtained by clustering meet the specified size, the clusters obtained by clustering are regarded as qualified sub-regions. Otherwise, it is considered that they do not meet the conditions to become the smallest regions, and the data they contain is regarded as noise points;

[0043] (3) Decrease the density threshold, and put the noise points into the next round of grid density clustering;

[0044] (4) Until the density threshold list is traversed to the end, all sub-regions are obtained.

[0045] Furthermore, in the above-mentioned Step S5, when merging sub-regions, calculate the intersection over union ratio CIOU(Ri, Rj) between the frequent commercial co-location patterns of two regions to determine whether the two regions are similar, that is:

[0046]

[0047] where {Ric} represents the set of frequent commercial co-location patterns in region Ri. If CIOU(Rj, Rj) is greater than the minimum similarity threshold min_simi of the given two sub-regions, then merge these two sub-regions, and the frequent commercial co-location pattern of the merged new sub-region takes the intersection of the frequent commercial co-location patterns of these two sub-regions.

[0048] Furthermore, the specific implementation process of the above-mentioned Step S6 is as follows:

[0049] Step S61: If there is no special specification, display all mining results;

[0050] Step S62: If the user specifies the specific location and scope for starting a business, then provide the frequent business co-location patterns within the specific location and scope.

[0051] Step S63: If the business category or business model to be engaged in is given, then search the mining results of all sub-regions, and provide the frequent regions that include the corresponding business category or business model.

[0052] The beneficial effects of the present invention are

[0053] 1) It can intuitively reflect the distribution of each business model, and the method can quickly obtain new mining results when changing the setting of the frequency threshold.

[0054] 2) The entire research area can be divided into sub-regions at different levels according to the distribution density of the maximum instance clusters. Compared with the sub-regions obtained by the clustering algorithm, the multi-level sub-regions can better reflect the distribution of business models in reality.

[0055] 3) It can correctly identify the regional frequent business co-location patterns, and the divided sub-regions can also be closer to the pattern distribution in reality. It can scientifically and effectively provide decision-making support for the user's business model selection or business location selection. Description of the Drawings

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0057] Figure 1 is the flowchart of the business decision-making method based on regional frequent co-location patterns in the embodiments of the present invention;

[0058] Figure 2 is the schematic diagram of the spatial business instance in the embodiments of the present invention.

[0059] Figure 3 is the schematic diagram of the storage structure CL-hash of the maximum instance cluster in the embodiments of the present invention.

[0060] Figure 4 is the schematic diagram of multi-density clustering based on the maximum instance cluster in the embodiments of the present invention, where (a) is to convert the maximum instance cluster into a geometric center to form a new data set, and (b) is the expected effect diagram of multi-density clustering based on the maximum instance cluster.

[0061] Figure 5 is the schematic diagram of the regional co-location pattern in the embodiments of the present invention.

[0062] Figure 6 These are two synthetic datasets of the embodiments of the present invention. Among them, (a) is the spatial distribution diagram of the synthetic dataset with a regional business model, and (b) is the spatial distribution diagram of the synthetic dataset with a global distribution.

[0063] Figure 7 This shows the influence of the embodiments of the present invention on synthetic datasets of different distribution categories under different distance thresholds. Among them, (a) is the bar chart of the influence on the number of generated maximum instance cliques, and (b) is the line chart of the influence on the running time.

[0064] Figure 8 These are multiple synthetic datasets of the embodiments of the present invention. Among them, (a) is the synthetic dataset containing multiple distribution densities, (b) is the synthetic dataset with an overlapping area of the spatial distribution of business categories, (c) is the synthetic dataset of the distribution of 4 groups of business categories with different degrees of sparsity, and (d) is the synthetic dataset with a global distribution.

[0065] Figure 9 This is the partitioning result diagram of the embodiments of the present invention on multiple synthetic datasets. Among them, (a) is Figure 8 the partitioning result diagram of (a) therein, (b) is Figure 8 the partitioning result diagram of (b) therein, (c) is Figure 8 the partitioning result diagram of (c) therein, and (d) is Figure 8 the partitioning result diagram of (d) therein.

[0066] Figure 10 These are two real datasets of the embodiments of the present invention. Among them, (a) is the spatial distribution diagram of the Beijing vegetation distribution dataset, and (b) is the spatial distribution diagram of the Shenzhen City Point of Interest (POI) distribution dataset.

[0067] Figure 11 This is the process diagram of multi-density grid clustering on the Beijing vegetation distribution dataset in the embodiments of the present invention. Among them, (a) is the clustering process diagram, and (b) is the clustering result diagram.

[0068] Figure 12 This is the bar chart of the running time of the embodiments of the present invention under different numbers of spatial business instances.

[0069] Figure 13 This shows the influence of different numbers of spatial business categories on the embodiments of the present invention. Among them, (a) is the bar chart of the influence on the running time, and (b) is the line chart of the influence on the number of generated regional frequent co-location patterns.

[0070] Figure 14 This is the line chart of the running time of the embodiments of the present invention under different minimum frequency thresholds. Detailed implementation manners

[0071] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0072] like Figure 1 As shown, the present invention provides a business decision support based on regional frequent co-location patterns, and the specific steps are as follows:

[0073] Step S1: First determine the global research area, then determine the category information of the commercial entities of interest in the global research area, encode each commercial entity category in uppercase letters, and obtain the spatial commercial category set F = {f1, f2, ..., f n}, n represents the number of business categories. Each business category contains a set of business entities distributed in different locations in space. All business entities constitute a space business instance, referred to as a business instance. The business instance set S = {s i}, 1≤i≤n, n represents the number of business categories. Each business entity represents a business instance of the corresponding business category. In order to represent and distinguish different entities, the business instances are numbered as the instance number of the business instance, denoted as o i . Business Examples i Represents any business entity to be mined, business instance o i It is described by the triple <business category, instance number, spatial location>. The instance number indicates the serial number of the instance among all business instances in the business category, and the spatial location refers to the spatial location of the business instance to be mined. Figure 2 As shown in the example given, four business categories are described, and the spatial business category set F = {A, B, C, D}. Among them, business category A has five business instances, and the corresponding business instance set S = {A1, A2, A3, A4, A5}. Taking business instance A1 as an example, it is represented in the data set as<A,1,[6,6]> , where [6,6] represents its spatial position coordinates.

[0074] Step S2: Arrange the business categories in alphabetical order, and sort the business instance set corresponding to each business category by instance number. Given a spatial proximity distance threshold d, calculate the spatial distance between business instances of different business categories based on the spatial position of the business instance. Two business instances that meet the distance threshold d are neighbors of each other, and the business instance set S = {s1, s2, ..., s n} corresponding commercial instance neighbor table. Then, all commercial instance maximal instance cliques are obtained by using the partitioned Bron–Kerbosch algorithm based on the commercial instance neighbor table. The specific steps are as follows:

[0075] Step S21: Given a spatial proximity relation distance threshold d, sort the commercial category set in alphabetical order, sort the commercial instances of each commercial space category according to the instance number, and calculate the Euclidean distance between any commercial instance in the commercial instance set S = {s1, s2,..., s n} and the commercial instances of other spatial commercial categories. If the Euclidean distance between any two participating commercial instances is less than the user-given spatial proximity distance threshold d, they satisfy the proximity relation to each other, and then the neighbors of each commercial instance are determined, obtaining the commercial instance neighbor table corresponding to the commercial instance set S = {s1, s2,..., s n}. For example, Figure 2 in, if there is a connection between two commercial instances, they are neighbors. For example, the neighbors of A4 are B1 and C3.

[0076] Step S22: The generation of maximal instance cliques is an NP-hard problem. To improve the algorithm efficiency, all commercial instances are partitioned and calculated according to the commercial instance neighbor table using the partition algorithm. The partition algorithm is described as follows:

[0077] 1) Select an unvisited node P from the commercial instance set S, and put the node P and its neighbor nodes into the set N.

[0078] 2) Then, take the set N and the instance nodes having a proximity relation with N as a partition.

[0079] 3) Finally, update the instance set S (S = S - N).

[0080] Step S23: Pass the partitioned instance sets into the Bron–Kerbosch algorithm in sequence to obtain all maximal instance cliques. The Bron–Kerbosch algorithm was proposed by Bron and Kerbosch in 1973 and is also the most widely used maximal instance clique mining algorithm.

[0081] The specific process is shown in Algorithm 1:

[0082] Algorithm 1: Bron–Kerbosch Algorithm with Pivot

[0083]

[0084] The algorithm process is explained as follows: First, if R ∪ X is empty, add Q to the set of maximum instance cliques (step 1), and select the vertex u with the largest number of adjacent vertices as the pivot element (step 2). For the non-neighbor vertex v of u in R, perform the following operations: put v into the set Q, and remove the vertices that are not adjacent to v from R and X, and then recursively set Q, R, and X (steps 3 and 4). Backtrack to the previous step, and move the selected vertex u from R to X (steps 5 and 6).

[0085] Step S3: According to all the obtained maximum instance cliques, use the two-stage mining method based on maximum instance cliques proposed by the present invention to calculate the participation degree PI(c) of all business co-location patterns. Classify the business co-location patterns that meet the minimum frequency threshold as global frequent business co-location patterns, and regard the patterns that do not meet the threshold as regional candidate business co-location patterns. The specific implementation process is as follows:

[0086] Step S31: In order to quickly find the participation examples of co-location patterns through maximum instance cliques and calculate the participation degree PI(c), a Hash table structure is designed to store all maximum instance cliques.

[0087] The storage method of the Hash table structure designed in the present invention is a double Hash structure <business co-location pattern, <business category, instance number>>, called CL-hash. This storage method can directly find all corresponding maximum instance cliques through the business model name. Among them, the instance numbers are stored in the form of a Set (collection), so that the same number will not be stored repeatedly. When calculating the participation degree PI(c) of business co-location patterns, the number of participation instances can be directly obtained by getting the size of the Set. Figure 3 Shows the Hash storage structure generated according to Figure 2 the maximum instance cliques in the dataset. For example, the key of {A.3, B.3, C.1, D.4} is {A, B, C, D}, and the value is also a set of key-value pairs

[0088] {<A, {3}>, <B, {3}>, <C, {1}>, <D, {4}>}.

[0089] Among them, the business co-location pattern is expressed as k is the number of business categories included in the business co-location pattern c, which is called the order of the business co-location pattern c. If a spatial set cl contains all the business categories in the business co-location pattern c, and no subset of cl contains all the business categories in c, then cl is called a row instance of the business co-location pattern c. All row instances of the business co-location pattern c constitute the table instance of the business co-location pattern c, denoted as T(c). A business category f i All non-repeated business instances in the table instance of the business co-location pattern c constitute the business category f iExample of participation in the business co-location pattern c.

[0090] In Figure 2 Among them, the business co-location pattern {A, B, C} is a 3-order business co-location pattern. The row instances of the business co-location pattern {A, B, C} are {A.3, B.3, C.1} and {A.4, B.4, C.3}. Then the table instance of the business co-location pattern {A, B, C} is {{A.3, B.3, C.1}, {A.4, B.4, C.3}}. The participation instances of business category A in the business co-location pattern {A, B, C} are {A.3, A.4}.

[0091] Step S32: Based on the CL-hash structure, use the two-stage mining method proposed by the present invention to quickly calculate the participation degree PI(c) of all business co-location patterns. Classify the business co-location patterns that meet the minimum frequency threshold as global frequent business co-location patterns, and regard the patterns that do not meet the threshold as regional candidate business co-location patterns.

[0092] The participation degree PI(c) of the business co-location pattern c is the minimum value of the participation rates of all business categories f i (1 ≤ i ≤ k), that is:

[0093]

[0094] To measure the frequency (interestingness) of a business co-location pattern, the participation degree is used to measure the frequency (interestingness) of a business co-location pattern in business co-location pattern mining, and the participation degree of a business co-location pattern depends on the participation rates of each business category in the pattern. In the business co-location pattern c, the participation rate of the business category f i (1 ≤ i ≤ k) is expressed as PR(c, f i ), which is to measure the participation status of the business category f i in the k-order business co-location pattern c = {f1, f2,..., f k}}. PR(c, f i ) is the ratio of the number of non-repeated instances of f i in the table instance T(c) to the total number of instances of the business category f i in the spatial dataset, that is

[0095]

[0096] Among them, represents the number of non-repeated occurrences of the business instances of the business category f i in the table instance of the business co-location pattern c. Its essence is the number of participation instances of the business category f i in the business co-location pattern c; |T{fi i}| represents the business category fi The total number of business instances; П is the projection operation of the relationship, indicating that the same business instance can only be counted once.

[0097] In Figure 2 there are 5 business instances in business category A, 4 business instances in business category B, 3 business instances in business category C, and 5 business instances in business category D. The table instances of business model c = {A, B} are {{A.1, B.2}, {A.3, B.3}, {A.4, B.4}}, then PR(c, A) = 3 / 5, PR(c, B) = 3 / 4, PI(c) = min{PR(c, A), PR(c, B)} = 0.75. The table instances of business model c = {A, B, C} are {{A.3, B.3, C.1}, {A.4, B.4, C.3}}, then PR(c, A) = 2 / 5, PR(c, B) = 2 / 4, PR(c, C) = 2 / 3, PI(c) = min{PR(c, A), PR(c, B), PR(c, C)} = 0.4. If the minimum frequency threshold min_prev is set to 0.6, then the business model c = {A, B} is a frequent business co-location pattern, while the business model c = {A, B, C} is not a frequent business co-location pattern.

[0098] Clique candidate pattern: The key in CL-hash (i.e., the business model) is called the clique candidate business model and can be directly obtained by collecting the maximum instance cliques.

[0099] Lemma 1. The row instances of any candidate business model must be included in the maximum instance clique in the space.

[0100] Proof: Since the row instance is actually a clique relationship that satisfies the spatial neighborhood relationship, and any clique relationship in the data set must be included in the maximum instance clique in the data set, any candidate row instance must be obtained from the maximum instance clique.

[0101] Lemma 2. For any remaining business model not in the clique candidate business model, its participating instances can be obtained from the participating instances of the clique candidate business model.

[0102] Proof: If there is no corresponding maximum instance clique for the candidate business model c, then according to Lemma 1, all row instances of c must be included in the maximum instance clique, and all maximum instance cliques are already included in the table instances of the clique candidate business model. Therefore, the participating instances of the remaining business model c not in the clique candidate can be obtained from the participating instances of the clique candidate business model.

[0103] Furthermore, a two-stage calculation method is designed according to Lemma 1 and Lemma 2.

[0104] The specific process is shown in Algorithm 2:

[0105] Algorithm 2: Two-stage Mining Algorithm

[0106]

[0107] The two-stage mining algorithm is described as follows: In the first stage, since CL-hash stores all table instances of clique candidates, the participation index PI(c) of all clique candidate business models can be directly calculated by looking up CL-hash, and the business co-location patterns that meet min_prev are added to the result set (Steps 1-4). For example, in Figure 3 , the table instances of the clique candidate business model {B, D} are {B1, D1}, {B2, D2}, so PR({B, D}, B) = 1 / 2, PR({B, D}, D) = 2 / 5, PI({B, D}) = min(PR({B, D}, B), PR({B, D}, D)) = 1 / 2. In the second stage, all remaining business models that are not in the clique candidate business models are generated, and the participation index PI(c) of all remaining business models is calculated based on the participation instances of business categories in the clique candidate business models (Steps 5-8). Then, frequent business co-location patterns are filtered according to the minimum frequency threshold min_prev (Step 9). For example, in Figure 3 , the business model {B, C} is not a clique candidate business model, but the participation instances of the business model {B, C} can be obtained by merging the participation instances in the supersets {A, B, C, D} and {A, B, C} of business category B and business category C. The table instances of the business models {A, B, C, D} and {A, B, C} are {<A, {3}>, <B, {3}>, <C, {1}>, <D, {4}>} and {<A, {4}>, <B, {4}>, <C, {3}>}. Therefore, the table instances of the business model {B, C} are {<B, {3, 4}>, <C, {1, 3}>}.

[0108] Note that this algorithm is applicable to both global business co-location pattern mining and regional business co-location pattern mining. For global business co-location pattern mining, the business co-location patterns that meet min_prev are added to the global frequent business co-location patterns, and the business co-location patterns that do not meet the criteria are regarded as regional candidate business models. For regional business co-location pattern mining, only the business co-location patterns that meet min_prev need to be found within the divided regions.

[0109] Step S4: Convert all the maximum instance cliques obtained in Step S2 into their geometric center points, and then project these geometric center points onto an N×N grid. Set multiple density thresholds and perform multi-density grid clustering on these points to subdivide the entire research area into multiple sub-regions.

[0110] The specific steps are as follows:

[0111] Step S41: Convert all the maximum instance clusters obtained in step S2 into their geometric center points. After the conversion, it is equivalent to converting all the maximum instance clusters into a new spatial data set. In this new spatial data set, construct an N×N grid, where the value of N is not fixed and depends on the scale of the data set. The larger the data set scale, the larger the value of N. Generally, the value of N is set between 10 and 50. Project the converted spatial data set onto the grid plane. The specific steps are as follows:

[0112] Step S411: Convert all the maximum instance clusters obtained in step S2 into their geometric center points. After the conversion, it is equivalent to converting all the maximum instance clusters into a new spatial data set. Figure 4 (a) of [reference] shows how to convert the maximum instance cluster into the geometric center.

[0113] Step S412: Construct an N×N grid, where the value of N is not fixed and depends on the scale of the data set. The larger the data set scale, the larger the value of N. Generally, the value of N is set between 10 and 50. Project the converted spatial data set onto the grid plane. As shown in Figure 4 (a) of [reference], a 9×6 grid is constructed. The red pentagrams in the figure are the converted geometric center points. Project these set center points onto the grid plane. Assuming the clustering density is set to 1, the yellow areas are the two formed partitions. Figure 4 (b) of [reference] shows the final effect that the grid density clustering based on the maximum instance cluster aims to achieve.

[0114] Step S42: Set multiple density thresholds and traverse from large to small according to the density thresholds. Each time, take out a density value and input it into the grid density clustering algorithm (CLIQUE algorithm) to cluster the converted spatial data set, so as to divide the entire research area into multiple sub - areas.

[0115] R.Agrawal, J.Gehrke, D.Gunopulos, P.Raghavan. Automatic Subspace Clustering of High Dimensional Data for Data Mining Applications. Proceedings of the 1998 ACM SIGMOD international conference on Management of data. proposed the grid density clustering algorithm CLIQUE. Its principle is to divide the data space into cells (grids), map data objects to cells, and calculate the density of each cell. According to the preset threshold, judge whether each cell is a high - density cell, and a "class" is composed of adjacent dense cells.

[0116] Further, the specific implementation of multi-density grid clustering is shown in Algorithm 3:

[0117] Algorithm 3: Multi-density Grid Clustering Algorithm

[0118]

[0119] Among them, a density threshold list densities is set, such as [4*density, 2*density, density, 0.5*density, 0.25*density], where the value of density is defaulted to the value obtained by mapping the scale of the spatial data set to each grid on average. clique_instance represents the clustering result of the CLIQUE algorithm, including the list of clustered clusters clique_cluster and the set of noise points noise.

[0120] The algorithm description is as follows: Set the grid size N (Step 1). Set a multi-density list densities (Step 2). Take a density threshold from the density list in descending order (Step 3), and perform grid-based density clustering on the current data set according to the current density (Step 4). Respectively take out the regions and noise points obtained by the CLIQUE algorithm at the current density (Steps 5 and 6). Traverse the current region. If the number of data points in the current region is greater than a certain number, usually set to 4 times the current density value, then the current region is considered reasonable and added to the final region list. Otherwise, the region is considered unreasonable, and the points in the region are added to the noise (Steps 7-10). Update the data set so that the noise points become the data set for the next round of clustering (Step 11), and then return to Step 3.

[0121] There is no strict definition on how to set the grid size. However, based on the experience of multiple experiments, it is recommended to set the grid size to N×N, where N ranges from 10 to 50. The specific value depends on the data volume of the data set. Here, two critical values are recommended. If the data volume of the data set is less than 10,000, it is set to 10. If the data volume of the data set exceeds 100,000, it is set to 50 or even larger.

[0122] Step S5: Regard each sub-region as an independent research region, and adopt a two-stage calculation method based on the maximum instance clique in each sub-region to calculate the participation degree PI(c) of all candidate business co-location patterns in the regional candidate business model set. The business co-location patterns that meet the minimum frequency threshold are the regional frequent business co-location patterns. The sub-regions where these frequent business co-location patterns are located are their frequent regions. Then merge these sub-regions, and merge the sub-regions with the same or similar semantic information into one sub-region.

[0123] The specific steps are as follows:

[0124] Step S51: After step 4, the global research area is divided into several non - overlapping sub - regions. In these sub - regions, the two - stage algorithm in step 2 is used to obtain the regional frequent business co - location patterns.

[0125] Step S52: After step S51, several non - overlapping sub - regions are obtained, and the regional frequent business co - location patterns corresponding to these sub - regions are mined. However, although these sub - regions do not overlap geographically, they may describe the same meaning, such as being all business districts, and their regional frequent business models are highly similar. Then such regions should be merged. The specific merging steps are as follows:

[0126] The merging of sub - regions uses the calculation of the intersection - over - union ratio CIOU(Ri, Rj) between the frequent business co - location patterns of two regions to determine whether the two regions are similar, that is:

[0127]

[0128] if CIOU(Ri, Rj)≥min_simi:

[0129] Ri_J = {Ric}∩{Rjc}

[0130] Theorem 1. The intersection of the frequent regional business co - location patterns of region Ri and region Rj is also frequent in the merged region Ri_j.

[0131] Proof: Regions Ri and Rj are sub - regions divided from the entire research area, and it can be determined that these sub - regions do not overlap. For any k (k > 2) - order regional business co - location pattern c in Ri, there is always:

[0132] RiPI(c)≥min_prev

[0133] For all RiPR(c, fi):

[0134]

[0135] Similarly, for region Rj:

[0136]

[0137] Since there is no overlap between region Ri and region Rj, for any common regional business co - location pattern c of the two regions, there is:

[0138]

[0139] Therefore, if regions Ri and Rj are merged into a new region Ri_j, the intersection of the frequent business co-location patterns of regions Ri and Rj must also be frequent.

[0140] Let {Ric} denote the frequent business co-location patterns in region Ri. If CIOU(Ri, Rj) is greater than the minimum similarity threshold min_simi of the two given regions, then these two regions are merged, and the frequent business co-location patterns of the merged new region are the intersection of the frequent business co-location patterns of these two regions. min_simi is the minimum similarity threshold between the two given regions.

[0141] In Figure 5 Suppose the minimum frequency threshold min_prev is set to 0.6, and the sets of business co-location patterns {R1c} and {R3c} of regions R1 and R3 are both {{A, C}}. In region R1, R1PR(c, A) = 3 / 4, R1PR(c, C) = 1, R1PI(c) = min{R1PR(c, A), R1PR(c, C)} = 0.75. In region R3, R3PR(c, A) = 1, R3PR(c, C) = 1, R3PI(c) = min{R3PR(c, A), R3PR(c, C)} = 1. If the region minimum similarity threshold min_simi is set to 0.8, then CIOU(R1, R3) = 1, CIOU(R1, R3) > min_simi, so regions R1 and R3 are merged into a new region, denoted as R1_3. The frequent business co-location patterns of the merged region are the intersection of the frequent business co-location patterns of the original two regions. For R1_3, R1_3PR(c, A) = 5 / 6, R1_3PR(c, C) = 1, R1_3PI(c) = min{R1_3PR(c, A), R1_3PR(c, C)} = 5 / 6, and 5 / 6 > 0.6.

[0142] Step S6: If there is no special specification, then display all global frequent and regional frequent business co-location patterns, as well as the regions where the regional frequent business co-location patterns are located; if the user specifies a search range, then display the frequent business co-location patterns within that range; if the user specifies an interested co-location pattern or business category, then display all frequent business co-location patterns and frequent regions that contain that co-location pattern or business category.

[0143] The specific steps are as follows:

[0144] Step S61: If the user has no special specification, then display all mining results.

[0145] Step S62: If the user specifies the specific location and range where they want to start a business, then give the frequent business co-location patterns within that range for the user's reference.

[0146] Step S63: If the user gives the business category or business model they want to engage in, search the mining results of all sub-regions, and display the frequent regions containing the business category or business model to the user for reference.

[0147] For example, in Figure 5 , assume that the minimum frequency threshold min_prev is set to 0.6, and the minimum regional similarity threshold min_simi is set to 1. Without special specification, the algorithm will output that the business co-location patterns {A, B} and {A, C} frequently appear in region R1, and the business co-location pattern {A, C} frequently appears in region R3; if the business activity range is specified as region R1 at the beginning, the business co-location patterns {A, B} and {A, C} will be output, indicating that the business co-location patterns frequently appearing in region R3 are {A, B} and {A, C}. If the user specifies at the beginning that it is related to business category A, the business co-location pattern {A, B}: {R1, R3} and the business co-location pattern {A, C}: {R1} will be output.

[0148] All steps of the present invention are as described above. For the convenience of understanding, a summary is made here and named the RCM-MCC algorithm. The RCM-MCC algorithm is the implementation of the complete process of the business decision-making method based on regional frequent co-location patterns proposed by the present invention, and is also Figure 1 the algorithm corresponding to the framework.

[0149] Given four parameter space datasets S, spatial proximity relationship d, minimum frequency threshold min_prev, and minimum regional similarity threshold min_simi. The specific implementation is shown in Algorithm 4:

[0150] Algorithm 4: RCM-MCC Algorithm

[0151]

[0152] Example:

[0153] In order to verify the correctness of the present invention, this embodiment is divided into two parts: synthetic data sets and real data sets. The synthetic data sets are also divided into four categories, which respectively verify the performance of the method in synthetic data sets with different instance distribution densities, the performance in synthetic data sets with overlapping areas in randomly generated spatial distributions of commercial categories, the performance in generating spatial distribution data sets of commercial categories with different sparsity, and the performance in randomly generated data sets containing only global business models. As for the real data set, this embodiment uses the urban point of interest (POI) data set of Shenzhen as the real embodiment of this time, in order to verify the correctness and effectiveness of the present invention. At the same time, this embodiment also expands the data set, and tests are also conducted on the vegetation data set in Beijing. Experiments show that the present invention can also run correctly on the expanded data set. All experiments in this embodiment are implemented in Java and performed on a computer configured with a Windows 10 operating system, Intel i5-11800H2.30GHz, and 16G memory.

[0154] This embodiment divides the experiment into two parts. First, the correctness and applicability of the present invention are verified on a synthetic data set. Second, actual tests are performed on a real data set and the results are compared with other algorithms to verify the practical applicability and advancement of the present invention.

[0155] 1. Experiments and Analysis on Synthetic Datasets

[0156] The purpose of synthesizing the data set is mainly to verify the correctness of the present invention and to derive the applicable scenarios of the present invention.

[0157] 1) Analysis of applicable scenarios

[0158] In order to verify the applicable scenarios of the present invention, two synthetic data sets with different distribution trends of business types are designed in this embodiment, namely a data set containing regionally distributed business categories and a data set containing only globally distributed business categories. Figure 6 The distribution characteristics of the two data sets are shown respectively. Figure 6 (a) in FIG. 5 is a data set containing regionally distributed business types, which is composed of 4 regionally distributed business categories and 1 globally distributed business. Figure 6 (b) in the figure is a dataset that only contains globally distributed commercial categories.

[0159] Figure 7 Shows Figure 6 The maximum number of instance clusters and the time required to generate them under different distance thresholds for two synthetic datasets with different distribution categories. Figure 7It can be seen that the number of maximum instance cliques generated on the globally distributed synthetic dataset and the time spent are much greater than those on the regionally distributed dataset. As the distance threshold increases, the difference becomes more and more obvious. The experimental results show that compared with the dataset showing a global distribution trend, the method of the present invention performs better on the dataset with regional distribution characteristics, which also indicates that the present invention is more suitable for discovering regional frequent business co-location patterns.

[0160] 2) Analysis of Partition Results

[0161] In addition, this embodiment also verifies the performance of the method of the present invention on synthetic datasets of various different distribution types. The embodiment designs 4 groups of synthetic datasets with different characteristics, such as Figure 8 shown. Among them Figure 8 (a) contains different distribution densities of certain business categories to verify the performance of the invention under different distribution densities; Figure 8 (b) describes that different business categories may have overlapping parts in spatial distribution to verify whether the invention can identify these overlapping parts; Figure 8 (c) describes that different business categories may have different degrees of sparsity in spatial distribution to verify whether the invention can identify the regional business models that are sparse but frequent in spatial distribution. Finally Figure 8 (d) is a globally distributed synthetic dataset for comparative analysis.

[0162] Figure 9 is the maximum instance clique clustering result generated by using the method of the present invention for the Figure 8 corresponding dataset. It can be seen that for the Figure 8 (a) with different distribution densities in the region, the present invention can identify it well and has the ability to subdivide the research area according to the prosperity degree. For the Figure 8 (b) synthetic dataset with overlapping regions of business category spatial distribution in the figure, the present invention can also correctly identify the overlapping regions of different scales. The partition results of 9(b) in the figure correspond to the Figure 8 (b) overlapping parts in the figure. Moreover, the present invention also identifies different overlapping levels, subdivides the regions with different numbers of overlapping business categories, and in addition excludes the influence of individual business categories. For the Figure 8 (b) regions with individual business category distributions are not divided into regions. It can be seen from the Figure 9 (c) in the figure that for the business model distributions with different degrees of sparsity, the present invention can also correctly identify them. As for the globally distributed business categories, Figure 9The partitioning result of (d) in [the dataset] is almost a complete region. For some small regions in the above results, they might originally belong to the adjacent partitions but are separated individually. This situation mainly occurs because the data distribution in the synthetic dataset is random. Random distribution does not mean uniform distribution, which leads to inconsistent distribution densities in some local regions, and thus they are split during density-based partitioning. However, this is only the result of the preliminary partitioning. These split small regions themselves have the same regional co-location pattern as the adjacent large regions. During the final region merging, these regions will be merged together again. Experiments show that, based on the partitioning results, the present invention can correctly identify local sub-regions and global regions.

[0163] 2. Experiments and Analyses on Real Datasets

[0164] In this embodiment, the present invention is verified on the urban Point of Interest (POI) dataset in Shenzhen. Meanwhile, to verify the extensibility and universality of the present invention, a vegetation dataset in Beijing is also introduced in this embodiment to test the actual performance of the present invention. Figure 10 The spatial distributions of two real datasets are respectively shown, where Figure 10 (a) is the vegetation distribution dataset in Beijing, Figure 10 (b) is the urban Point of Interest (POI) distribution dataset in Shenzhen. To verify the mining results and algorithm performance of the present invention, the multi-level method proposed in M.Deng, J.Cai, Q.Liu, Z.He, J.Tang. Multi-level method for discovery of regional co-location patterns. International Journal of Geographical Information Science is selected as the comparison method. (Since the usage scenarios of the invention are expanded in the real dataset and are not only applied to the commercial field, the description methods of commercial categories and commercial instances are no longer adopted in the following experimental processes, and the expressions of spatial features and spatial instances are used instead.)

[0165] 1) Comparative Analysis of Mining Results

[0166] For the urban Point of Interest (POI) dataset in Shenzhen, the minimum frequency threshold min_prev is set to 0.5, and the distance threshold d is set to 120. The minimum region similarity threshold min_simi required in the present invention is set to 0.8, and the abundance threshold mentioned in the multi-level method is set to 0.05.

[0167] Table 1. Mining Results of the Shenzhen Point of Interest (POI) Dataset

[0168]

[0169] For the extended Beijing vegetation distribution dataset, the minimum frequency threshold min_prev is set to 0.6, and the distance threshold d is set to 300. The minimum regional similarity threshold min_simi required in the present invention is set to 0.8, and the abundance threshold mentioned in the Multi-level method is set to 0.05.

[0170] Table 2. Mining results of the Beijing vegetation distribution dataset

[0171]

[0172] Table 1 and Table 2 are respectively the mining results of two real datasets, where the dashed lines indicate no mining results. Since both the RCM-MCC method designed in the present invention and the comparative algorithm Multi-level algorithm first mine the global frequent co-location patterns and then use the globally infrequent co-location patterns as the regional candidate co-location patterns, the global co-location patterns mined by the RCM-MCC method and the Multi-level method are the same. For the regional co-location patterns, it can be seen from the table that most of the regional co-location patterns mined by the two algorithms overlap, which also demonstrates the feasibility of the RCM-MCC method and the correctness of the mining results. At the same time, in order to more intuitively understand the effect that the multi-density grid clustering algorithm based on the maximum instance clique wants to achieve, in this embodiment, a visualization display of the grid-based multi-density clustering process is made on the Beijing vegetation distribution dataset, and the clustering process and the final effect are as Figure 11 shown. In addition, it can also be seen from Figure 11 that the clustering method adopted in the present invention is different from the traditional clustering methods. The method of the present invention not only integrates the relevant advantages of the maximum instance clique, but also fully considers the density distribution of real-life datasets: the present invention takes the maximum instance clique as the minimum element, first selects the region with the highest density in the dataset, and then determines whether the selected region meets the requirements of the minimum partition. If the conditions are met, the region is added to the final region list; if not, the region is returned to the original dataset, and then the regional division of the next round of density is carried out.

[0173] For the different parts mined by the two methods, this embodiment also conducts relevant analyses: 1) The starting ideas of the two mining methods are different. The RCM-MCC method proposed in the present invention mines regionally frequent business co-location patterns based on maximum instance cliques, while the Multi-level method focuses on considering the adjacency relationships between different business instances. 2) The region division methods are different. The RCM-MCC method proposed in the present invention starts from the semantic perspective. According to actual investigations, it is found that the entire business field is hierarchically distributed, roughly divided into prosperous areas, sub-prosperous areas, ordinary living areas, suburbs, and towns. Therefore, the present invention uses different density thresholds to fit the real distribution. The Multi-level method uses the Delaunay triangulation method to adaptively cluster the same patterns into one region. This results in great differences in the finally divided sub-regions. The region division result of the RCM-MCC method proposed in the present invention is centered on high-density regions, presenting a hierarchical structure, with a large area and a moderate number of sub-regions. However, the sub-regions divided by the Multi-level method have no specific rules, with small area and a large number of regions. The experimental results show that the present invention is superior to the Multi-level method and is more interpretable.

[0174] 2) Analysis of the influence of dataset scale on algorithm efficiency

[0175] To prove the scalability of the present invention, this embodiment conducts analysis experiments on datasets of different scales, mainly including the influence of the number of instances on the algorithm, the influence of the number of features on the algorithm, and the influence of different minimum frequency thresholds on the algorithm.

[0176] Influence of the number of instances on the algorithm:

[0177] To test the influence of the number of instances on the running time of the RCM-MCC method proposed in the present invention, this embodiment compares the running times of the two methods (the RCM-MCC method and the Multi-level algorithm) under different numbers of instances: In this comparative experiment, to discover effective regional co-location patterns, this embodiment uses a data distribution category similar to Figure 6 in (a). A synthetic dataset is generated using a data generator, with the number of features fixed at 5 and the number of instances set to 5000, 10000, 15000, 20000, and 25000 respectively. Figure 12The running times of two methods under different numbers of instances are described. It can be seen that the running times of the two methods under different numbers of instances are almost the same, and sometimes the performance of the Multi-level algorithm is better. However, this is only the result of fixing the minimum frequency threshold. If the minimum frequency threshold is changed, the RCM-MCC proposed by the present invention can quickly obtain new mining results without re-running the algorithm, which is incomparable to the Multi-level algorithm. Considering this point, the RCM-MCC method of the present invention is generally superior to the Multi-level algorithm.

[0178] Effect of the number of features on the algorithm:

[0179] To more intuitively study the impact of the increase in the number of features on the algorithm efficiency, in this embodiment, the worst-case scenario is considered, and the experiment is carried out under the data distribution similar to that in Figure 6 (b) above. The number of commercial instances is fixed at 20,000, and the number of spatial features is set to 6, 8, 10, 12, 14, and 16 respectively for the experiment. Figure 13 Describes the running times of the two algorithms under different numbers of spatial features and the number of regional frequent co-location patterns generated. As can be seen from Figure 13 (a) above, when the number of commercial instances is constant and the number of spatial features is small, increasing the number of spatial features has little impact on the efficiency of the RCM-MCC method and the Multi-level method proposed by the present invention. However, as the number of spatial features continues to increase, the method proposed by the present invention begins to gradually lag behind the Multi-level method in terms of efficiency. According to the experimental analysis, this is mainly because the present invention is designed based on the maximum instance clique, aiming to mine regional frequent business models. When the number of spatial features increases, the number of maximum instance cliques will also increase rapidly, thus affecting the running time of the present invention. Although the performance of the present invention seems slightly worse than that of the Multi-level method for the time being, there is an interesting phenomenon in the experimental results: as the number of spatial features increases, the error rate of the mining results of the Multi-level method also shows a rapid growth trend. In this comparative experiment, the synthetic data used in this embodiment is similar to the data distribution in Figure 6 (b) above, that is, the mining results should not contain regional frequent co-location patterns. But from Figure 13As can be seen from (b) of , as the number of spatial features increases, the number of region frequent co-location patterns obtained by the Multi-level method grows rapidly. The number of region frequent co-location patterns mined by the present invention is always 0, which also shows that the correctness of the present invention is higher than that of the Multi-level method. And this comparative experiment is the performance on the dataset of the worst global distribution category. In real life, the distribution of business models cannot all be global. Therefore, the actual performance of the present invention should be much higher than this comparative experiment.

[0180] The influence of different minimum frequency thresholds on the algorithm:

[0181] In the previous comparative experiment, it was repeatedly mentioned that the mining framework based on the maximum instance clique is insensitive to the frequency threshold. When the minimum frequency threshold changes, the present invention can quickly respond to obtain new mining results. Figure 14 shows the running times of the two methods under the minimum frequency threshold. From Figure 14 it can be easily seen that the method of the present invention only consumes more time in the first run, and hardly consumes time when the subsequent minimum frequency threshold changes, while the Multi-level method needs to run again every time the minimum frequency threshold changes.

[0182] The comprehensive experimental results show that the present invention is higher than the latest Multi-level method in terms of correctness, scalability and performance. This is sufficient to show the innovation and superiority of the present invention.

[0183] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.

Claims

1. A business decision-making method based on regional frequent co-location patterns, characterized in that, Including the following steps: Step S1: Determine the global mining area and all the business categories of interest within the mining range, and obtain the spatial business category set F = {f1, f2,..., f n}, where n represents the number of business categories; each business category consists of multiple business instances, and the business instance set of a business category is represented as S = {S i}, 1 ≤ i ≤ n; Step S2: Calculate the proximity relationship between business instances based on a given distance threshold d to obtain the spatial instance neighbor table of each business instance in the business instance set S, and then use the partitioned Bron–Kerbosch algorithm to calculate all maximal instance cliques; Step S3: According to all the maximal instance cliques obtained in S2, use a two-stage mining method based on maximal instance cliques to calculate the participation degrees of all business co-location patterns, classify the business co-location patterns that meet the minimum participation degree threshold as globally frequent business co-location patterns, and regard the patterns that do not meet the threshold as regional candidate business co-location patterns; Step S4: Convert all the maximal instance cliques obtained in Step S2 into corresponding geometric center points, and then project the geometric center points onto an N×N grid; Set multiple density thresholds, and perform multi-density grid clustering on the geometric center points to subdivide the entire region into multiple sub-regions; Step S5: Treat each sub-region as an independent research area, use the two-stage mining method in each sub-region to calculate the participation degrees of all patterns in the set of regional candidate business co-location patterns, and the business co-location patterns that meet the minimum participation degree threshold are the regional frequent business co-location patterns; the sub-region where the regional frequent business co-location pattern is located is the corresponding frequent region; then merge the sub-regions with the same or similar semantic information into one sub-region and find the regional frequent business co-location patterns of the merged region; Step S6: Give corresponding decision results according to different requirements.

2. A business decision-making method based on regional frequent co-location patterns according to claim 1, characterized in that In the step S1, the global mining area is a spatial coordinate system, and the spatial business categories are represented in the form of a business category list {A, B,..., Z}, where each letter represents a business category; the business instance O i is represented in the form of a triple <belonging business category, instance number, spatial position>, where the instance number represents the serial number of the business instance among all instances of the belonging business category, and the spatial position is composed of a two-dimensional array [x, y], where x and y respectively represent the abscissa and ordinate of the business instance in the global spatial coordinate system.

3. A commercial decision-making method based on regional frequent co-location patterns according to claim 1, characterized in that, The specific implementation process in the said Step S2 is as follows: Step S21: Given a proximity relation spatial proximity distance threshold d, two business instances have a neighbor relationship when the Euclidean distance between them is less than the threshold d; sort the business categories in alphabetical order, and sort the business instances included in each business category by instance number; calculate the Euclidean distance between any two business instances O i and O j in the business instance set S, where O i and O j belong to different business instance sets S; two business instances that satisfy being less than the distance threshold d are neighbors of each other, and obtain the business instance neighbor table corresponding to the business instance set S = {s1, s2,..., s n}; Step S22: Perform partition operation on all business instances according to the business instance neighbor table. The partition steps are as follows: Perform partition operation on all business instances. First, select an unvisited business instance P from the instance set S, put the business instance P and its neighbor business instances into the set N, then regard the set N and the business instances having a proximity relationship with N as a partition, and finally update the instance set S; Step S23: Pass the partitioned business instance sets into the Bron–Kerbosch algorithm in turn to obtain all maximal instance cliques.

4. A business decision-making method based on regional frequent co-location patterns according to claim 3, characterized in that The specific steps of the Bron–Kerbosch algorithm in the said Step S23 are as follows: First, initialize three disjoint sets Q, R, X, where Q represents the set of generated maximal instance cliques, R represents the set of instances to be processed, and X represents the set of instances that have been processed. The initialized Q and X are empty sets, and R is obtained from the partitioned business instance set; First, select the vertex u with the largest number of adjacent vertices in P as the pivot element, put each non-neighbor vertex v of u in R into the set Q, and remove the vertices that are not adjacent to v from R and X, then recursively set the sets Q, R, X, and move the selected vertex u from R to X; Repeat the above steps until the union of R and X is an empty set, and at this time Q is the finally obtained maximal instance clique.

5. A business decision-making method based on regional frequent co-location patterns according to claim 1, characterized in that The specific steps of step S3 are as follows: Step S31: Store all maximal instance cliques through a Hash table structure, and find the participating examples of business co-location patterns; Step S32: Based on the CL-hash structure, calculate the participation degree PI(c) of all business co-location patterns through a two-stage mining method based on maximal instance cliques, and first find the globally frequent business co-location patterns according to the given minimum frequency threshold min_prev. The globally frequent business co-location patterns will no longer participate in the mining of regional frequent business co-location patterns. The patterns that do not meet the threshold are used as regional candidate business co-location patterns. The regional candidate business co-location patterns will participate in the subsequent mining of regional frequent business co-location patterns, and finally obtain the frequent regional business co-location patterns; In step S31, the Hash table structure storage method is a double Hash structure <business parity mode, <business category, instance number>>, called CL-hash, where the business parity mode is represented by k is the number of business categories contained in the business homology pattern c, which is called the order of the business homology pattern c. i All non-repeated business instances in the table instance of business homography pattern c constitute business category f i Participation example in commercial parity pattern c, where 1≤i≤k; The table instance is all row instances of the business co-location pattern c; The row instance means that when a spatial set cl contains all business categories in the business co-location pattern c, and any subset of cl does not contain all business categories in c, then the spatial set cl is called a row instance of the business co-location pattern c.

6. A business decision-making method based on regional frequent co-location patterns according to claim 5, characterized in that In the step S32, the participation index PI(c) of the business co-location mode c is the minimum value of the participation rates of all business categories f i (1 ≤ i ≤ k), that is: Where where PR(c,f i ) is the participation rate of business category f i ; represents the number of participation examples of business category f i in this co-location pattern c; |T{f i}| represents the total number of business instances of business category f i ; Π is the projection operation of the relationship, indicating that the same business instance can only be counted once to remove redundancy; T(c) represents the table instance of pattern c; The specific process of the two-stage mining method based on maximal instance cliques is as follows: In the first stage, the PI(c) values of all clique candidate business models are directly calculated through CL-hash, and the business co-location patterns with PI(c) values greater than or equal to min_pre are used as globally frequent business co-location patterns, and those less than min_pre are used as regional candidate business models; In the second stage, the table instances of all remaining business models that are not in the clique candidate business models can be obtained by collecting all supersets containing the remaining business models, and the PI(c) values of all remaining business models are calculated according to the participating examples in the table instances, and those with PI(c) values greater than or equal to min_pre are used as globally frequent business co-location patterns, and those less than min_pre are used as regional candidate business models.

7. A business decision-making method based on regional frequent co-location patterns according to claim 1, characterized in that The specific implementation process of step S4 is as follows: Step S41: Convert all the maximal instance cliques obtained in step S2 into corresponding geometric center points. After the conversion, a new spatial data set is formed; in the new spatial data set, an N×N grid is constructed, and the value of N is not fixed and depends on the scale of the data set. The larger the data set scale, the larger the value of N. Project the converted spatial data set onto the grid plane; Step S42: Set multiple density thresholds, and perform multi-density grid clustering on the converted spatial data set from large to small according to the density thresholds to divide the entire region into multiple sub-regions.

8. A business decision-making method based on regional frequent co-location patterns according to claim 7, characterized in that In step S42, the specific process of multi-density grid clustering is as follows: (1) Set a list of density thresholds density, where the default value of density is the value obtained by mapping the scale of the spatial dataset to each grid on average; (2) Sort the dataset according to the density thresholds from large to small and perform grid density clustering on the dataset. If the clusters obtained by clustering meet the specified size, the clusters obtained by clustering are regarded as qualified sub-regions; otherwise, it is considered that they do not meet the conditions to become the smallest regions, and the data they contain is regarded as noise points; (3) Decrease the density threshold and put the noise points into the next round of grid density clustering; (4) Until the list of density thresholds is traversed, all sub-regions are obtained.

9. The commercial decision-making method based on regional frequent co-location patterns according to claim 1, wherein in step S5, when merging sub-regions, the intersection-over-union ratio CIOU(Ri, Rj) between the frequent commercial co-location patterns of two regions is calculated to determine whether the two regions are similar, that is: where {Ric} represents the set of frequent commercial co-location patterns in region Ri. If CIOU(Ri, Rj) is greater than the minimum similarity threshold min_simi of the two given sub-regions, then these two sub-regions are merged, and the frequent commercial co-location patterns of the merged new sub-region are the intersection of the frequent commercial co-location patterns of these two sub-regions.

10. The commercial decision-making method based on regional frequent co-location patterns according to claim 1, wherein the specific implementation process of step S6 is as follows: Step S61: If there is no special specification, display all mining results; Step S62: If the user specifies the specific location and range where they want to start a business, give the frequent commercial co-location patterns within the specific location and range; Step S63: If the commercial category or business model that the user wants to engage in is given, search the mining results of all sub-regions and give the frequent regions containing the corresponding commercial category or business model.