A city interest point homotopy pattern mining method and device
By utilizing topic models and network distance matrices for isomorphic pattern mining at the urban functional area scale, the errors caused by neglecting facility differences in existing technologies are resolved, achieving more accurate isomorphic pattern recognition for urban POIs.
Patent Information
- Application Number
- CN202310168009.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-24
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2043-02-24
AI Technical Summary
Existing isomorphic pattern mining methods fail to consider factors such as facility size, management methods, and operating hours in urban facility research, leading to the neglect of instance differences within the same type of facility and potentially causing problems such as 'variable facets' and 'ecological fallacies'.
A topic model based on urban residents' movement trajectories is used to identify urban functional areas, and a network distance matrix is used to mine co-location patterns between points of interest (POIs). Spatial neighborhood relationships are constructed through the DMR topic model and the network distance matrix to identify frequent co-location patterns.
It can more accurately reflect the actual situation of the city, reduce the amount of calculation, improve the accuracy and efficiency of isotope pattern mining, and identify heterogeneous patterns under different functional areas.
Smart Images

Figure CN116401470B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of urban management technology, and in particular to a method and apparatus for mining urban point of interest co-location patterns. Background Technology
[0002] With the development of computer technology, geospatial big data has expanded rapidly, exceeding people's ability to interpret it. In order to identify the hidden information in the massive amount of data, a brand-new research field—spatial data mining—has emerged, completing the cross-integration of multiple disciplines and technologies such as data mining, spatial databases, statistics, 3S (spatial information, data storage, and image processing), and machine learning.
[0003] Isotopic pattern mining is an emerging research hotspot in spatial data mining. The concept originates from ecology: isotopic patterns are subsets of spatial features whose instances frequently lie within their spatial neighborhoods. For example, ecologists identify the frequent coexistence of Nile crocodiles and Egyptian plovers; based on this isotopic rule, they can predict that Egyptian plovers exist in areas inhabited by Nile crocodiles. Spatial isotopic patterns can provide important insights for many applications. For instance, criminologists might be interested in high-risk facilities (for a group of facilities, such as bars, parking lots, and supermarkets, where a small portion of the facilities cover the majority of crime), obtaining isotopic patterns based on the spatial adjacency of crime locations with various facilities to support decision-making in urban crime prevention. Other application areas include earth sciences, public health, biology, and transportation.
[0004] Improvements to existing homonym mining methods primarily focus on two aspects: algorithm efficiency and research scope. Regarding algorithm efficiency, efforts are made to reduce computational load and increase candidate example generation efficiency, including methods based on partial connections, connectionsless connections, and cliques and trees. In terms of research scope, the scope of homonym mining is expanded to include one-dimensional and two-dimensional objects, and spatial location-based homonym mining is extended to spatiotemporal dimensions. However, all these studies presuppose an important premise: that the research area is homogeneous.
[0005] In summary, existing technologies have obvious defects and shortcomings. Taking urban facilities as spatial features, without considering other factors such as facility size, management method, operating hours, and target audience, instances of the same type of facility in different areas are not different. This may lead to problems such as "variable surface" and "ecological fallacy". Summary of the Invention
[0006] To address the aforementioned problems, this invention provides a method and apparatus for mining urban point of interest co-location patterns. Considering the shortcomings of existing co-location pattern mining, this invention proposes to mine co-location patterns at the scale of urban functional zones, taking into account the characteristics of urban functional distribution.
[0007] A method for mining urban point of interest co-location patterns, the method comprising:
[0008] Urban functional zones are identified using a topic model based on the movement trajectories of urban residents.
[0009] Given a city road network, each POI is projected onto the road network. The distance between POIs is calculated by finding the shortest connection path and stored in the distance matrix to establish the network distance matrix between city POIs.
[0010] Based on the network distance matrix between POIs in the urban road network, we perform POI co-location pattern mining within various functional zones to identify co-location patterns.
[0011] Furthermore, the identification of urban functional zones based on the movement trajectories of urban residents using a topic model specifically includes:
[0012] Generate unit analysis documents for each analysis unit based on the movement trajectory, and generate metadata for the documents based on the number of various POI facilities in the unit analysis documents;
[0013] The DMR topic model is used to identify urban functional areas. The number of topic models is determined by a combination of perplexity and topic consistency index. Perplexity is the uncertainty of whether a model trained in text analysis can identify which topics a document contains.
[0014] Functional areas are identified by combining representative vocabulary, frequency density, internal ranking, and POIs for each theme.
[0015] Furthermore, the step of generating unit analysis documents for each analysis unit based on the movement trajectory specifically includes:
[0016] A dataset is established based on the movement trajectory data of urban residents. The movement trajectory data includes departure time, arrival time, departure point and destination.
[0017] Documents are constructed for each region based on the dataset. The words in the documents are in the form of Symbol_Date_Moment, where Symbol represents the arrival or departure trajectory, Date represents weekdays or holidays, and Moment represents the arrival or departure time period.
[0018] Furthermore, the step of generating metadata for the document based on the number of various POI facilities in the unit analysis document specifically includes:
[0019] area The metadata for the corresponding document is the frequency density of POIs, which is represented as: , where F is the number of categories of POI, and for Frequency density of the i-th POI category Calculated using the following formula:
[0020]
[0021] in For the region The number of facilities of type i in the middle For the region The area.
[0022] Furthermore, the use of DMR topic models for urban functional area identification specifically includes:
[0023] Based on the DMR topic model, input documents and metadata from each region, and select the number of topics according to perplexity and topic consistency indicators;
[0024] The DMR topic model is trained to obtain representative words and their corresponding probabilities under each topic. The POIs and areas of each region are statistically analyzed, the frequency density is calculated, and the internal ranking is obtained.
[0025] In each region, the frequency density of different POIs is calculated, which is the ratio of the number of POI facilities in the region to the area of the region, and the ranking of the region is calculated based on the frequency density.
[0026] Furthermore, the functional area identification, which combines representative vocabulary, frequency density, internal ranking, and POIs for each topic, specifically includes:
[0027] The region is described from the perspective of POI by frequency density and internal ranking; the region's functions are reflected from a dynamic perspective by taxi trajectory data; the region's functions are reflected from a static perspective by the proportion of various types of facilities in the region; and the functions of each urban functional area are identified from the perspectives of POI, dynamic perspective, and static perspective.
[0028] Furthermore, based on the network distance matrix between POIs in the urban road network, POI co-location pattern mining is performed within various functional zones, specifically including:
[0029] Based on the network distance matrix between city POIs, spatial neighborhood relationships of spatial element instances are constructed within different functional zones to generate candidate spatial co-location pattern instances.
[0030] Based on candidate spatial co-occurrence pattern instances, the participation degree of candidate spatial co-occurrence patterns is calculated, and spatial co-occurrence patterns with participation exceeding a given threshold are extracted.
[0031] Furthermore, the calculation of the participation degree of the candidate spatial co-location pattern specifically includes:
[0032] For N objects in the space, each object has two possibilities in the candidate patterns: it appears or it does not appear. The number of candidate patterns is 2^N. N One; use participation to determine whether candidate patterns form frequent patterns; the participation formula is:
[0033]
[0034] In the formula, For specific spatial objects, A spatial alignment pattern consisting of k spatial objects. For relational projection, For table instances; if a set of unique spatial instances are close to each other and contain... An instance in the array is called a row instance. All row instances are A table instance.
[0035] Furthermore, the step of constructing spatial neighborhood relationships for spatial element instances within different functional zones based on the network distance matrix between city POIs, and generating candidate spatial co-location pattern instances, specifically includes:
[0036] The spatial dataset is materialized into a star-shaped neighbor materialization model: for each instance, instances within a specified threshold range that have star-shaped neighbor relationships are identified, each instance is taken as the center instance, and a set is formed with the neighbor instances that have star-shaped neighbor relationships, and the feature type of the neighbor instances is required to be larger than that of the center instance in lexicographical order.
[0037] Generate k+1 order candidate spatial isomorphic patterns based on the generated k-order frequent spatial isomorphic patterns.
[0038] Generate star instances of candidate spatial isomorphic patterns from the star neighbor materialization model;
[0039] Generate coarse frequent spatial co-location patterns using star instances of spatial co-location patterns, i.e., calculate whether each spatial object in the spatial co-location pattern has reached the participation threshold.
[0040] For third-order or higher star schemas, check if the star schema is a clique instance; delete instances that are not clique. Second-order star schemas are clique instances, so there is no need to check if they are clique.
[0041] The patterns that meet the above filtering conditions are frequent co-location patterns, and all spatial co-location patterns of order k+1 are generated.
[0042] Iterate until no new candidate homonyms are generated.
[0043] A device for mining urban point of interest co-location patterns includes: a functional area identification unit, a network distance matrix establishment unit, and a co-location pattern identification unit connected in sequence;
[0044] The functional area identification unit is used to identify urban functional areas based on the movement trajectories of urban residents using a topic model.
[0045] The network distance matrix building unit is used to project each POI onto the road network given a city road network, calculate the distance between POIs by finding the shortest connection path, and store it in the distance matrix to build the network distance matrix between city POIs.
[0046] The same-position pattern recognition unit is used to mine the same-position patterns of POIs in various functional areas based on the network distance matrix between POIs in the urban road network, so as to identify the same-position patterns.
[0047] An electronic device includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;
[0048] Memory, used to store computer programs;
[0049] When the processor executes the program stored in memory, it implements the above-mentioned method for mining urban point of interest co-location patterns.
[0050] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method for mining urban point of interest co-location patterns.
[0051] The present invention has at least the following beneficial effects:
[0052] Compared to traditional isotope pattern mining, this invention takes into account the existence of spatial heterogeneity and performs isotope pattern mining under the premise of functional area identification, thus reflecting the actual situation more comprehensively and accurately.
[0053] In the process of constructing spatial neighborhood relationships, this invention uses network distance instead of Euclidean distance to measure the spatial proximity between instances, which is a more practical and effective measure.
[0054] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures pointed out in the description and the drawings. Attached Figure Description
[0055] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0056] Figure 1 This is a flowchart of the mining method according to an embodiment of the present invention;
[0057] Figure 2 This is a diagram illustrating spatial proximity.
[0058] Figure 3 This is a schematic diagram of the excavation device structure according to an embodiment of the present invention;
[0059] Figure 4 This is a schematic diagram of the functional area identification results based on the DMR topic model;
[0060] Figure 5 This diagram illustrates the impact of network distance and Euclidean distance metrics on the number of adjacency relationships. Detailed Implementation
[0061] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0062] Existing technologies, when considering urban facilities as spatial features, often fail to account for variations in facility size, management methods, operating hours, target audience, and other factors. This results in inconsistencies in instances of the same type of facility across different regions, potentially leading to issues such as "variable facets" and "ecological fallacies." To address this, this invention proposes a method and apparatus for mining urban point-of-interest (POI) co-location patterns, comprising a method, an apparatus, an electronic device, and a computer-readable storage medium.
[0063] Compared to traditional isotope pattern mining, this invention takes into account the existence of spatial heterogeneity and performs isotope pattern mining under the premise of functional area identification, thus reflecting the actual situation more comprehensively and accurately.
[0064] Firstly, such as Figure 1 As shown, this invention provides a method for mining urban point-of-interest co-location patterns, the method comprising:
[0065] Urban functional zones are identified using a topic model based on the movement trajectories of urban residents.
[0066] Given a city road network, each POI (generally an abbreviation for Point of Interest, also known as Point of Information, usually referred to as a point of interest, referring to data that can be abstracted as points in Internet electronic maps) is projected onto the road network. The distance between POIs is calculated by finding the shortest connecting path and stored in the distance matrix to establish the network distance matrix between city POIs.
[0067] Based on the network distance matrix between POIs in the urban road network, we perform POI co-location pattern mining within various functional zones to identify co-location patterns.
[0068] In this embodiment, the process of identifying urban functional zones using a topic model based on the movement trajectories of urban residents specifically includes:
[0069] Generate unit analysis documents for each analysis unit based on the movement trajectory, and generate metadata for the documents based on the number of various POI facilities in the unit analysis documents;
[0070] The DMR topic model is used to identify urban functional areas. The number of topic models is determined by a combination of perplexity and topic consistency index. Perplexity refers to the uncertainty in text analysis regarding which topics a trained model contains.
[0071] Functional areas are identified by combining representative vocabulary, frequency density, internal ranking, and representative POIs for each theme.
[0072] In this embodiment, generating unit analysis documents for each analysis unit based on the movement trajectory specifically includes:
[0073] A dataset is established based on the movement trajectory data of urban residents. The movement trajectory data includes departure time, arrival time, departure point and destination.
[0074] Documents are constructed for each region based on the dataset. The words in the documents are in the form of Symbol_Date_Moment, where Symbol distinguishes between arrival and departure trajectories, Date distinguishes between weekdays and holidays, and Moment represents the time period of arrival or departure.
[0075] In this embodiment, the step of generating metadata for the document based on the number of various POI facilities in the unit analysis document specifically includes:
[0076] area The metadata of the corresponding document is the frequency density of POIs, represented as , where F is the number of categories of POI, and for Frequency density of the i-th POI category Calculated using the following formula:
[0077]
[0078] in For the region The number of facilities of type i in the middle For the region The area.
[0079] In this embodiment, the use of DMR topic model for urban functional area identification specifically includes:
[0080] Based on the DMR topic model, input documents and metadata from each region, and select the number of topics according to perplexity and topic consistency indicators;
[0081] The DMR topic model is trained to obtain representative words and their corresponding probabilities for each topic. The POIs and areas of each region are statistically analyzed, frequency density is calculated, and internal rankings are obtained.
[0082] Within each region, calculate the frequency density of different Points of Interest (POIs), which is the ratio of the number of POI facilities in the region to the area of that region. Rank the regions based on their frequency density.
[0083] For example, if an entertainment area contains three facilities—a bar, a supermarket, and a library—with frequency densities of 0.5, 0.3, and 0.2 respectively, then their internal rankings would be 1, 2, and 3.
[0084] In this embodiment, the functional area identification, which combines representative vocabulary, frequency density, internal ranking, and representative POIs for each topic, specifically includes:
[0085] Frequency density and internal ranking describe a region from the perspective of POI; taxi trajectory data reflects the region's functions from a dynamic perspective; and the proportion of different types of facilities in a region reflects its functions from a static perspective.
[0086] The functional identification results for each urban functional area are completed from the perspectives of POI, dynamic perspective, and static perspective.
[0087] In this embodiment, based on the network distance matrix between POIs in the urban road network, POI co-location pattern mining is performed within various functional zones, specifically including:
[0088] Based on the network distance matrix between city POIs, spatial neighborhood relationships of spatial element instances are constructed within different functional zones to generate candidate spatial co-location pattern instances.
[0089] Based on candidate spatial co-occurrence pattern instances, the participation degree of candidate spatial co-occurrence patterns is calculated, and spatial co-occurrence patterns with participation exceeding a given threshold are extracted.
[0090] In this embodiment, the calculation of the participation degree of the candidate space co-position pattern specifically includes:
[0091] For N objects in the space, each object has two possibilities in the candidate patterns: it appears or it does not. Therefore, the number of candidate patterns is 2^N. N One; use participation to determine whether candidate patterns form frequent patterns; the participation formula is:
[0092]
[0093] In the formula, For specific spatial objects, A spatial alignment pattern consisting of k spatial objects. It is a relational projection. For table instances; if a set of unique spatial instances are close to each other and contain... An instance in the array is called a row instance, and All row instances are A table instance;
[0094] When PI(c) >= min_prev, where min_prev is the minimum engagement threshold given by the user, then the same-position pattern c is said to be frequent.
[0095] In this embodiment, the step of constructing spatial neighborhood relationships of spatial element instances within different functional zones based on the network distance matrix between city POIs and generating candidate spatial co-location pattern instances specifically includes:
[0096] The spatial dataset is materialized into a star-shaped neighbor materialization model: for each instance, instances within a specified threshold range that have star-shaped neighbor relationships are identified, each instance is taken as the center instance, and a set is formed with the neighbor instances that have star-shaped neighbor relationships, and the feature type of the neighbor instances is required to be larger than that of the center instance in lexicographical order.
[0097] Based on the generated k-order frequent spatial co-location patterns, generate k+1-order candidate spatial co-location patterns, where the 1-order co-location patterns (each spatial object) are frequent and are used as initialization.
[0098] Generate star instances of candidate spatial isomorphic patterns from the star neighbor materialization model;
[0099] Generate coarse frequent spatial co-location patterns using star instances of spatial co-location patterns, i.e., calculate whether each spatial object in the spatial co-location pattern has reached the participation threshold.
[0100] For order 3 or higher, it is necessary to check whether the star instance is a clique instance, that is, whether each pair of them is adjacent. Instances that are not clique are deleted. Since the spatial proximity relationship is symmetrical, the order 2 star instance is a clique instance and there is no need to check whether it is clique.
[0101] The patterns that meet the above filtering conditions are frequent co-location patterns, and all spatial co-location patterns of order k+1 are generated.
[0102] Iterate until no new candidate homonyms are generated.
[0103] In specific implementation, such as Figure 2 As shown, spatial proximity describes the spatial relationship between spatial instances. This can be topological, distance-based, or a hybrid relationship. Here, proximity is a distance relationship; for example, B2 and B5 are two instances with a proximity relationship because the distance between them is less than the proximity distance threshold.
[0104] The star-shaped neighbor materialization model materializes the spatial proximity relationships of a spatial dataset using star-shaped proximity relationships. Simply put, the user specifies a proximity distance threshold as the radius to draw a circle; instances within the circle have star-shaped proximity relationships. Instances with star-shaped proximity relationships to the center instance are highlighted with solid black lines. For example, A1, B1, and C1 in the diagram have star-shaped proximity relationships. Here, A1 is the center instance, and its neighbors B1 and C1 within the circle are its neighbor instances.
[0105] Star-shaped instances are the results of filtering candidate pattern instances using the star-shaped neighbor materialization model. For example, for the candidate co-occurring pattern {A,B}, if A has the smallest lexicographical order, then the instances of this co-occurring pattern are collected from the star-shaped neighbors of feature A: {A1,B1}, {A2,B4}, {A3,B3}.
[0106] Based on whether each spatial object in the star schema instance spatial primacy pattern reaches the participation threshold, a coarse filtering can be performed on the candidate spatial primacy patterns to obtain a rough spatial primacy pattern. In the candidate primacy pattern {A,B}, object A contains instances A1, A2, and A3, accounting for 3 / 4 of the total number of instances, meaning object A's participation is 0.75; object B contains B1, B3, and B4, accounting for 3 / 5 of the total number of instances, meaning object B's participation is 0.6. When the threshold is set to 0.6 or below, the candidate primacy pattern passes the coarse filtering and can proceed to the next step of selection.
[0107] This filtering step checks whether star instances are cluster instances, meaning whether all pairs of them are adjacent. Assuming the candidate co-occurrence pattern {A,B,C} passes the coarse filtering, {A1,B1,C1}, {A2,B4,C2}, and {A3,B3,C1} are star instances of this pattern. Since B1 and C1 are not adjacent, the instance {A1,B1,C1} needs to be deleted.
[0108] Secondly, such as Figure 3 As shown, the present invention provides an urban point of interest co-location pattern mining device, comprising: a functional area identification unit, a network distance matrix establishment unit, and a co-location pattern identification unit;
[0109] The functional area identification unit is used to identify urban functional areas based on the movement trajectories of urban residents using a topic model.
[0110] The network distance matrix building unit is used to project each POI onto the road network given a city road network, calculate the distance between POIs by finding the shortest connection path, and store it in the distance matrix to build the network distance matrix between city POIs.
[0111] The same-position pattern recognition unit is used to mine the same-position patterns of POIs in various functional areas based on the network distance matrix between POIs in the urban road network, so as to identify the same-position patterns.
[0112] Thirdly, the present invention provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;
[0113] Memory, used to store computer programs;
[0114] When the processor executes the program stored in memory, it implements the above-mentioned method for mining urban point of interest co-location patterns.
[0115] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the above-described method for mining urban point of interest co-location patterns.
[0116] The computer-readable storage medium may be included in the device / apparatus described in the above embodiments; or it may exist independently and not assembled into the device / apparatus. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.
[0117] According to embodiments of this disclosure, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0118] To enable those skilled in the art to better understand the present invention, the principles of the present invention are explained below in conjunction with the accompanying drawings:
[0119] Improvements to existing isomorphic pattern mining methods primarily focus on two aspects: algorithm efficiency and research scope. Regarding algorithm efficiency, efforts are made to reduce computational load and increase candidate example generation efficiency, including methods based on partial connections, connectionsless connections, and cliques and trees. In terms of research scope, the scope of isomorphic pattern mining is expanded to include one-dimensional and two-dimensional objects, and spatial location-based isomorphic pattern mining is extended to spatiotemporal dimensions. However, these studies all assume a crucial premise: the research area is homogeneous. In other words, taking urban facilities as a spatial feature, without considering factors such as facility size, management methods, operating hours, and object orientation, instances of the same type of facility in different areas will not be differentiated. This could lead to problems such as "variable facets" and the "ecological fallacy."
[0120] This invention is a method for mining urban point of interest co-location patterns that takes into account functional distribution characteristics.
[0121] Point of Interest (POI) is a concept in Geographic Information Systems (GIS), referring to geographic objects that can be abstracted into points, especially geographic entities closely related to people's lives, such as schools, banks, restaurants, gas stations, hospitals, supermarkets, and bus stops. It also includes abstract concepts such as hotspots of human activity and crime. In cities, exploring the implicit spatial configurations between POIs, such as the frequent co-occurrence of schools and robberies, or restaurants and bus stops, helps to understand some problems in urban development and provides theoretical support. For example, we can reduce urban robberies by strengthening surveillance near schools.
[0122] Traditional spatial co-location pattern mining is based on the Euclidean spatial assumption. Researchers assume that the patterns of interest occur within a uniform plane, and the spatial proximity between two instances is measured by the Euclidean straight-line distance. However, this method is unsuitable because when studying city POIs, most instances are confined to a planar network. Therefore, network distance is a more practical and effective metric. The technical approach for city POI co-location pattern mining is as follows:
[0123] 1. Establish a network distance matrix between city POIs: Given a city road network, project each POI onto the road network, calculate the distance between POIs by finding the shortest connection path, and store it in the distance matrix to facilitate the determination of proximity relationships in the same position pattern and reduce the amount of computation.
[0124] 2. Urban POI co-location pattern mining is conducted within different functional zones. Compared with global co-location pattern mining, this invention can take into account the environmental characteristics behind the patterns. For example, restaurants in residential areas may mainly serve the surrounding residents, while restaurants in entertainment areas have the function of attracting crowds.
[0125] This invention addresses the shortcomings of existing isomorphic pattern mining methods by proposing isomorphic pattern mining at the urban functional area scale. The main technical solution includes the following steps:
[0126] I. Identification of Urban Functional Zones Based on Dirichlet-Multinomial Regression (DMR) Model
[0127] Similar to LDA topic modeling, DMR topic modeling is a document generation model. It assumes that an article has multiple topics, each corresponding to different words. The article construction process first selects a topic with a certain probability, then selects a word under that topic with a certain probability, iteratively generating the vocabulary in the document, and finally generating the entire article. Compared to LDA, DMR topic modeling adds metadata as a hyperparameter to modify the log-linear prior of the topic document distribution. Therefore, based on DMR topic modeling, the identification of urban functional areas follows these three steps:
[0128] 1. Generate unit analysis documents for each analysis unit based on the movement trajectory, and generate metadata for the documents based on the number of various POI facilities in the unit analysis documents;
[0129] 2. DMR topic models are used for urban functional area identification. The number of topic models is determined by a combination of confusion level and topic consistency index.
[0130] 3. Identify functional areas by combining representative vocabulary, frequency density, internal ranking, and characteristic POIs for each theme.
[0131] To identify urban functional zones using the DMR topic model, the first step is to define the analysis unit, obtain the various regions within that unit, and generate documents and their metadata for each region. Document generation is based on the movement trajectories of urban residents, including parking record data, taxi trajectory record data, and SCD (Smart Card Data). Taking taxi trajectory data as an example, the data should include pick-up and drop-off times and locations. Documents are constructed for each region based on the dataset. The words in these documents are in the form of Symbol_Date_Moment, where Symbol distinguishes between arrival and departure trajectories, Date distinguishes between weekdays and holidays, and Moment represents the arrival (departure) time period.
[0132] area The metadata of the corresponding document is the frequency density of POIs, represented as , where F is the number of categories of POI, and for Frequency density of the i-th POI category Calculated using the following formula:
[0133]
[0134] in For the region The number of facilities of type i in the middle For the region The area.
[0135] The DMR topic model takes documents and metadata from each region as input and selects the number of topics based on perplexity and topic consistency metrics. Perplexity, as the reciprocal of sentence probability, theoretically indicates a better topic model fit with lower perplexity. Topic consistency measures the coherence of words within a topic; theoretically, greater topic coherence leads to a better topic fit. Here, the number of topics is comprehensively evaluated based on these two metrics.
[0136] The model is trained to obtain representative words and their corresponding probabilities for each topic. The POIs and areas of each region are statistically analyzed, and their frequency densities are calculated. Then, within each region, the internal ranking of different POI frequency densities is calculated.
[0137] Functional areas are identified by combining representative keywords, frequency density, internal ranking, and representative Points of Interest (POIs). The representative keywords describe the areas based on taxi trajectories, while frequency density and internal ranking describe the areas from the perspective of POIs. Generally, knowing the proportion of each type of facility in an area allows for a static assessment of its function, while taxi trajectory data provides a dynamic assessment. Additionally, for areas that are difficult to understand or distinguish, representative POIs may be used for supplementary identification.
[0138] II. Join-less Co-location Pattern Mining Based on Network Distance to Identify City POI Co-location Patterns in Different Functional Zones
[0139] The spatial co-location pattern mining method mainly includes two steps: ① constructing spatial neighborhood relationships of spatial feature instances to generate candidate spatial co-location pattern instances; ② calculating the participation degree of candidate spatial co-location patterns and extracting spatial co-location patterns with participation degrees exceeding a given threshold.
[0140] Step 1: Calculate network distances within the road network. Compared to Euclidean distance, using network distance to measure the spatial proximity between POIs can reduce interference from non-existent related patterns in co-location pattern mining. To increase efficiency, the calculated network distances between POIs are stored in a matrix. The calculation process consists of two steps: moving the POI locations onto the road network. Since most POIs in a city may not accurately fall on the road network, and network distance calculation is based on the shortest path on the road network, it is necessary to project the POIs onto the road network; calculating the network distances between POIs and storing them in a symmetric matrix.
[0141] In different functional areas, same-location pattern mining based on network distance becomes extremely computationally expensive as the order increases (i.e., the number of POIs included in the pattern increases). Therefore, we will adopt connectionless same-location pattern mining, and the mining process consists of the following seven steps:
[0142] 1. Materialize the spatial dataset into a star-shaped neighbor materialization model: Instances within a specified threshold range have star-shaped neighbor relationships. Secondly, the star-shaped neighbors of each instance are a set consisting of the central instance and neighbor instances that are adjacent to it. It is required that the feature type of the neighbor instances is larger than that of the central instance in the lexicographical order.
[0143] 2. Generate k+1 order candidate spatial symmetry patterns based on the generated k-order frequent spatial symmetry patterns, where the first-order symmetry patterns are all frequent and are used as initialization.
[0144] 3. Generate star instances of candidate spatially corresponding patterns from the star neighbor set;
[0145] 4. Generate coarsely filtered candidate spatial alignment patterns using star instances of spatial alignment patterns, i.e., calculate whether each spatial object in the spatial alignment pattern reaches the participation threshold.
[0146] 5. For star schemas of order 3 or higher, it is necessary to check whether the star schema instances are clique instances, i.e., whether every pair of adjacent instances is a clique instance. Instances that are not clique are deleted. Since spatial proximity is symmetrical, a star schema of order 2 is a clique instance, and there is no need to check whether it is clique.
[0147] 6. Generate frequent spatial co-location patterns;
[0148] 7. Iterate through steps 2-6 until no new candidate homonyms are generated.
[0149] Step two, the concept of engagement is similar to support in data mining, and its formula is:
[0150]
[0151] In the formula, For a specific spatial object, c represents a spatial alignment pattern composed of k spatial objects. It is a relational projection. For a table instance, a row instance represents a set of instances that contains all features in c, and a subset of those instances does not contain all features in c. The set of all row instances in equation c is a table instance. Participation is defined to calculate the minimum participation of each object in a pattern, thereby determining whether each pattern occurs frequently.
[0152] The connectionless method introduces the concept of a star-shaped neighbor materialization model: spatial proximity relationships in a spatial dataset are materialized using star-shaped proximity relations, meaning instances within a specified threshold range have star-shaped proximity relationships. Furthermore, each instance's star-shaped neighbors consist of a central instance and its neighboring instances, requiring that the feature types of the neighboring instances are lexicographically larger than those of the central instance. This ensures that when calculating the adjacency relationships of instances, each instance is close to its star-shaped neighbors, avoiding redundant calculations.
[0153] In a scenario where New York City is the analysis area and taxi areas are the analysis units, functional area identification is as follows: Figure 4 As shown.
[0154] The number of topic models was set to 9 based on the combined confusion level and topic consistency index. Functional areas were identified according to the topic's representative words, frequency density, internal ranking, and representative POI facilities, and were divided into entertainment area, residential area, public area, central business district, municipal area, office area, commuter area, commercial area, and airport.
[0155] Identify POI co-location patterns in different functional zones:
[0156] Table 1. City POI Co-location Patterns under Different Functional Zones (Only Co-location Patterns Including Category A POIs are Shown)
[0157]
[0158] Among them, A: property crimes; B: schools; C: tourist attractions; D: hospitals; E: bars; F: hotels; G: vending machines; H: parks; I: cinemas; J: police departments; K: fire stations; L: banks; M: supermarkets; N: restaurants; O: sports fields; P: museums; Q: cafes; R: shops.
[0159] Crime, as an abstract POI, is represented by points containing location coordinates and attributes, consistent with other entity POI types. In the table above, the co-occurrence pattern is shown as AB, indicating that POIs A and B frequently co-occur. The participation level is consistent with what was described above, representing the minimum participation of A and B in the AB pattern.
[0160] Taking functional zones into account, urban POI isomorphism mining takes into account spatial heterogeneity: compared to isomorphism mining at the city scale, more different patterns, even patterns of different orders, can be identified in some functional zones; even for the same isomorphism, the level of participation varies in different functional zones; and for isomorphism obtained in a specific functional zone, we can also analyze it in conjunction with the specific attributes of that functional zone. Therefore, urban POI isomorphism mining technology that takes functional zones into account is effective and important.
[0161] Secondly, by using network distance instead of Euclidean distance to measure the spatial proximity between instances, the adjacency relationship between instances is reduced, thus improving computational efficiency and the accuracy of co-location patterns.
[0162] This article uses 23,726 POIs, such as Figure 5 As shown in the figure, the upper curve represents Euclidean distance, and the lower curve represents network distance. Within a distance threshold of 100m-1000m, the network distance metric significantly reduces spatial adjacency relationships. In constructing the star-shaped neighbor materialization model, it reduces adjacency relationships by 66.08%-98.69%, greatly improving computational efficiency and avoiding the generation of non-existent association patterns that could affect the accuracy of the mining results.
[0163] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for urban point of interest co-locating pattern mining, characterized in that, The method comprises: Based on the mobile trajectory of urban residents, the city functional area is identified using a topic model; In the case of a given city road network, each POI is projected onto the road network, the distance between POIs is calculated by finding a shortest connection path, and is stored in a distance matrix to establish a network distance matrix between the city POIs; Based on the network distance matrix between the POIs in the city road network, the co-occurrence pattern of POIs is mined within various functional areas to identify co-occurrence patterns; Based on the network distance matrix between the POIs in the city road network, the co-occurrence pattern of POIs is mined within various functional areas, specifically including: Based on the network distance matrix between the POIs in the city, the spatial neighborhood relationship of spatial element instances is constructed within different functional areas to generate candidate spatial co-occurrence pattern instances; Based on the candidate spatial co-occurrence pattern instances, the participation of the candidate spatial co-occurrence pattern is calculated, and the spatial co-occurrence pattern whose participation exceeds a given threshold is extracted; Based on the network distance matrix between the POIs in the city, the spatial neighborhood relationship of spatial element instances is constructed within different functional areas to generate candidate spatial co-occurrence pattern instances, specifically including: Materialize the spatial data set into a star neighbor materialization model: for each instance, specify the instances within the threshold range that have a star neighbor relationship with it, and group the neighbor instances with the center instance to form a set, and require the feature type of the neighbor instance to be larger than the center instance in lexicographic order; Generate k+1 order candidate spatial co-occurrence patterns according to the generated k order frequent spatial co-occurrence patterns; Generate star instances of candidate spatial co-occurrence patterns from the star neighbor materialization model; Use the star instances of spatial co-occurrence patterns to generate rough frequent spatial co-occurrence patterns, that is, calculate whether each spatial object in the spatial co-occurrence pattern reaches the participation threshold; For three or higher orders, check whether the star instance is a team instance, and delete the instances that do not form a team; the second order star instance is a team instance, and does not need to be checked whether it forms a team; The ones that meet the filtering conditions are frequent co-occurrence patterns, and all spatial co-occurrence patterns of k+1 order are generated; Iterate until no new candidate co-occurrence patterns are generated.
2. The city interest point co-occurrence pattern mining method according to claim 1, wherein: The city functional area is identified based on the mobile trajectory of urban residents using a topic model, specifically including: Generate a unit analysis document for each analysis unit according to the mobile trajectory, and generate metadata of the document according to the number of POIs in the unit analysis document; Identify the city functional area using a DMR topic model, and determine the number of topic models by comprehensively considering the perplexity and topic consistency indicators; the perplexity is the uncertainty of the model trained in the text analysis in identifying whether some documents contain which topics; Identify the functional area in combination with the topic representative words, frequency density, internal ranking, and POIs.
3. The city interest point co-occurrence pattern mining method according to claim 2, wherein: The unit analysis document is generated for each analysis unit according to the mobile trajectory, specifically including: Based on the mobile trajectory data of urban residents, a dataset is established, and the mobile trajectory data includes departure time, arrival time, departure place and destination; According to the dataset, a document is constructed for each region, and the form of the words in the document is Symbol_Date_Moment, wherein Symbol represents a trajectory of arrival or departure, Date represents a working day or a holiday, and Moment represents an arrival or departure time period.
4. The city interest point co-location pattern mining method according to claim 2 or 3, characterized in that, the unit for generating metadata of the document according to the number of POIs of each type in the document specifically includes: Region The metadata corresponding to the document is the frequency density of POIs, denoted as where F is the number of categories of POIs, and for the frequency density of the i-th POI category in is calculated by the following equation: wherein is the area of the region is the number of i-th type of setting in the region is the area of the region is the area of the region 5. The city interest point co-location pattern mining method according to claim 2, characterized in that, the city functional area identification by using the DMR topic model specifically includes: inputting the documents and metadata of each region into the DMR topic model, and selecting the number of topics according to perplexity and topic consistency indexes; training the DMR topic model to obtain representative words and corresponding probabilities under each topic, and statistically analyzing POIs and areas of each region to calculate frequency density and obtain internal ranking; calculating the frequency density of different POIs under each region, and calculating the intra-regional ranking according to the frequency density in each region.
6. The city interest point co-location pattern mining method according to claim 2, characterized in that, the functional area identification by combining the representative words of each topic, frequency density, internal ranking and POIs specifically includes: describing the region from the POI perspective by frequency density and internal ranking, reflecting the function of the region from the dynamic perspective by taxi trajectory data, reflecting the function of the region from the static perspective by the proportion of each type of facility in the region, and completing the functional identification from the POI perspective, dynamic perspective and static perspective for each city functional area.
7. The city interest point co-location pattern mining method according to claim 1, characterized in that, the calculation of the participation degree of the candidate spatial co-location pattern specifically includes: For N objects in space, each object has two cases in candidate pattern, which is appeared and not appeared, the number of candidate pattern is 2 N N; using participation to judge whether the candidate pattern forms frequent pattern; the participation formula is: wherein, is a concrete spatial object, is a spatial isomode consisting of k spatial objects, is a relation projection, is a table instance; if a set of non-repeating spatial instances are close to each other and contain instances, then it is called a row instance, all row instances of are table instances of 8. An urban point of interest co-located pattern mining apparatus, comprising: including: a functional area identification unit, a network distance matrix establishment unit and a co-location pattern identification unit connected in sequence; the functional area identification unit is configured to identify city functional areas by using a topic model based on mobile trajectory data of urban residents; the network distance matrix establishment unit is configured to project each POI onto a road network under a given city road network, calculate the distance between POIs by finding a shortest connection path, and store the distance in a distance matrix to establish a network distance matrix between city POIs; the co-location pattern identification unit is configured to mine co-location patterns of POIs within various functional areas based on the network distance matrix between POIs in the city road network to identify co-location patterns; mining co-location patterns of POIs within various functional areas based on the network distance matrix between POIs in the city road network specifically includes: constructing spatial neighborhood relationships of spatial element instances within different functional areas based on the network distance matrix between city POIs to generate candidate spatial co-location pattern instances; Based on the candidate spatial co-location pattern instances, the participation of the candidate spatial co-location pattern is calculated, and the spatial co-location pattern whose participation exceeds a given threshold is extracted; Based on the network distance matrix between the city POIs, the spatial neighborhood relationship of the spatial feature instances is constructed within the range of different functional areas, and candidate spatial co-location pattern instances are generated, specifically including: Materializing the spatial data set into a star neighbor materialization model: for each instance, specify the instances within the threshold range that have star neighbor relationship with it, and group the instances as center instances with neighbor instances having star neighbor relationship, and require the feature type of the neighbor instances to be larger than the center instance in lexicographic order; According to the generated k-order frequent spatial co-location pattern, a k+1-order candidate spatial co-location pattern is generated; The star instance of the candidate spatial co-location pattern is generated from the star neighbor materialization model; The rough frequent spatial co-location pattern is generated using the star instance of the spatial co-location pattern, that is, whether each spatial object in the spatial co-location pattern reaches the participation threshold is calculated; For three orders or higher orders, check whether the star instance is a team instance, and delete the instances that are not in the team; the second-order star instance is a team instance, and there is no need to check whether it is in the team; The spatial co-location pattern that meets the filtering condition is the frequent co-location pattern, and all k+1-order spatial co-location patterns are generated; Iterate until no new candidate co-location pattern is generated.
9. An electronic device, comprising: The computer program is executed by the processor to realize the city interest point co-location pattern mining method in any one of claims 1-7. The computer program is executed by the processor to realize the city interest point co-location pattern mining method in any one of claims 1-7. 10. A computer-readable storage medium having stored thereon a computer program, characterized in that,