A method for identifying, classifying, and verifying urban public activity centers based on multi-source data.

By combining multi-source data with the Getis-Ord Gi* index method and information entropy theory model, and dynamically verifying the data based on commercial facility building volume and OD data, the accuracy problem of identifying and classifying urban public activity centers in existing technologies has been solved, thus achieving refined urban construction management.

CN115905934BActive Publication Date: 2026-08-04SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SOUTH CHINA UNIV OF TECH
Filing Date
2022-04-18
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing methods for identifying urban public activity centers lack self-verification and fail to accurately identify and classify them from a spatiotemporal four-dimensional perspective.

Method used

We employ multi-source data combined with the Getis-Ord Gi* index method based on local spatial autocorrelation to conduct facility clustering analysis. We also use the information entropy theory model to calculate the degree of mixing, identify public activity centers by superimposing commercial facility building volumes, and perform dynamic verification through OD data. This results in the construction of a spatiotemporal index measurement system that includes preliminary identification, supervised classification, and secondary verification.

Benefits of technology

It enables accurate identification and classification of urban public activity centers, promotes the refinement of urban construction and management, and improves the reliability and accuracy of identification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115905934B_ABST
    Figure CN115905934B_ABST
Patent Text Reader

Abstract

This invention discloses a method for identifying, classifying, and verifying urban public activity centers based on multi-source data. The method includes: acquiring data from a target area through a multi-source data open platform; dividing and numbering the target area into grids based on spatial analysis methods such as GIS spatial autocorrelation analysis, and constructing an identification index system for urban public activity centers based on the acquired data; preprocessing the data from the target area, including data cleaning and coordinate unification; linking the divided grids according to the preprocessed data, statistically analyzing the values ​​of each influencing factor in the index system, and initially identifying urban public activity centers; constructing a classification index system based on Dianping (a Chinese review platform) data to initially identify the functions of the activity centers; constructing a verification index system for urban public activity centers based on dynamic data, further verifying the identified grids, and obtaining the final identification result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for identifying, classifying, and verifying urban public activity centers based on multi-source data. Background Technology

[0002] The urban public activity center, as described in this invention, is a key focus in evaluating urban spatial structure and urban spatial planning. Public activity centers host the most frequent public activities and social life of citizens, serving as core areas of the urban structure and important components of urban functions. They are also crucial urban spaces for safeguarding people's aspirations for a better life.

[0003] Currently, existing methods for identifying urban public activity centers lack self-verification of the identification results and fail to identify and classify urban public activity centers from a four-dimensional perspective of time and space. Summary of the Invention

[0004] To address the problems existing in the prior art, the purpose of this invention is to provide a method for identifying, classifying, and verifying urban public activity centers based on multi-source data.

[0005] To solve the above problems, the present invention adopts the following technical solution.

[0006] A method for identifying, classifying, and verifying urban public activity centers based on multi-source data includes: A: Facility spatial clustering analysis, which includes the following steps: The POI data of four types of public activity centers—parks and plazas, shopping plazas, sports centers, and science, education, and cultural institutions—were assigned to each standard spatial unit. The Getis-Ord Gi* index method of local spatial autocorrelation was used for analysis, and the values ​​were divided into low-value, medium-value and high-value areas according to significance, and the high-value area of ​​facility agglomeration was obtained respectively. By superimposing high-value areas with different facilities according to their different vitality levels, a high-value area for a comprehensive public activity center can be obtained. B: Based on the number and proportion of different facilities, areas with low, medium, and high levels of mixing are divided, including the following steps: The nature of the public center is identified based on four types of POI facilities. The proportion of each type of POI in each unit is calculated. When the proportion of a certain type is greater than 80%, the unit is determined to be a single-function area. When the proportion of all types of POI in the unit is less than 80%, the unit is determined to be a mixed-function area. Based on the identified mixed functional zones, the functional zones are rated using the degree of facility mixing and the popularity of local life service platforms, further narrowing down the identification scope; The functional mixing degree of a public activity center is calculated using the information entropy theory model. The higher the probability of a certain type of facility appearing, the lower the uncertainty of the information. The formula is as follows: H= In the formula, H represents the degree of mixing, and Pi represents the proportion of a certain type of facility to the total number of facilities. The smaller the sum of the proportions of each type of facility, the higher the degree of mixing of facilities in the public activity center.

[0007] Furthermore, based on the criteria of H < 0.4 being the low mixing region, 0.4 ≤ H < 0.7 being the medium mixing region, and H ≥ 0.7 being the high mixing region, the medium and high mixing regions are selected for the next step of verification and testing. C: Based on the superimposed identification of commercial facility building volume, the identification scope is further narrowed down, including the following steps: Overlay and label POIs onto the building outline data layer to filter out commercial facility buildings; The calculation of commercial facility building volume is based on the building area used for commercial functions. When the building has 1 to 5 floors, the formula for calculating the commercial facility building volume is as follows: S= When the number of floors in a building exceeds 5, the formula for calculating the commercial building volume is as follows: S= in, The first floor area is denoted as n, and the number of floors is denoted as n. The intersection tabulation method is used to count the number of commercial facilities within the grid, and public centers are further identified in five levels, with redder colors representing higher levels of public centers. D: Overlay the medium-to-high value areas identified through facility agglomeration, functional classification rating, and mixed-use characteristics with the high value areas of commercial facility building volume; For a given grid, if two high values ​​overlap, the center can be preliminarily identified as a municipal-level center; if only one value is high and the other is medium, the center can be preliminarily identified as a sub-municipal-level center; if neither value is high, or even if one value is low, the center can be preliminarily identified as a regional center; areas outside the medium-high value zone do not possess the characteristics of a public activity center. E: Based on the identification results of the distribution of public activity centers, a preliminary classification of urban public activity centers is performed using POI data for dining, leisure, and shopping, including the following steps: Using POIs converted from catering data from local life service platforms as static classification data, and through GIS visualization, by overlaying and comparing with the above identification results, urban public activity centers with catering functions as the main function are identified. Using POIs converted from leisure data from local life service platforms as static classification data, and through GIS visualization, by overlaying and comparing with the above identification results, urban public activity centers with leisure functions are identified. Using POIs converted from shopping data from local life service platforms as static classification data, and through GIS visualization, by overlaying and comparing with the above identification results, urban public activity centers with shopping functions are identified. F: The classification results of the distribution of public activity centers are further verified and checked using dynamic data, including the following steps: The OD data is sliced ​​into time segments based on the departure time of the trip, and the time consumed for leaving the trip in each time period is counted to calculate the time of the outflow function when leaving the grid in different time periods. Similarly, by slicing the travel data according to the arrival time, the time consumed by the arrival of each time period is calculated, and the time of inflow of the arrival grid is calculated. Calculate the time variation of grid functions: Time variation of functions in a time period = Time period inflow of functions - Time period outflow of functions; Calculate the cumulative time value of the grid function, which is the sum of the time changes of the function in all time periods before this time period.

[0008] Beneficial effects of the present invention Compared with the prior art, the advantages of this invention are: Compared with the prior art, the present invention has the following beneficial effects: Based on the spatiotemporal index measurement system composed of "preliminary identification-supervised classification-secondary verification", it can more accurately identify urban public activity centers and effectively classify them by combining dynamic and static open source data of time and space, thus promoting the government's refined management of urban construction in the era of stock. Attached Figure Description

[0009] Figure 1 OD data example diagram Figure 2 A grid map of high-value areas of facility agglomeration in six districts of a certain city.

[0010] Figure 3 Functional classification map of land parcels in six districts of a certain city.

[0011] Figure 4 A diagram showing the proportion of mixed-use facilities in the six districts of a certain city.

[0012] Figure 5 Commercial building volume diagram of six districts in a certain city.

[0013] Figure 6 Preliminary identification form for urban public activity centers.

[0014] Figure 7 Map showing the distribution of urban dining centers.

[0015] Figure 8 Map showing the distribution of urban leisure centers.

[0016] Figure 9 Map showing the distribution of city shopping centers.

[0017] Figure 10 Preliminary classification table of urban public activity centers.

[0018] Figure 11 Example chart of net asset value changes over time.

[0019] Figure 12 Schematic diagram of grid inflow and outflow time.

[0020] Figure 13 Example chart of cumulative function value data.

[0021] Figure 14 Functional value for every two hours on rest days in the six districts of a certain city.

[0022] Figure 15 Verification form for urban public activity centers.

[0023] Figure 16 This is a flowchart of the method of the present invention. Detailed Implementation

[0024] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0025] Example This embodiment takes a city with six districts (administrative district A, administrative district B, administrative district C, administrative district D, administrative district E, and administrative district F) as an example and provides a city public activity center based on multi-source open source data. The embodiment of the present invention is described below with reference to the accompanying drawings.

[0026] like Figure 16 As shown, S101, data from the target area is obtained through a multi-source data open platform.

[0027] The data includes public activity POI data, commercial facility AOI data, local life service platform data, and mobile signaling OD data. Among them, public activity POI data includes shopping-related commercial POIs, green park-related POIs, science, education and culture-related POIs, and sports and leisure-related POI data. The OD data is the data for the rest day of September 5, 2020, with 2-hour time slices. By writing SQL language code, four types of data are obtained: the time consumed for arriving at the grid, the time consumed for leaving the grid, the arrival time, and the departure time.

[0028] In the implementation case, the specific implementation method for obtaining data in step S101 is as follows: Public activity POI data and commercial facility AOI data were obtained from the GD map database; local life service platform data were obtained from local life service platform websites. Python web crawlers were used to crawl the XY coordinates of the corresponding facilities on the public websites of local life service platforms. In GIS, the coordinates of all data were uniformly converted to the WGS1984 geographic coordinate system to obtain POI data for three categories: catering, leisure, and shopping; mobile phone signaling OD data for a certain city on a certain rest day was provided by the Smart Footprint DaaS BI platform developed by LT Company on September 5, 2020.

[0029] All mobile signaling OD data preprocessing steps are performed on the data platform by writing SQL language code, and the target data is downloaded after data preprocessing.

[0030] The starting grid number of the trip is denoted as point O, the ending grid number of the trip is denoted as point D, the departure time of the trip is denoted as time_O, the arrival time of the trip is denoted as time_D, and the average time of the entire trip is denoted as time_OtoD. All data is expanded according to the weights provided by the data platform.

[0031] The OD data is categorized. Travel data is divided into departure and arrival segments. The raster numbers for departure and arrival are designated as the IDs of the travel data. The travel time (time_OtoD) is used as the value, and the data is categorized at two-hour intervals based on its respective time_O and time_D values. For example, the travel times for departures and arrivals between 0:00 and 2:00 are recorded as time_O_2 and time_D_2, respectively; the travel times for departures and arrivals between 10:00 and 12:00 are recorded as time_O_1012 and time_D_1012, respectively.

[0032] Match the OD data. Use the merge method in the pandas package to match the departure and arrival data for different time periods in (3) one by one according to the ID value. This will give you a data matching list for 12 time periods of the day. Taking 10-12 o'clock as an example, we get Figure 1 The list shows a one-to-one correspondence.

[0033] This completes the acquisition of POI data for public activities, AOI data for commercial facilities, data from local life service platforms, and OD data for mobile signaling.

[0034] S102. Divide the target area into grids and number them, and construct a "preliminary identification-supervision and classification-secondary verification" system for urban public activity centers based on the acquired data.

[0035] In one implementation case, the specific implementation method for step S102, which involves gridding the target area according to administrative boundaries, numbering the grids, and constructing a clustering index system, is as follows: The administrative boundaries of six districts in a city were extracted and imported into GIS software. A 300m*300m grid was created to cover the entire administrative boundary, and then clipped according to the boundary to obtain 6272 grids with corresponding numbers.

[0036] Based on the various factors that need to be considered in the identification, classification and verification of urban public activity centers, the identification indicators are divided into: public activity facility indicators, plot function indicators and mixed-use indicators; the classification indicators are divided into: catering function, leisure function and shopping function; and the cumulative value of function time is used as the verification indicator.

[0037] S103. Preprocess the data of the target area, link the divided grids respectively, and count the value of each influencing factor in the "identification" indicator system.

[0038] In one implementation example, the preprocessing of the acquired public activity POI data and commercial facility AOI data in step S103 is carried out as follows: For public activity point of interest (POI) data, all POI raster data were imported into GIS software. Spatial clustering analysis was performed using the Getis-Ord Gi* index method based on local spatial autocorrelation. The "reclassification" tool was selected, and the data was divided into three categories based on the natural discontinuity classification method. Corresponding intervals were assigned values ​​from low to high: "low activity zone," "medium activity zone," and "high activity zone," resulting in a raster map of high-value areas of facility agglomeration in six districts of a certain city. Figure 2 .

[0039] Hotspot analysis (Getis-Ord Gi*) calculates the z-score and p-value for each feature in a dataset using Getis-Ord Gi* statistics (called Gi-asters), revealing the spatial clustering locations of high- or low-value features. It identifies statistically significant hotspots by examining each feature in its neighborhood. To be considered a statistically significant hotspot, a feature should have a high value and be surrounded by other features with similarly high values. The local sum of a feature and its neighbors is compared to the sum of all features. A statistically significant z-score is generated when the local sum differs significantly from the expected local sum, making it impossible to determine if it is a random result.

[0040] The Getis-Ord local statistical formula is as follows:

[0041] Where x j It is the attribute value of element j. is the spatial weight between elements i and j, n is the total number of elements, and:

[0042]

[0043] The statistics are based on the z-score, therefore no further calculations are needed. The natural discontinuity classification method is based on the inherent natural grouping of data. It identifies classification intervals, enabling the most appropriate grouping of similar values ​​and maximizing the differences between classes. This grouping method divides the data into multiple classes, and for each class, its boundaries are set at locations where the differences in data values ​​are relatively large.

[0044] Calculate the sum of squared total biases (SDAM) for an array of classification results for a specific class. Let the set of results be denoted as array A, and its mean be: Then its total sum of squared deviations (SDAM) is: In equations (1) and (2), n is the number of elements in the array; Xi is the value of the i-th element.

[0045] For each combination of ranges in the classification results, calculate the sum of squared total biases (SDCM) and find the minimum value, denoted as SDCM min. Divide n elements into k classes, resulting in k subsets. One possible subset is [X1X2...Xi], [Xi+1Xi+2...Xj], ..., [Xj+1Xj+2...Xn]. Calculate the sum of squared total biases (SDAMi, SDAMj, ..., SDAMn) for each subset and sum them, SDCM1: SDCM 1=SDAM i +SDAM j +...+SDAM n (3) Similarly, the classification results can also be divided into other cases of k classes. The values ​​of SDCM2, ..., SDAMn are calculated in sequence, and the smallest value is selected as the final result SDCMmin. The result is then verified by the goodness of fit.

[0046] The gradient gvf i for each classification is calculated as follows: The gradient ranges from 1 (perfect fit) to 0 (poor fit). A higher gradient indicates greater inter-class differences. Experiments show that the classification obtained by SDCM min has the largest gradient value, leading to the conclusion that the natural discontinuity classification method yields relatively ideal results.

[0047] For public activity facility indicator data, namely shopping-related commercial POIs, green park-related POIs, science, education, and culture-related POIs, and sports and leisure POIs, ArcGIS software was used to spatially link them with a 300m*300m grid of six districts in a city, so that each raster carries a different proportion of POI data. The data was then exported to Excel software, and the proportion of the four types of POIs within each raster was calculated and classified. The classification method is as follows: When the proportion of a single type of POI exceeds 80%, the unit is designated as a single-function area (shopping, science and education, park, or sports). Conversely, when the proportion of any type of POI within the unit is less than 80%, the unit is designated as a mixed-function area. This results in the final land parcel function classification map. Figure 3 .

[0048] Furthermore, mixed-use zones within the functional areas were identified, and the mixing degree of the four types of POIs and the mixing popularity of local life service platforms were calculated to reflect the completeness and diversity of facilities in urban public activity centers, resulting in a facility mixing degree ratio map for six districts of a certain city. Figure 4 .

[0049] Facility Mixing Degree Calculation. The functional mixing degree of a public activity center is calculated using the information entropy theory model. The higher the probability of a certain type of facility appearing, the lower the uncertainty of information. Here, information uncertainty is interpreted as the facility mixing degree (H), and probability is interpreted as the proportion (Pi) of a certain type of facility within the public center. Therefore, the smaller the sum of the proportions of each type of facility, the higher the facility mixing degree of the public activity center. H < 0.4 is considered a low mixing degree region, 0.4 ≤ H < 0.7 a medium mixing degree region, and H ≥ 0.7 a high mixing degree region. The medium and high mixing degree regions are selected for further verification. The formula for calculating the H value is as follows: (i=1, ...,13) For processing commercial facility AOI data, ArcGIS software was used to spatially connect shopping-related commercial POI data, cleaning and filtering out the polygon features of commercial facility buildings. These polygon features were then spatially linked to a 300m*300m grid of six districts in a city. Intersection tabulation was used to statistically analyze the commercial building area data for each raster, and the commercial building volume was calculated from the commercial building area, thus obtaining a commercial building volume map. Figure 5 .

[0050] Intersection tabulation is used to calculate the intersection between two feature classes and to measure the area, length, or number of intersecting features. It is used to calculate the size (area and percentage of area) of the intersecting region for each class. A region consists of all features with the same region field value in the input region features: when both the input region features and the input class features are polygons, the output table will be based on the area calculation results; when the input class features are lines, the output table will be based on the linear calculation results; when the input class features are points, the output table will be based on the feature count.

[0051] The calculation of commercial building volume is based on the building area used for commercial functions. When the building has 1 to 5 floors, the formula for calculating the commercial building volume is as follows: S= When the number of floors in a building exceeds 5, the formula for calculating the commercial building volume is as follows: S= in, The first floor area is denoted as n, and the number of floors is denoted as n. The result obtained in step S103 Figure 2 , Figure 3 , Figure 4 , Figure 5 The initial identification of various public activity centers categorized them into different levels. High-value areas identified through facility clustering, functional classification and rating, and mixed-use characteristics were overlaid with high-value areas of commercial facility construction volume to further identify the public center system. For a given grid, if two high-value areas overlap, the center can be preliminarily identified as a municipal-level center; if only one is high-value and the other is medium-value, the center can be preliminarily identified as a municipal-level sub-center; if neither is high-value, or even if one is low-value, the center can be preliminarily identified as a regional center. Areas outside the high-value areas do not possess the characteristics of public activity centers. This completes the preliminary identification of urban public activity centers in six districts of a certain city, resulting in the preliminary identification scope of urban public activity centers and the preliminary identification table of urban public activity centers. Figure 6 .

[0052] S104. Preprocess the data of the target area, link the divided grids respectively, and count the value of each influencing factor in the "classification" index system.

[0053] In one implementation example, the preprocessing of the acquired local life service platform data in step S101 is specifically implemented as follows: For the local life service platform data, the POI location data for dining, leisure, and shopping obtained in step S102 are imported into ArcGIS software and spatially linked with a 300m*300m grid of six districts of a city, so that each grid carries a different proportion of POI data, which serve as indicators of dining, leisure, and shopping functions respectively.

[0054] The system classifies and grades the data based on natural breakpoints. By iteratively comparing the sum of squared differences between the mean and observed values ​​of each group and the elements in each group, the system determines the optimal arrangement of values ​​within the group, thus completing the grading of the functional indicators for catering, leisure, and shopping.

[0055] Furthermore, the natural breakpoints are divided into levels, and the optimal arrangement of values ​​within each group is determined by iteratively comparing the sum of squared differences between the mean and the observed values ​​of each group and the mean of the elements in that group. Specifically, this includes: Convert the quantitative data of the three types of functional indicators into arrays respectively. ,and ; Calculate array Sum of squares of deviations from the mean As shown in the following formula:

[0056] in, , For array The mean, For array Length; Iterate through each range combination, calculating the sum of squared deviations of the class means. And find the minimum value, assuming the current range is an array. array and array As shown in the following formula:

[0057] in, , , , , and arrays array and array The mean, , and arrays array and array Length; The smallest Marked as This makes the range combination at this time labeled as an array. array and array This is to determine the best arrangement of values ​​in the grouping.

[0058] Based on the natural discontinuity grading method, the functional indicators of catering, leisure, and shopping are divided into three categories, resulting in density-grading raster maps of the number of facilities in these three categories: urban catering centers, urban leisure centers, and urban shopping centers. Figure 7 , Figure 8 , Figure 9 .

[0059] This completes the preliminary classification of urban public activity centers in the six districts of a certain city, yielding the preliminary classification scope and a preliminary identification table for urban public activity centers. Figure 10 .

[0060] S105. Preprocess the data of the target area, link the divided grids respectively, and calculate the value of each influencing factor in the "verification" indicator system.

[0061] In one implementation example, the preprocessing of the acquired mobile signaling OD data in step S102 is specifically implemented as follows: For OD data, calculate the net functional time change of a specific grid using the data obtained in step S101. Record the net functional time change of a specific grid over a given time period as `time_result`. For example, from 10:00 to 12:00: time_result_1012 = time_D_1012 - time_O_1012 Calculate the time_result for each of the 12 time periods. Use the `merge` method from the pandas package to match the time_results of the 12 time periods based on their ID values. This results in a list where the ID values ​​of a given grid correspond one-to-one with the time_results of the twelve time periods. Figure 11 : Define the functional value of a land parcel. The difference between the "inflow functional time" and the "outflow functional time" of the target land parcel is defined as the "functional time change," while the sum of the functional time changes of all time periods before the current time period is used as the "cumulative functional time value," i.e., the functional value. For example, if user A leaves grid A at 6:50, and the trip takes 120 minutes, arriving at grid B at 9:00, then the 120 minutes consumed by user A's trip are considered to be the time "taken" by user A from A during the 6-8 o'clock time period. All the time "taken" by users leaving A during this time period is the outflow functional time of A. This part of the time also belongs to the inflow functional time of grid B during the 8-10 o'clock time period. Figure 12 As shown, the functional value reflects the probability that the plot of land will be chosen by someone, thereby measuring the effectiveness of the plot's function and verifying the identified public activity center.

[0062] Calculate the functional value of the plot. According to the different ID values, accumulate the net value of the time change of socially necessary functions in (5) time_result, and record the accumulated value of each time period as process_a certain time period. Use the df.min(axis=1) method of pandas package to retrieve the minimum value of process for each ID in the twelve time periods, and record the minimum value of result for that ID as df['min_val']. Record the net value of function for each ID in each time period as result_a certain time period. Let the net value of function result for each time period be equal to the process of that time period minus the minimum value of process of that ID df['min_val'], calculate the result of each ID in the 12 time periods respectively, and obtain a one-to-one corresponding list, that is Figure 13 .

[0063] ArcGIS software was used to classify and grade the accumulated functional time values ​​for different time periods on rest days using natural breakpoints. The optimal arrangement of values ​​within each group was determined by iteratively comparing the sum of squared differences between the mean and observed values ​​of each element in each group. Further, [the following steps were performed]. Figure 14 Visual representation.

[0064] It is believed that during the peak travel period on weekends, from 10:00-12:00 to the return peak period, from 20:00-22:00, the site exhibits a high probability of becoming an urban public activity center, exhibiting a functional "beginning of construction - completion of construction - gradual deconstruction" phenomenon. It is also believed that during the same period, the site exhibits a high probability of becoming a residential area.

[0065] Will Figure 14The city public activity center identification and classification diagrams obtained from the "identification-classification" steps in steps s103 and s104 are classified, and the results are verified.

[0066] This completes the verification of the urban public activity centers in the six districts of a certain city, resulting in the verified scope of the urban public activity centers, as follows: Figure 15 As shown.

[0067] It should be noted that although the method operations of the above embodiments are described in a specific order, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. On the contrary, the described steps may be performed in a different order. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, or one step may be broken down into multiple steps.

[0068] The above description is merely a preferred embodiment of the present invention; however, the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and its improved concepts, should be covered within the scope of protection of the present invention.

Claims

1. A method for identifying, classifying, and verifying urban public activity centers based on multi-source data, characterized in that, include: A: Facility spatial clustering analysis includes the following steps: The POI data of four types of public activity centers—parks and plazas, shopping plazas, sports centers, and science, education, and cultural institutions—were assigned to each standard spatial unit. The Getis-Ord Gi* index method of local spatial autocorrelation was used for analysis, and the area was divided into low-value area, medium-value area and high-value area according to significance, and the high-value area of ​​facility agglomeration was obtained. By superimposing high-value areas with different facilities according to their different vitality levels, a high-value area for a comprehensive public activity center can be obtained. B: Based on the number and proportion of different facilities, areas with low, medium, and high levels of mixing are divided, including the following steps: The nature of the public center is identified based on four types of POI facilities. The proportion of each type of POI in each unit is calculated. When the proportion of a certain type is greater than 80%, the unit is determined to be a single-function area. When the proportion of all types of POI in a unit is less than 80%, the unit is determined to be a mixed-function area. Based on the identified mixed functional zones, the functional zones are rated using the degree of facility mixing and the popularity of local life service platforms, further narrowing down the identification scope; The functional mixing degree of a public activity center is calculated using the information entropy theory model, as follows: H= In the formula, H represents the degree of mixing, and Pi represents the proportion of a certain type of facility to the total number of facilities; Furthermore, based on the criteria of H < 0.4 being the low mixing region, 0.4 ≤ H < 0.7 being the medium mixing region, and H ≥ 0.7 being the high mixing region, the medium and high mixing regions are selected for the next step of verification and testing. C: Based on the superimposed identification of commercial facility building volume, the identification scope is further narrowed down, including the following steps: Overlay and label POIs onto the building outline data layer to filter out commercial facility buildings; The calculation of commercial facility building volume is based on the building area used for commercial functions. When the building has 1 to 5 floors, the formula for calculating the commercial facility building volume is as follows: S= When the number of floors in a building exceeds 5, the formula for calculating the commercial building volume is as follows: S= in, The first floor area is denoted as n, and the number of floors is denoted as n. The intersection tabulation method is used to count the number of commercial facilities within the grid, and public centers are further identified in five levels, with redder colors representing higher levels of public centers. D: Overlay the medium-to-high value areas identified through facility agglomeration, functional classification rating, and mixed-use characteristics with the high value areas of commercial facility building volume; For a given grid, if two high values ​​overlap, the center can be preliminarily identified as a municipal-level center; if only one value is high and the other is medium, the center can be preliminarily identified as a sub-municipal-level center; if neither value is high, or even if one value is low, the center can be preliminarily identified as a regional center; areas outside the medium-high value zone do not possess the characteristics of a public activity center. E: Based on the identification results of the distribution of public activity centers, a preliminary classification of urban public activity centers is performed using POI data for dining, leisure, and shopping, including the following steps: Using POIs converted from catering data from local life service platforms as static classification data, and through GIS visualization, by overlaying and comparing with the above identification results, urban public activity centers with catering functions as the main function are identified. Using POIs converted from leisure data from local life service platforms as static classification data, and through GIS visualization, by overlaying and comparing with the above identification results, urban public activity centers with leisure functions are identified. Using POIs converted from shopping data from local life service platforms as static classification data, and through GIS visualization, by overlaying and comparing with the above identification results, urban public activity centers with shopping functions are identified. F: The classification results of the distribution of public activity centers are further verified and checked using dynamic data, including the following steps: The OD data is sliced ​​into time segments based on the departure time of the trip, and the time consumed for leaving the trip in each time period is counted to calculate the time of the outflow function when leaving the grid in different time periods. Similarly, by slicing the travel data according to the arrival time, the time consumed by the arrival of each time period is calculated, and the time of inflow of the arrival grid is calculated. Calculate the time change of grid functions: Time change of function in a time period = Time inflow of function in a time period - Time outflow of function in a time period; Calculate the cumulative value of grid function time, which is the sum of the time change values ​​of all time periods before this time period.