Community facility use space-time state identification method based on resident activity and attribute clustering
By constructing a three-dimensional coupled analysis framework and a stepwise clustering method, and using questionnaire survey data to identify the spatiotemporal status of community facility use, this study solves the problem of lack of systematic integration in existing research and enables refined planning support for the use of community facilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-02
- Publication Date
- 2026-04-07
AI Technical Summary
Existing research lacks a composite analytical framework that systematically integrates 'individual attributes, activity patterns, and facility use', resulting in reduced interpretability of results and difficulty in supporting the needs of refined planning at the micro and meso scales. In particular, it is difficult to reveal the intrinsic relationship between behavioral patterns and population attributes in areas with limited data.
A three-dimensional coupled analysis framework of 'individual attributes - activity patterns - facility usage' was constructed. Stepwise clustering was performed using K-means and DBSCAN algorithms, combined with random forest and Cohen's f effect size analysis, and dimensionality reduction was performed layer by layer using questionnaire survey data to identify typical user characteristics and facility usage patterns.
It enables accurate identification of the spatiotemporal status of community facilities, provides refined and human-centered support for community living circle planning, and improves the interpretability of the model and the applicability of the data, making it particularly suitable for rural and suburban areas at the micro- and meso-scale.
Smart Images

Figure CN121808440A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of urban planning, geography, sociology and management, in particular to a method for identifying the spatio-temporal state of 15-minute life circle community facility use based on resident activity patterns and individual attribute clustering. BACKGROUND
[0002] The use of community facilities is closely related to residents' daily life and the high-quality operation of the city. Currently, sociology mainly analyzes the accessibility differences of facility services from the perspectives of social networks and fairness, while transportation engineering focuses on the interaction between commuting behavior and facility layout. Geographical and urban planning fields mainly explore facility use characteristics from the perspectives of spatial form and behavior patterns through empirical methods. In this context, this research is based on the professional perspectives of urban planning and behavioral geography, focusing on the internal relationship between "individual attributes-activity patterns-facility use", and providing fine and humanized decision support for community life circle planning.
[0003] In the field of facility use research, geographers and urban planners have adopted empirical methods. In the early stage, they relied on micro-scale questionnaire surveys and field interviews to obtain data. With the development of big data technology, the research scale has expanded to the city level, and the data sources have expanded from traditional questionnaires (Wan Chengwei et al., 2020) to multi-source information such as location-based service data (LBS), point of interest data (POI), and remote sensing images (Matuszewska et al., 2023; Yang Zhen-shan et al., 2021). The use of mobile phone signaling identification and GPS tracking to analyze public space use patterns has emerged (Sun Daosheng et al., 2017; Wang Wei et al., 2024). In terms of technical approach, empirical methods focusing on feature induction such as hierarchical clustering, K-means clustering, self-organizing mapping (SOM), time series analysis, spatial statistics, and behavior modeling are relied upon (Hussaini et al., 2022; Yhee et al., 2023; Zhang Jingxiang et al., 2012). These improvements in data and technology have directly promoted the spatio-temporal behavior analysis to a deeper level of quantitative exploration and rule revelation. At the same time, in terms of topic selection, some studies have gradually shifted from isolated behavior analysis to systematic exploration of multiple types of facilities (Geng Jian et al., 2013) and complex activity patterns (Li Ying et al., 2019), providing important academic support for understanding the spatio-temporal rules of facility use.
[0004] However, there are three obvious limitations in existing research: first, most studies focus on the binary relationship between "people-activities" or "activities-facilities", lacking a comprehensive analysis framework that integrates "individual attributes-activity patterns-facility use"; second, the method often leads to a decrease in result interpretability due to the expansion of feature dimensions; third, there is a blind spot in revealing the internal relationship between behavior patterns and population attributes, especially in supporting the detailed planning needs of areas with weak data at the meso-micro scale.
[0005] The advantages and innovations of the method are as follows: (1) At the topic level, a three-dimensional coupled analysis framework of "individual attributes-activity patterns-facility use" is established, systematically revealing the internal laws of user attributes, activity chain characteristics, and facility use; (2) At the technical method level, through step-by-step clustering (K-means and DBSCAN algorithms) and association filtering (random forest and Cohen's f effect size analysis), dimensionality is reduced layer by layer, effectively controlling the exponential growth of feature quantity, and improving model interpretability; (3) At the data level, a one-questionnaire covering multiple dimensions of individual attributes is established as the core technical path, effectively making up for the shortcomings of relying solely on trajectory big data in identifying attribute-behavior correlations, especially suitable for areas with weak data but significant planning needs at the meso-micro spatial scale, such as rural and suburban areas.
[0006] The key problem addressed by the invention is the coupling relationship between user population attributes and facility use. Further dividing it into the following sub-problems: What are the typical characteristics of the daily activity patterns of users with different individual attributes? What is the internal relationship between individual attributes and activity patterns? Based on the "group-activity" sub-type, what is the spatio-temporal state of community facility use? To systematically address the above problems, it is necessary to clarify the core concepts of "community facilities and life circle", "daily activities and activity chain", and "attribute clustering and individual attributes".
[0007] First, community facilities and life circle. According to the "Urban Residential District Planning and Design Specification" (GB50180-2018) and "Community Life Circle Planning Technical Guide" (TD / T 1062-2021), "community facilities" refer to facilities that serve residents' basic needs such as "clothing, food, housing, transportation, culture, education, sports, and health". The configuration emphasizes two levels of life circle: 15 minutes and 5-10 minutes. In this invention, the concept specifically refers to the collection of physical places within the above-mentioned life circle range, whose use state can be indirectly reflected by users' spatio-temporal activity behavior. The core is to analyze users' activity patterns to identify the real use state of facilities in the time (frequency, time period) and space (distance, accessibility) dimensions.
[0008] Secondly, daily activities and activity chains. In time geography and behavioral geography, "daily activities" are often defined as the sum of a series of purposeful behaviors engaged in by residents in a specific spatio-temporal environment, with three characteristics of time limitation, place uniqueness and time interval of activity transfer, and these activities are connected in time sequence in the individual's daily itinerary, which is called "activity chain". Based on the systematic analysis of activity classification of urban and rural residents, combined with the unique agricultural and agricultural production properties of rural communities, the invention classifies it into seven categories of production, sleep, private affairs, housework, shopping, leisure and entertainment, and travel, and the production activities further cover agriculture, forestry, animal husbandry, fishing, work, learning and other contents. The activity chain is described by three attributes of activity type, space (facility type) and time, and the typical spatio-temporal mode of the user can be identified by clustering analysis, and the collective rules of facility use are revealed.
[0009] Thirdly, attribute clustering and individual attributes. As a key step of population segmentation, the basic variables of "attribute clustering" cover multi-dimensional attributes such as population statistics, social economy and family environment, such as age, income, education level, family structure, etc. Specifically, the "individual attributes" focused on by the invention mainly include three categories: (1) personal basic attributes, such as age, gender, education level, etc.; (2) family economic and living conditions, such as total family income, number of employed persons, number of people living together, housing area, etc.; (3) core life needs, such as rigid service needs due to family structure (such as whether there are school-age children) or life stage (such as old age). These attributes are used as basic variables for clustering analysis to construct stable and interpretable population portraits, and then coupled with spatio-temporal activity patterns for coupled analysis. SUMMARY
[0010] To solve the problems mentioned in the background art, the invention provides a method for revealing the internal correlation of "people-facility use", which provides an operable path for spatio-temporal state recognition of 15-minute community life circle facility use based on resident activity patterns and individual attribute clustering.
[0011] The purpose of the invention can be achieved by the following technical solution: a method for recognizing the spatio-temporal state of community facility use based on resident activity and attribute clustering, the method comprising the following steps:
[0012] First-hand data is collected through a questionnaire survey, and the determination method of the questionnaire distribution range is as follows: taking the geographic coordinates of any community facility in a community in a city as the center, combining the actual city road network, and using the network analysis of the Arcgis platform to generate a 15-minute (1000 meters) polygon service area, and calculating the union of all community facility service areas in the community to serve as the questionnaire distribution range.
[0013] The data collected by the questionnaire includes: individual attribute data of the user and activity log of a typical day, the activity type and the corresponding activity facility of the user are counted every two hours as a time unit. In addition, the housing location of the respondent is recorded and the corresponding geographic coordinates are queried. The questionnaire data is screened for effectiveness, and the individual attribute and activity log data are filled in completely. It is suggested that the number of questionnaires should not be less than 2‰~3‰ of the number of residents in the questionnaire distribution range, so as to ensure the sample size of the reorganization (the rural and town areas can be appropriately reduced).
[0014] The individual attributes are clustered using the obtained effective questionnaire data to preliminarily screen the group characteristics;
[0015] The activity mode of the user is clustered and the main features are extracted using the activity log data in the obtained effective questionnaire;
[0016] The individual attribute clustering and activity mode clustering results are coupled to identify specific type of "group-activity" association characteristics;
[0017] The "group-activity" association characteristics are used to identify the space-time state of the user's use of community facilities, with frequency+time (time sequence change of frequency) and frequency+space (circle layer change of frequency) as core indicators.
[0018] Preferably, the basis for screening the effective questionnaire data is as follows:
[0019] Since 0:00-4:00 is basically in the sleep period of residents, and it is difficult to reflect the differences in individual activities, the questionnaire with no blank information in the activity log of a typical day between 4:00-24:00 and complete individual attribute information is effective data. The individual attribute information includes but is not limited to: gender, age, nationality, education level, occupation type, unit nature, family structure, income level, consumption level, employment status, housing condition, housing location, number of commonly used transportation tools and their use frequency, etc.
[0020] Preferably, the calculation process of the individual attribute clustering and group characteristic preliminary screening is as follows:
[0021] Firstly, the clustering result of individual attribute is calculated by K-means algorithm; then, the group characteristics are extracted by random forest algorithm.
[0022] Preferably, the K-means algorithm handles the categorical variables by converting them into dummy variables (0-1 encoding) to make them into virtual variables that can be quantitatively calculated, and then brings them into calculation;
[0023] Preferably, the determination method of the K-means algorithm classification cluster number (k) is as follows:
[0024] First, all individual attribute variables (quantitative + categorical variables after dummy variable transformation) are included in the cluster analysis. Then, the elbow rule, cluster profile coefficient, and group sample size balance are comprehensively applied. The elbow rule determines the number of candidate clusters based on the inflection point (from steep to gentle) of the sum of squared errors as a function of k. The cluster profile coefficient is used to evaluate the separation quality of each cluster (ideally >0.7). Group sample size balance is assessed by calculating the sum of the absolute values of the differences in sample size among clusters (ideally smaller sums). The final k value is the number of clusters corresponding to the optimal combination of clusters located near the inflection point (within 2 units before and after the inflection point), with a higher cluster profile coefficient (>0.7) and lower group sample size balance. Simultaneously, a significance test (p < 0.05 is considered significant, otherwise p ≥ 0.05 is not significant) is used to identify individual attribute variables that show significant differences between the clustered categories. Only variables with significant differences are retained for subsequent analysis. Finally, the number of clusters k and the individual attribute variables with significant differences are determined.
[0025] Preferably, the method for initial screening of group characteristics is as follows: using the clustering analysis type classification result as the dependent variable and the aforementioned significantly different individual attributes as independent variables, a random forest algorithm is used for classification calculation. The node splitting evaluation criterion is MSE (mean squared error), the number of decision trees is 100, sampling with replacement is used, the minimum number of samples for internal node splitting is 2, the minimum number of samples for leaf nodes is 1, the maximum number of leaf nodes is 50, and other parameters are default. The weight contribution of feature indicators in the random forest algorithm results (…) Sort the indicators in descending order and calculate the sum of their cumulative weighted contributions. ),by (85%≤) When the percentage of each of the above n indicators is ≤95%, the n indicators are used as the representation dimensions of the group characteristics. Then, the proportion of the secondary indicators of each of the above n indicators is calculated. ), and sort them from highest to lowest according to their proportion, and calculate the sum of their cumulative proportions ( ),by (80%≤) When the percentage of individuals is ≤90%, the meaning represented by the j secondary indicators is the group characteristic described by that indicator. Based on this, the results of individual attribute clustering are named according to the following rule: n indicators are arranged in descending order of their weight contribution, and the meaning under any indicator is determined by the j secondary indicators. For example, using the random forest algorithm, age, income, and education level are selected as the feature indicators describing the population attributes (i.e., n=3), and ordered by their weight as "age-income-education level". For age, if the cumulative proportion of those aged 45-70 is higher than 80%, then the specific meaning under the age indicator is middle-aged and elderly people aged 45-70. Similarly, the income indicator represents middle-to-high income (5000-10000 yuan), and the education level indicator represents low education level (mostly primary and secondary school students). Therefore, the final name is "Middle-aged and elderly (45-70 years old), middle-to-high income (5000-10000 yuan), low education level (mostly primary and secondary school students)". This yields the main characteristics of clustering based on individual attributes.
[0026] Preferably, the calculation process for activity pattern clustering and main feature extraction is as follows:
[0027] First, the activity log data is encoded to form a continuous activity chain containing "time_space_activity". Then, the DBSCAN algorithm is used to cluster the user's spatiotemporal activity chain to obtain activity patterns; finally, the random forest algorithm is used to identify the main features of any activity pattern.
[0028] Preferably, the coding rule for the activity log is as follows: using 2-hour time units, starting from 4:00, the 10 time periods are sequentially named... Naming of type i daily activities Naming of type m activity spaces (facility types) And so on, the activity chain is encoded as " ".
[0029] Preferably, the clustering algorithm for the activity pattern is to cluster the "time_activity" (or "time_activity") data in any activity log data. "Time_Space" "Time_Space_Activity" The user's activity pattern is obtained using the DBSCAN algorithm with multiple combinations of features as independent variables. Specifically, the minimum number of samples within the neighborhood of the DBSCAN algorithm should be between "feature count + 1" and "twice the feature count". For data with k dimensions, the distance from any point in the data to its k-nearest neighbor is calculated. All points are sorted by k-distance in descending order and plotted as a line graph. The k-distance value corresponding to the inflection point from flat to steep is set as the initial value of the neighborhood radius. For variables with large differences in scale, Z-score standardization is performed first, and then the values are substituted into the calculation. Subsequently, through multiple iterations and by combining the proportion of cluster noise points (the smaller the better) and the cluster silhouette coefficient (>0.7 and the higher the better), the optimal clustering result is determined.
[0030] The formula for Z-score standardization is:
[0031]
[0032] in, The standardized value. The original data values, The average value of the data. denoted as the standard deviation of the data.
[0033] Preferably, the method for identifying the main features of any type of activity pattern is as follows: using the clustering result as the dependent variable, and the "time_activity" participating in the clustering ( "Time_Space" "Time_Space_Activity" The independent variable is , and the random forest algorithm is used for classification calculation. The node splitting evaluation criterion is MSE (mean squared error). The number of decision trees is 100, sampling with replacement is used, the minimum number of samples for internal node splitting is 2, the minimum number of samples for leaf nodes is 1, the maximum number of leaf nodes is 50, and other parameters are default. The weight contribution of the feature index (independent variable) is also calculated. ). Identification Given w independent variables (i.e., the average weight contribution of the independent variables), calculate the sum of the weight contributions of these w independent variables. ),by (80%≤) When the percentage of activities is ≤90%, the intersection of "space_activity" among the v independent variables is taken as the main feature of the activity pattern, and the "time" distribution feature is added to the naming. That is, the naming rule is to list the corresponding "time_space_activity" in order from morning to night. If multiple intersections exist, they are also listed in order from morning to night. If multiple time periods have the same activities and spaces, the time periods can be combined for description. For example, "time period..." space Activity - Time period Space Activity ” is named as “Morning ( ) Home ( ) Production ( ) - Afternoon ( ) Shopping in commercial facilities ( )”; “Time period Space Activity Activity - Time period Space Activity ” is named as “All day ( ) Home ( ) Production ( )”. Thus, the main features of clustering based on human activity patterns are obtained.
[0034] Preferably, the method for identifying the association features of the specific type “group - activity” is as follows:
[0035] First, use one - way ANOVA. For any population clustering type, take all secondary indicators under the n indicators reflecting group characteristics and the clustering results of activity patterns as variables for calculation; then, through one - sample variance test and Welch's variance test, identify the associated secondary indicators that have a significant differential effect on activity patterns; then, calculate Cohen's f value to measure the degree of association; finally, finely describe the main features of the specific type of “group - activity”.
[0036] Preferably, the selection of the variance test method for one - way ANOVA is judged according to the homogeneity of variance. If the homogeneity of variance is satisfied, use one - sample variance test; if the homogeneity of variance is not satisfied, use Welch's variance test. Both are based on as the basis for judging whether there is an association. If then there is an association, otherwise there is no association.
[0037] Preferably, the method for measuring the degree of association by calculating Cohen's f value is judged by the following rules: 0 < Cohen's f < 0.1, the degree of association is extremely small; 0.1 ≤ Cohen's f < 0.25, the degree of association is small; 0.25 ≤ Cohen's f < 0.4, the degree of association is medium; Cohen's f ≥ 0.4, the degree of association is large.
[0038] Preferably, the method for finely characterizing the main features of a specific type of "group-activity" is as follows: For the population clustering results under any activity pattern, extract the primary individual attribute indicators corresponding to the secondary indicators with moderate correlation (Cohen's f≥0.25) or higher in the aforementioned steps; for any primary indicator, calculate the proportion of all its secondary indicators (…). And calculate the sum of their proportions in descending order. );by ≥ (65%≤) The secondary indicators contained in (≤90%) are used as the characteristics of the indicator; by traversing all primary indicators, the main characteristics of the "group + activity" coupling under the human element are obtained.
[0039] Preferably, the process for identifying the spatiotemporal status of the user community facilities is as follows:
[0040] First, we analyze the spatiotemporal state of any facility's use from two aspects: temporal change characteristics and spatial concentric change characteristics. Then, we use the same method to traverse other facilities and summarize the individual attributes of the users of any community facility and their corresponding spatiotemporal usage status.
[0041] Preferably, the method for analyzing the spatiotemporal usage status of a facility is as follows:
[0042] For any "group-activity" sub-category, user data from those who used the facility at least once within a single day's activity chain are extracted and analyzed. Regarding the characterization of temporal variation features, based on the main features of the "group-activity" sub-pattern described above, the time periods involved in the activity chains from the main features are extracted. The temporal distribution of the proportion of user facility usage frequency during these time periods (the proportion of users using the facility in each time period out of the total daily users) is statistically analyzed to identify peak periods. Regarding the characterization of spatial concentric variation features, with the facility as the center, a radius of 100 meters (a typical 1-minute walk distance for community residents), and a boundary of 1000 meters within a 15-minute community living circle, ArcGIS network analysis is used to form a multi-ring service area integrated with the urban road network (i.e., the radius interval between each service area ring is 100 meters). The proportion distribution of users accessing the facility within each ring (the proportion of users using the facility within a certain ring out of all users using the facility) is calculated to identify concentrated distance ranges. Following the above method, the spatiotemporal usage status of any facility and the main individual characteristics of its user groups are traversed and summarized. The final result is a summary table containing groups with different individual attribute characteristics, the concentrated time periods and distances of activities performed using any facility, for example:
[0043]
[0044] The beneficial effects of this invention are as follows: This invention provides a method for relatively accurately identifying the spatiotemporal status of community facility use based on questionnaire survey data and resident activity and attribute clustering. First, the obtained valid questionnaire data is used to cluster individual attributes to initially screen group characteristics. Second, activity log data from the obtained valid questionnaires is used to cluster user activity patterns and extract key features. Subsequently, the results of individual attribute clustering and activity pattern clustering are coupled to identify specific types of "group-activity" association features. Finally, using the obtained "group-activity" association features, frequency + time (temporal changes in frequency) and frequency + space (layered changes in frequency) are used as core indicators to identify the spatiotemporal status of user use of community facilities. This invention can effectively identify the spatiotemporal status of facility use within a 15-minute community living circle based on resident activity patterns and individual attribute clustering, providing refined and human-centered decision support for community living circle planning. Attached Figure Description
[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0046] Image title list:
[0047] Figure 1 This is a flowchart illustrating a method for identifying the spatiotemporal status of community facility use based on resident activity and attribute clustering.
[0048] Figure 2 It is a scatter plot and type naming of the mathematical structure of K-means user attribute clustering.
[0049] Figure 3 This refers to the scatter plot of the mathematical structure of DBSCAN user activity pattern clustering and its type naming. Detailed Implementation
[0050] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0051] like Figure 1 As shown, a method for identifying the spatiotemporal status of community facility use based on resident activity and attribute clustering is described. The method includes the following steps:
[0052] Step 1: Obtain residents' individual attributes and daily activity data through questionnaire distribution. It is recommended that the number of questionnaires be no less than 2‰ to 3‰ of the number of residents within the questionnaire distribution scope to ensure the sample size for reorganization (this can be appropriately reduced in rural areas). In this example, targeting villages and towns in location X, questionnaires are distributed at a rate of 1‰ of the total number of residents. The individual attribute data of residents in the selected valid questionnaires are clustered to initially screen the group characteristics.
[0053]
[0054] Step 1.1: Use the K-means algorithm to cluster the selected individual attribute variables (including: type of production activity, education level, number of people living together, shared arable land, facilities, housing type, number of rooms, age, total housing area, number of family members employed, family living expenses, family monthly income, gender, marital status, and number of family members engaged in agricultural activities). The elbow principle, cluster profile coefficient, and group sample size balance are comprehensively applied. The elbow principle determines the number of candidate clusters based on the inflection point (from steep to gentle) of the sum of squared errors as a function of k. The cluster profile coefficient... The cluster profile coefficient (CPC) is used to assess the separation quality of each cluster (ideally >0.7). Group size balance is assessed by summing the absolute values of the differences in sample size between clusters (ideally smaller sums). The final value of k is the number of clusters corresponding to the optimal combination located near the inflection point (within 2 units before and after the inflection point), with a higher CPC (>0.7) and lower group size balance. Simultaneously, a significance test (p < 0.05 is considered significant, p ≥ 0.05 is considered insignificant) is used to identify individual attribute variables that show significant differences between the clustered categories. Based on these principles, this example determines the number of clusters to be 4 (k=4). Eight categories of individual attribute variables (including: education level, number of people living together, number of rooms, age, total housing area, number of family members employed, family living expenses, and family monthly income) that passed the significance test (p < 0.05) are identified and retained for subsequent analysis.
[0055] Step 1.2: Using the clustering analysis results as the dependent variable and the significantly different individual attributes as independent variables, the random forest algorithm is used for classification calculation. The node splitting evaluation criterion is MSE (mean squared error), the number of decision trees is 100, sampling with replacement is used, the minimum number of samples for internal node splitting is 2, the minimum number of samples for leaf nodes is 1, the maximum number of leaf nodes is 50, and other parameters are default. The weight contribution of the feature indicators in the random forest algorithm results (…) Sort the indicators in descending order and calculate the sum of their cumulative weighted contributions. ),by Five (n=5) indicators (including: total housing area, number of rooms, number of people living together, monthly household income, and age) are used as the representation dimensions of group characteristics when the percentage is 90%.
[0056] Step 1.3: Calculate the weight of the secondary indicator for each of the above five indicators. ), and sort them from highest to lowest according to their proportion, and calculate the sum of their cumulative proportions ( ),by When (=80%), the meaning represented by j secondary indicators is the group characteristic described by that indicator. Based on this, the results of individual attribute clustering are named as follows:
[0057] Group Category 1: Medium to high total housing area (200-300m²) 2 Number of rooms: medium to high (4-8 rooms); Education level: low (junior high school); Monthly family income: medium to high (3000-10000+); Age: middle-aged to elderly (45-70 years old); Number of family members engaged in agricultural activities: few (0-2 people).
[0058] Group Category 2: Low to medium total housing area (80-200m²) 2 Number of rooms: low to medium (2-6 rooms); Education level: medium (college diploma); Monthly family income: low to medium (500-10000+); Age: middle-aged to elderly (40-67 years old); Number of family members engaged in agricultural activities: 0-2 people.
[0059] Group Category 3: Low total housing area (60-150m²) 2 Low to medium number of rooms (2-5 rooms), low level of education (primary school), low monthly family income (0-8000 yuan), middle-aged and elderly (50-75 years old), few family members engaged in agricultural activities (0-2 people).
[0060] The final scatter plot of the clustering mathematical structure and its feature recognition results are attached. Figure 2 As shown.
[0061] Step 2: Based on the activity log data from the valid questionnaires, cluster the users' activity patterns and extract the main features.
[0062] Step 2.1: Using 2-hour time units, starting from 4:00, name the 10 time periods sequentially. Naming of type i daily activities Naming of type m activity spaces (facility types) And so on, the activity chain is encoded as " ".
[0063] Step 2.2: Extract the "Time_Activity" data from any activity log entry. "Time_Space" "Time_Space_Activity" The user's activity pattern is obtained using the DBSCAN algorithm with multiple combinations of features as independent variables. Specifically, the minimum number of samples within the neighborhood of the DBSCAN algorithm should be between "feature count + 1" and "twice the feature count". For data with k dimensions, the distance from any point in the data to its k-nearest neighbor is calculated. All points are sorted by k-distance in descending order and plotted as a line graph. The k-distance value corresponding to the inflection point from flat to steep is set as the initial value of the neighborhood radius. For variables with large differences in scale, Z-score standardization is performed first, and then the values are substituted into the calculation. Subsequently, through multiple iterations and by combining the proportion of cluster noise points (the smaller the better) and the cluster silhouette coefficient (>0.7 and the higher the better), the optimal clustering result is determined.
[0064] The formula for Z-score standardization is:
[0065]
[0066] in, The standardized value. The original data values, The average value of the data. denoted as the standard deviation of the data.
[0067] In this example, the optimal clustering result for user activity patterns is determined to be 5 classes.
[0068] Step 2.3: Using the above activity clustering results as the dependent variable, and the "time_activity" participating in the clustering ( "Time_Space" "Time_Space_Activity" The independent variable is , and the random forest algorithm is used for classification calculation. The node splitting evaluation criterion is MSE (mean squared error). The number of decision trees is 100, sampling with replacement is used, the minimum number of samples for internal node splitting is 2, the minimum number of samples for leaf nodes is 1, the maximum number of leaf nodes is 50, and other parameters are default. The weight contribution of the feature index (independent variable) is also calculated. ).
[0069] Step 2.4: Identification Given w independent variables (i.e., the average weight contribution of the independent variables), calculate the sum of the weight contributions of these w independent variables. ),by When the percentage is 80%, the intersection of "space_activity" among the v independent variables is the main feature of the activity pattern, and the "time" distribution feature is added to the name.
[0070] If multiple overlaps exist, the corresponding "Time_Space_Activity" are listed sequentially from earliest to latest, serving as the main features and names of the corresponding activity patterns. The main features of this human activity pattern clustering are as follows:
[0071] Activity Category 1: Morning (8-12 am) at public facilities ( )Production( - In the afternoon and evening (2-6 pm) at public facilities ( )Production( )
[0072] Activity Category 2: Morning (10–12 AM) at Home )housework( —Evening (4–6 PM) at home )housework( )
[0073] Activity Category 3: Morning (8-12 am) in the production space ( )Production( - In the production space during the afternoon and evening (2-6 pm) )Production( )
[0074] Activity Category 4: All-day (8:00-22:00) at home )Production( )
[0075] Activity Category 5: Morning (8-12 PM) outside the village ( )Production( - In the afternoon and evening (2-6 pm) outside the village ( )Production( )
[0076] The final scatter plot of the clustering mathematical structure and its feature recognition results are attached. Figure 3 As shown.
[0077] Step 3: Using the results of individual attribute clustering and activity pattern clustering, couple the two to identify specific types of "group-activity" association features.
[0078] Step 3.1: Taking Group 1 as an example, use one-way ANOVA, and use all the secondary indicators and activity patterns under the five indicators reflecting the characteristics of the group (including: total housing area, number of rooms, number of people living together, monthly family income, and age) as variables in the calculation. Through one-sample ANOVA and Welch's ANOVA, identify the associated secondary indicators that have significant differential effects on the five activity patterns of Group 1.
[0079] Step 3.2: Calculate the Cohen's f values of the associated secondary indicators with significant differential effects, and extract the primary individual attribute indicators corresponding to secondary indicators with moderate correlation (Cohen's f ≥ 0.25) or higher; for any primary indicator, calculate the weight of all its secondary indicators (…). And calculate the sum of their proportions in descending order. );by ≥ The secondary indicators included in (=70%) serve as the characteristics of this indicator; traversing all primary indicators yields group category 1, "Middle to High Total Housing Area (200-300m²)". 2 The group category is categorized as follows: High number of rooms (4-8), low education level (junior high school), high monthly family income (3000-10000+), middle-aged and elderly (45-70 years old), and few family members engaged in agricultural activities (0-2 people). After coupling the "group + activity" model, indicators with moderate correlation (education level, age) are retained. Taking activity patterns 1-3 of this group category as an example, the main characteristics after coupling are summarized as follows:
[0080] People with low levels of education (primary / junior high school) and middle-aged (30-57 years old) + morning (8-12 pm) in public facilities ( )Production( - In the afternoon and evening (2-6 pm) at public facilities ( )Production( )
[0081] Low level of education (primary / junior high school) middle-aged and elderly (50-70 years old) + morning (10-12 pm) at home )housework( —Evening (4–6 PM) at home )housework( )
[0082] People with low levels of education (primary / junior high school) and middle-aged and elderly people (30-63 years old) + morning (8-12 o'clock) in the production space ( )Production( - In the production space during the afternoon and evening (2-6 pm) )Production( )
[0083] Step 4: For individuals with low levels of education (primary / junior high school) who are middle-aged (30-57 years old) and spend the morning (8-12 am) at public facilities ( )Production( - In the afternoon and evening (2-6 pm) at public facilities ( )Production( Taking the sub-category of ")" as an example, extract those who have used public facilities at least once in a single day's activity chain. Analyze user data to identify the spatiotemporal status of users' use of community facilities.
[0084] Step 4.1: Extract "low level of education (primary / junior high school), middle-aged (30-57 years old) + morning (8-12 am) in public facilities ( )Production( - In the afternoon and evening (2-6 pm) at public facilities ( )Production( The ")" segmentation mode's main feature is the activity chain involving time periods, which is statistically analyzed for user public facilities during these time periods. The time distribution of usage frequency indicates that the peak usage period for this facility is in the afternoon (14:00-16:00); similarly, the following groups are identified: those with low educational attainment (primary / junior high school), middle-aged and elderly (30-63 years old), and those using the facility in the morning (8-12:00) in production spaces. )Production( - In the production space during the afternoon and evening (2-6 pm) )Production( "Production space" The peak usage time for the facility is in the morning (8:00-10:00).
[0085] Step 4.2: Using the above two types of facilities ( Centered on a 100-meter spatial unit, and using ArcGIS network analysis, a multi-ring service area is formed that integrates with the urban road network. The accessibility of the two types of facilities within each ring is calculated. The distribution of user numbers was analyzed, identifying concentration ranges of 0-300 meters and 100-400 meters. Following the above method, the two types of facilities were traversed and summarized. The spatiotemporal usage status of ) and the main characteristics of the individual attributes of its user groups.
[0086] The final result is a summary table containing different individual attribute characteristics, the concentrated time periods and concentrated distances for the use of any facility for various activities. In this example, we take the use of public facilities (plazas, green spaces) by the "low-education (primary / junior high school) middle-aged (30-57 years old)" group, and the use of production spaces (agricultural, forestry, and aquaculture land) by the "low-education (primary / junior high school) middle-aged and elderly (30-63 years old)" group as examples. The final identification results are shown in the table below:
[0087]
[0088] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0089] The foregoing has shown and described the basic principles, main features, and advantages of this disclosure. Those skilled in the art should understand that this disclosure is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of this disclosure. Various changes and modifications can be made to this disclosure without departing from its spirit and scope, and all such changes and modifications fall within the scope of this disclosure as claimed.
Claims
1. A method for identifying the spatiotemporal state of community facility use based on resident activity and attribute clustering, characterized in that, Includes the following steps: First-hand data was collected through a questionnaire survey. The scope of the questionnaire distribution was determined as follows: taking the geographical coordinates of any community facility in a certain community in the city as the center, combined with the actual urban road network, and using network analysis on the ArcGIS platform, a polygonal service area of 1000 meters or 15 minutes' walk was generated. The union of all community facility service areas in the community was calculated, and this was used as the scope of the questionnaire distribution. The data collected by the questionnaire survey specifically includes: individual attribute data of users and activity logs of a typical day, with statistical units of two hours to count the types of user activities and their corresponding activity facilities; in addition, the housing location of the respondents is recorded and the corresponding geographical coordinates are queried; the questionnaire data is screened for validity, and data with complete individual attributes and activity logs are retained; the number of questionnaires is no less than 2‰ to 3‰ of the number of resident users within the scope of questionnaire distribution; The obtained valid questionnaire data were used to cluster individual attributes in order to initially screen group characteristics; Using the activity log data from the valid questionnaires, we clustered the users' activity patterns and extracted the main features. By using the results of individual attribute clustering and activity pattern clustering, the two are coupled to identify specific types of "group-activity" association features; By utilizing the obtained "group-activity" correlation features, the spatiotemporal status of users' use of community facilities is identified using frequency + time and frequency + space as core indicators.
2. The method for identifying the spatiotemporal status of community facility use based on resident activity and attribute clustering according to claim 1, characterized in that, The criteria for selecting valid questionnaire data are as follows: Valid data are questionnaires with no blank information between 4:00 and 24:00 on a typical day in the activity log and with complete individual attribute information. Individual attribute information includes, but is not limited to: gender, age, ethnicity, education level, occupation type, employer type, family structure, income level, consumption level, employment status, housing situation, housing location, number of commonly used transportation tools and their usage frequency.
3. The method for identifying the spatiotemporal state of community facility use based on resident activity and attribute clustering according to claim 1, characterized in that, The calculation process for individual attribute clustering and initial screening of group characteristics is as follows: First, the clustering results of individual attributes are calculated using the K-means algorithm; then, the group features are extracted using the random forest algorithm.
4. The method for identifying the spatiotemporal status of community facility use based on resident activity and attribute clustering according to claim 3, characterized in that, The calculation process for the initial screening of group features after clustering individual attributes is as follows: The weight contribution of feature indicators in the random forest algorithm results Sort the indicators in descending order and calculate the sum of their cumulative weighted contributions. ,by The n indicators at time represent the dimensions of group characteristics, where 85% ≤ ≤95%; then, calculate the proportion of the secondary indicator for each of the above n indicators. They are then sorted from highest to lowest according to their proportion, and the cumulative proportion is calculated. ,by The meaning represented by each of the j secondary indicators is the group characteristic described by the indicator, where 80% ≤ ≤90%.
5. The method for identifying the spatiotemporal state of community facility use based on resident activity and attribute clustering according to claim 1, characterized in that, The calculation process for activity pattern clustering and main feature extraction is as follows: First, the activity log data is encoded to form a continuous activity chain containing "time_space_activity"; then, the DBSCAN algorithm is used to cluster the user's spatiotemporal activity chain to obtain activity patterns; finally, the random forest algorithm is used to identify the main features of any activity pattern.
6. The method for identifying the spatiotemporal state of community facility use based on resident activity and attribute clustering according to claim 5, characterized in that, The coding rules for activity logs are as follows: Using 2-hour time units, starting from 4:00, the 10 time periods are named sequentially. ; Naming of Class i daily activities Naming the m-class activity space Similarly, the activity chain is encoded as " ".
7. The method for identifying the spatiotemporal state of community facility use based on resident activity and attribute clustering according to claim 5, characterized in that, The calculation method for extracting the main features after activity pattern clustering is as follows: The weight contribution of feature indicators is calculated using the random forest algorithm. ; Identification Given w independent variables, calculate the sum of the weighted contributions of these w independent variables. ,by The intersection of "space_activity" among the v independent variables at time is the main feature of the activity pattern, where 80% ≤ ≤90%.
8. The method for identifying the spatiotemporal state of community facility use based on resident activity and attribute clustering according to claim 1, characterized in that, The method for identifying the association features of specific types of "groups-activities" is as follows: First, using one-way ANOVA, for any population cluster type, all secondary indicators and activity patterns under the aforementioned n indicators reflecting group characteristics are used as variables in the calculation. Then, through one-sample ANOVA and Welch's ANOVA, the associated secondary indicators that have significant differential effects on activity patterns are identified. Next, Cohen's f-value is calculated to measure the degree of association. Finally, the main characteristics of the specific type of "group-activity" are finely characterized.
9. The method for identifying the spatiotemporal state of community facility use based on resident activity and attribute clustering according to claim 8, characterized in that, The following are methods for finely characterizing the main features of specific types of "group-activity": For the population clustering results under any activity pattern, extract the primary individual attribute indicators corresponding to the secondary indicators with moderate or higher correlation in the aforementioned steps; for any primary indicator, calculate the weight of all its secondary indicators. And calculate the sum of their proportions in descending order. ;by ≥ The sub-indicators included serve as characteristics of this indicator, where 65% ≤ ≤90%; Traverse all primary indicators to obtain the main characteristics of the "group + activity" coupling under the human element.
10. The method for identifying the spatiotemporal state of community facility use based on resident activity and attribute clustering according to claim 1, characterized in that, The process for identifying the spatiotemporal status of user community facilities is as follows: First, we analyze the spatiotemporal state of any facility's use from two aspects: temporal change characteristics and spatial concentric change characteristics. Then, we use the same method to traverse other facilities and summarize the individual attributes of the users of any community facility and their corresponding spatiotemporal usage status. The method for analyzing the spatiotemporal usage status of a facility is as follows: For any "group-activity" sub-category, user data that has used the facility at least once in a single day's activity chain is extracted and analyzed. Regarding the characterization of temporal variation features, based on the main features of the "group-activity" sub-pattern described in the previous steps, the time periods involved in the activity chains from the main features are extracted, and the temporal distribution of the proportion of user facility usage frequency during these time periods is statistically analyzed to identify peak periods. Regarding the characterization of spatial concentric variation features, with the facility as the center, a radius of 100 meters, and a boundary of 1000 meters within a 15-minute walk from the community living circle, ArcGIS network analysis is used to form a multi-ring service area integrated with the urban road network. The radius interval between each service area ring is 100 meters. The proportion distribution of the number of users accessing the facility within each ring is calculated, i.e., the proportion of users using the facility within a certain ring out of all users using the facility, thus identifying the concentration distance range.