A method for identifying crowd types based on venue code data and clustering integration
Through multi-feature extraction and clustering integration methods, combined with site code data and three classic clustering algorithms, the problem of relying on prior knowledge in the existing technology is solved, and efficient and accurate population type recognition is achieved.
Patent Information
- Application Number
- CN202211608733.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-15
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-12-15
AI Technical Summary
When using place code data to identify population types, the prior art relies on high-cost prior knowledge and low recognition efficiency, making it difficult to accurately and timely analyze the regional characteristics and travel rules of the population.
A multi-feature extraction method based on site code data is adopted, and three classic clustering algorithms (K-means, DBSCAN, and condensate hierarchical clustering) are combined for cluster analysis, and the feature importance is calculated by the integration method of weighted voting to generate the final clustering result.
It improves the accuracy and stability of population type identification, reduces the dependence on prior knowledge, realizes efficient unsupervised learning, and can be widely used in multiple cities.
Smart Images

Figure CN116070124B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data mining, and relates to a method for identifying population types based on venue code data and clustering integration. Background Art
[0002] In recent years, with the rapid economic development, activities such as business trips, tourism, and study among people between cities and between districts and counties have been increasing day by day, posing higher requirements for urban management issues such as urban infrastructure resource allocation, urban public security, and crowd evacuation, and also bringing a series of problems; how to utilize big data resources and accurately and timely grasp the regional characteristics of different populations based on machine learning technology, and analyze the travel patterns of the population is of great significance for improving the management ability of smart cities and has become one of the current research and application hotspots.
[0003] With the popularization of venue codes, venue codes have been deployed and applied in front of venues such as catering and accommodation, shopping, hospitals, schools, communities, cultural tourism, transportation, and office buildings; due to the daily living needs of residents, venue code data contains a large number of residents' travel records, which has the advantages of high sample volume, low cost, and wide coverage, and can relatively completely record the travel information of users every day; therefore, by mining the user information contained in venue code data and analyzing the travel patterns of users to identify the population types of users is a low-cost and high-efficiency means and has broad application prospects. Summary of the Invention
[0004] In order to overcome the problems existing in the above-mentioned prior art, the purpose of the present invention is to provide a method for identifying population types based on venue code data and clustering integration; based on venue code data, this method first extracts relevant features of population travel from different perspectives and analyzes the characterization ability of different features for population types; then uses three classic clustering algorithms to perform clustering analysis on different features respectively, selects the best clustering algorithm for each feature, and generates base clustering members; finally, through the integration method of weighted voting, assigns corresponding weights to the base clustering members according to the importance of the features, and integrates the results of the base clustering to obtain the final clustering result to achieve the identification of population types; the present invention uses machine learning methods to divide and identify population types, greatly utilizes the user information contained in venue code data, reduces the need for prior knowledge, reduces human intervention, and improves the objectivity and stability of the method.
[0005] A method for identifying population types based on venue code data and clustering integration includes the following steps:
[0006] A Data Preprocessing and Feature Extraction
[0007] S1: Data Preprocessing
[0008] (1) Desensitization processing: Personal information such as machine numbers, ID numbers, and home addresses is relatively sensitive, and such data is removed from the original data;
[0009] (2) Data cleaning: Data with the same scanning time or within a 5-minute time interval is merged into one piece of data; Data with missing fields is removed;
[0010] S2. Feature extraction
[0011] After the above data preprocessing steps, the venue code data mainly consists of user ID, scanning time, scanning type, venue type, and geographical location;
[0012] (1) Statistical feature extraction: The features are scanning frequency, travel frequency, and entry frequency;
[0013] (2) Time feature extraction: A day can be divided into eight time periods: early morning, morning, forenoon, noon, afternoon, evening, night, and late at night;
[0014] (3) Spatial feature extraction: The longitude and latitude are classified by features to obtain category features; Then use the features after the Cartesian product of the living space; Finally, the subspaces are encoded to obtain category features;
[0015] B. Selection of base clusterers
[0016] Three classic clustering algorithms, namely K-means based on partitioning, DBSCAN based on density, and agglomerative hierarchical clustering based on hierarchical clustering, are used to perform clustering analysis on different features respectively, and the corresponding best base classifiers are selected for different features through comparative analysis;
[0017] C. Identification of crowd types based on clustering ensemble algorithms
[0018] (1) Calculation of integration weights based on F-Ratio
[0019] F-Ratio measures the linear discriminant ability of feature X by the square ratio of the within-class difference to the between-class difference of the feature j The calculation formula is shown in Equation (1);
[0020]
[0021] In the formula, m j(c) and are the mean and variance of the c-class feature X j respectively, which means that the larger the F-Ratio score of the feature, the stronger the discrimination ability of the feature;
[0022] Based on obtaining the F-Ratio scores of all features, the weights of different categories of features are calculated to reflect the representation ability of the road surface condition. Assume the base model is M j (j = 1, 2,..., m), which uses k (k = 1, 2,..., n) features, and the F-Ratio score of the k-th feature is FS i =(i = 1, 2,..., k); The six calculation methods are as follows:
[0023] a) Summation method: Simply add the F-Ratio scores of all features, and its calculation formula is as follows:
[0024]
[0025] b) Logarithmic summation method: First sum the F-Ratio scores of all features, and then perform a logarithmic transformation. The calculation formula is as follows:
[0026]
[0027] c) Sum of squares method: Square the F-Ratio score of each feature first and then sum them. Its calculation formula is as follows:
[0028]
[0029] d) Logarithmic sum of squares method: Similar to the logarithmic summation method, first obtain the sum of squares of the F-Ratio scores of all features, and then perform a logarithmic transformation on it. Its calculation formula is as follows:
[0030]
[0031] e) Square root summation method: First obtain the square root of the F-Ratio score of each feature, and then sum them. Its calculation formula is as follows:
[0032]
[0033] f) Logarithmic square root summation method: Perform a logarithmic transformation on the sum of the square roots of the F-Ratio scores of the features. Its calculation formula is as follows:
[0034]
[0035] Obtain w j After that, perform a normalization process on it to obtain the weights of the base clustering model M j
[0036]
[0037] (2) Clustering ensemble strategy
[0038] Using the Bagging algorithm with the weighted voting integration method, first input the three types of samples into their corresponding base clusterers for training in parallel, and then, based on the weights obtained above, use the weighted voting method to integrate the results of each base clusterer to obtain the final result of crowd type recognition.
[0039] The specific steps of the clustering integration strategy are as follows:
[0040] Step 1: Input the three different sample data sets into their corresponding base clusterers for training respectively;
[0041] Step 2: After training, output the recognition results of different features for the crowd type respectively, and use to represent the probability that the nth type of feature identifies the crowd type as i;
[0042] Step 3: Take the feature weight calculated based on the F-Ratio as the weight of the base clusterer;
[0043] Step 4: Perform a linear weighted combination of the results of each base clusterer with their corresponding weights, synthesize the characterization capabilities of different features for the crowd type, and achieve the recognition of different crowds.
[0044] Compared with the existing technologies, the beneficial effects of the present invention are as follows:
[0045] (1) Wide data source: The data source of the present invention is the venue code data, which has the advantages of high sample volume, low cost, and wide coverage. Moreover, the acquisition method is stable and mature, and it can relatively completely record the travel information of users every day. It has general applicability to multiple cities and high accuracy.
[0046] (2) High accuracy in crowd type recognition: Clustering is a classic machine learning method with strong learning ability, which can discover different crowd types. The clustering integration method integrates the base clustering members, which can combine the advantages of single features and single clustering. The clustering integration algorithm proposed by the present invention can achieve a relatively high accuracy in crowd type recognition by fully exploring the characterization capabilities of different features for the crowd type.
[0047] (3) Low dependence on prior knowledge: It is difficult and time-consuming to obtain rich and complete prior knowledge of users. The clustering integration used in the present invention is an unsupervised learning method that does not require prior knowledge and can effectively distinguish unlabeled data, solving the problem of high dependence on prior knowledge and low efficiency in existing methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is the overall framework diagram of the crowd type recognition method based on venue code data and clustering integration;
[0049] Figure 2 It is the overall flowchart of the clustering ensemble algorithm. Specific implementation manner
[0050] The present invention will be described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0051] As Figure 1 shown, the method for identifying population types based on venue code data and clustering ensemble according to the embodiment of the present invention includes the following steps:
[0052] Perform preprocessing such as desensitization processing and data cleaning on the original data, and adopt three complementary feature extraction techniques to comprehensively mine the feature information of the venue code data, which maximally characterizes the association between travel habits and population types;
[0053] Use three classic clustering algorithms to perform clustering analysis on the extracted features, select the best clustering algorithm for each type of feature, and generate base clustering members;
[0054] Design a clustering ensemble method based on weighted voting according to the base clustering results; first, obtain the importance scores of the features through the F-Ratio method; then propose six integration weight calculation methods to obtain weights for the base clustering members; finally, integrate the results of the base clustering through the integration strategy of weighted voting to obtain the final clustering result, realizing the identification of population types.
[0055] The first part: Data preprocessing and feature extraction
[0056] The venue code data includes the user's personal basic information (mobile phone number, ID number, home address, household register), scanning time, venue type, scanning type, and geographical location, etc.; among them, the scanning time refers to the specific time when the user enters each venue; the venue type refers to the type to which the current venue belongs, mainly including office buildings, communities, schools, shopping malls, hospitals, etc.; the scanning type is information unique to the community, mainly including entering the community and leaving the community; the geographical location refers to the longitude and latitude information of the current scanned venue;
[0057] 1. Data preprocessing
[0058] In order to protect the privacy of users and facilitate subsequent work, the following preprocessing operations are performed on the original data;
[0059] (1) Desensitization processing; personal information such as mobile phone numbers, ID numbers, and home addresses is relatively sensitive, so such data is removed from the original data;
[0060] (2) Data cleaning; data with the same scanning time or within a 5-minute time interval are merged into one data record; data with missing fields are excluded.
[0061] 2. Feature extraction
[0062] After the above data preprocessing steps, the venue code data mainly consists of user ID, scanning time, scanning type, venue type, and geographical location. Due to the complex data type and large scale of this data, directly performing clustering analysis on the original data not only has low efficiency but also seriously affects the clustering effect. Therefore, how to effectively represent the original data and improve the expressiveness of the data for the real situation is of great significance for the identification of population types. The present invention performs feature extraction from three different perspectives, fully excavating the internal information of the venue code data and maximizing the representation of the association between travel habits and population types. Each feature extraction method is as follows:
[0063] (1) The purpose of statistical feature extraction is to reflect the activity patterns of different users through statistical analysis of the scanning data of a user over a period of time. Such features mainly include three types: scanning frequency, travel frequency, and entry frequency. Scanning frequency refers to the total number of scans of a user over a period of time. The appearance frequency refers to the number of times the user's scanning type corresponds to travel, and the entry frequency is the number of times the user's scanning type is entry. Among them, the sum of the travel frequency and the entry frequency is equal to the scanning frequency.
[0064] (2) The purpose of time feature extraction is to identify the population type by mining the time pattern of user activities. According to the extraction method of discrete value time features, a day can be divided into eight time periods: early morning, morning, morning, noon, afternoon, evening, night, and late at night. The present invention converts the original scanning time data into corresponding data segments, which is beneficial to the subsequent experiments on the one hand and can better reflect the travel patterns of different populations on the other hand.
[0065] (3) The purpose of spatial feature extraction is to identify the population type by analyzing the activity trajectories of users. First, the longitude and latitude are classified into feature categories to obtain category features; then the features after the Cartesian product of the living space are used; finally, the subspaces are encoded to obtain category features.
[0066] Part Two: Selection of Base Clusterers
[0067] In order to fully utilize the characterization ability of different features for population types and the advantages of each clustering algorithm, three classic clustering algorithms, namely partition-based K-means, density-based DBSCAN, and hierarchical agglomerative clustering based on hierarchical clustering, are used to perform clustering analysis on different features respectively, and the best base classifier corresponding to different features is selected through comparative analysis.
[0068] Part III: Identification of Crowd Types Based on Clustering Ensemble Algorithm
[0069] After the above-mentioned feature extraction on the original data, a sample data set containing three different types of features, namely statistical features, time features, and spatial features, is constructed. To comprehensively represent different crowds, a method for identifying crowd types based on weighted voting ensemble is designed. The core is to determine the weights of the base clusters in the weighted voting ensemble according to the importance of the features, avoiding the limitations of manually assigning weights;
[0070] (1) Calculation of Ensemble Weights Based on F-Ratio
[0071] The essence of F-Ratio is to measure the linear discriminant ability of feature X by the square ratio of the within-class difference to the between-class difference of the feature j , and its calculation formula is shown as Equation (1);
[0072]
[0073] In the formula, m j(c) and are the mean and variance of feature X in c classes respectively, which means that the larger the F-Ratio score of the feature, the stronger the discriminant ability of the feature; j Based on obtaining the F-Ratio scores of all features, six methods for calculating ensemble weights are explored. By calculating the weights of different types of features, the representational ability for road surface conditions is reflected. Assume the base model is M
[0074] (j = 1, 2,..., m), which uses k (k = 1, 2,..., n) features, and the F-Ratio score of the k-th feature is FS j =(i = 1, 2,..., k); The six calculation methods are as follows: i g) Summation method: Simply add up the F-Ratio scores of all features, and its calculation formula is as follows:
[0075]
[0076]
[0077] h) Logarithmic summation method: First sum up the F-Ratio scores of all features, and then perform a logarithmic transformation. The calculation formula is as follows:
[0078]
[0079] i) Sum of squares method: Square the F-Ratio score of each feature first and then sum them up. The calculation formula is as follows:
[0080]
[0081] j) Logarithmic sum of squares method: Similar to the logarithmic summation method, first calculate the sum of squares of the F-Ratio scores of all features, and then perform a logarithmic transformation on it. The calculation formula is as follows:
[0082]
[0083] k) Square root summation method: First, calculate the square root of the F-Ratio score of each feature, and then sum them up. The calculation formula is as follows:
[0084]
[0085] l) Logarithmic square root summation method: Perform a logarithmic transformation on the sum of the square roots of the F-Ratio scores of the features. The calculation formula is as follows:
[0086]
[0087] Obtain w j After that, perform normalization on it to obtain the weights of the base clustering model M j of
[0088]
[0089] (2) Clustering ensemble strategy
[0090] Drawing on the idea of the Bagging algorithm, a weighted voting ensemble method is adopted to propose a simple and effective clustering ensemble strategy. First, the three types of samples are respectively input into the corresponding base clusterers for training in parallel, and then, according to the weights obtained above, the results of each base clusterer are integrated by weighted voting to obtain the final result of crowd type recognition;
[0091] Participate in Figure 2 , the specific steps of the above clustering ensemble strategy are as follows:
[0092] The first step: Input the three different sample data sets into their corresponding base clusterers for training respectively;
[0093] The second step: After training, output the recognition results of different features for the crowd type respectively. Use C n i to represent the probability that the nth type of feature recognizes the crowd type as i;
[0094] The third step: Use the feature weights calculated based on F-Ratio as the weights of the base clusterers;
[0095] Step 4: Linearly weight and combine the results of each base clusterer with their corresponding weights, synthesize the representation capabilities of different features for population types, and achieve the recognition of different populations;
[0096] Embodiment
[0097] To verify the effectiveness and reliability of the present invention, this example analyzes and illustrates the data of the venue codes in the Lanzhou market in June 2022;
[0098] Step 1: Data preprocessing and feature extraction
[0099] After the original data undergoes preprocessing operations, a sample data set consisting of user IDs, code scanning times, code scanning types, code scanning location information, etc. is obtained, as shown in Table 1;
[0100] Table 1 Basic attributes of venue code data
[0101]
[0102]
[0103] After the original data undergoes preprocessing operations and the above-mentioned feature extraction method, three different types of features, namely statistical features, time features, and spatial features, are extracted from the original data; the specific feature information is shown in Table 2;
[0104] Table 2 Related feature types and interpretations of venue code data
[0105]
[0106] Second: Selection of base clusterers
[0107] Analyze the results of the base clustering, obtain the category induction table of the base clustering members, and select the corresponding best base clusterer for different features; as shown in Table 3, the K-means method has a better clustering effect on statistical features and has good clustering results for categories 1, 2, and 3; the agglomerative hierarchical clustering has the best effect on time features and has a good clustering effect on categories 2 and 3; the DBCAN method has the best effect on spatial features and has a good clustering effect on categories 1 and 3; Table 4 shows the clustering accuracies of several methods on different features, and the best-performing one is the K-means method;
[0108] Table 3 Accurate category induction of the division of base clustering members
[0109]
[0110]
[0111] Table 4 Accuracies of different clustering algorithms on different features
[0112]
[0113] Third: Clustering integration
[0114] To comprehensively consider the characterization ability of various features for pavement anomalies, this paper explored six methods for calculating the weights of base clusterers based on the importance of features. As shown in Table 5, among the six weight calculation methods, the Square method has the best clustering integration effect, with an accuracy rate reaching 95%, which is 5.1% higher than the best K-means method among them.
[0115] Table 5 Accuracy Rates of Six Weight Calculation Methods
[0116]
[0117] Table 6 shows the weights assigned to different features (base clusterers) by the six weighting methods. Among them, the weight distribution calculated by the Square method is relatively uniform, with statistical features accounting for 41%, time features accounting for 29%, and spatial features accounting for 30%. This shows the effectiveness of the method of the present invention. The clustering results of population types are not determined solely by a certain type of feature. Considering the performance of various features comprehensively can effectively improve the recognition of population types.
[0118] Table 6 Weight Assignment Results of Six Weighting Methods for Different Features
[0119]
[0120] The key technical points of the present invention are as follows:
[0121] (1) Multi-feature extraction technology
[0122] The multi-feature extraction methods disclosed in the present invention have strong complementarity with each other, excavate the internal information contained in the original data from different angles, restore the association between the user's scanning code rules and the population to the greatest extent, and have good anti-interference ability. Compared with the existing methods, the method disclosed in the present invention is a non-linear fusion of multi-features. Three classic clustering algorithms are used to perform initial clustering analysis on different features. By analyzing and comparing, the most suitable clustering algorithm is selected for each type of feature, and the base clustering results are integrated into an optimal clustering integration result through an integrated manner. The existing multi-feature extraction and fusion technology simply linearly combines multiple features to generate a high-dimensional feature vector, and performs clustering analysis on this feature vector. It does not fully excavate the differences between different features, and the dimension of the feature vector is relatively high, resulting in unsatisfactory computational consumption, accuracy, and stability of the clustering analysis.
[0123] (2) Clustering integration technology
[0124] The weighted clustering ensemble technology disclosed in the present invention determines the ensemble weights of the base clusterers according to the importance degree of features, avoiding the limitations of assigning weights manually; through the ensemble method of weighted voting, the initial clusters are reasonably integrated to obtain the final result, improving the stability and accuracy of the clustering analysis result; in addition, among the six ensemble weight algorithms discussed, the square (square root) application increases (decreases) the importance difference between each feature, which makes the importance of the base model more (less) dependent on the first few top-ranked features; similarly, the logarithm also helps to reduce the importance of the base model being dominated by the first few top-ranked features.
Claims
1. A method for identifying crowd types based on venue code data and clustering integration, characterized in that It includes the following steps: A. Data preprocessing and feature extraction S1: Data preprocessing (1) Desensitization processing: Personal information such as machine numbers, ID numbers, and home addresses is relatively sensitive and is removed from the original data; (2) Data cleaning: Data with the same scanning time or a time interval within 5 minutes is merged into one piece of data; data with missing fields is removed; S2. Feature extraction After the above data preprocessing steps, the venue code data mainly consists of user ID, scanning time, scanning type, venue type, and geographical location; (1) Statistical feature extraction: The features are scanning frequency, travel frequency, and entry frequency; (2) Time feature extraction: A day can be divided into eight time periods: early morning, morning, forenoon, noon, afternoon, evening, night, and late at night; (3) Spatial feature extraction: The longitude and latitude are classified by features to obtain category features; then the features after the Cartesian product of the living space are used; finally, the subspaces are encoded to obtain category features; B. Selection of base clusterers Three classic clustering algorithms, namely partition-based K-means, density-based DBSCAN, and agglomerative hierarchical clustering based on hierarchical clustering, are used to perform clustering analysis on different features respectively. Through comparative analysis, the corresponding best base classifier is selected for different features; C. Identification of population types based on the clustering ensemble algorithm (1) Calculation of integration weights based on F-Ratio The F-Ratio measures the linear discriminant ability of feature X by the squared ratio of the within-class difference to the between-class difference of the feature, and its calculation formula is shown in Equation (1); j The linear discriminant ability of feature X is measured by the squared ratio of the within-class difference to the between-class difference of the feature, and its calculation formula is shown in Equation (1); where m j(c) and are the mean and variance of the c-type feature X j respectively, which means that the larger the F-Ratio score of the feature, the stronger the discrimination ability of the feature; Based on obtaining the F-Ratio scores of all features, the weights of different categories of features are calculated to reflect the representation ability of the road surface condition. Assume the base model is M j (j = 1, 2,..., m), which uses k (k = 1, 2,..., n) features, and the F-Ratio score of the k-th feature is FS i =(i = 1, 2,..., k); The six calculation methods are as follows: a) Summation method: Simply add the F-Ratio scores of all features, and its calculation formula is as follows: b) Logarithmic summation method: First sum the F-Ratio scores of all features, and then perform logarithmic transformation. The calculation formula is as follows: c) Sum of squares summation method: Square the F-Ratio score of each feature first and then sum them. The calculation formula is as follows: d) Logarithmic sum of squares method: Similar to the logarithmic summation method, first obtain the sum of squares of the F-Ratio scores of all features, and then perform logarithmic transformation on it. The calculation formula is as follows: e) Square root summation method: First obtain the square root of the F-Ratio score of each feature, and then sum them. The calculation formula is as follows: f) Logarithmic square root summation method: Perform logarithmic transformation on the sum of the square roots of the F-Ratio scores of the features. The calculation formula is as follows: Obtain w j After that, perform normalization on it to obtain the weight of the base aggregation model M j (2) Clustering integration strategy Using the Bagging algorithm with the weighted voting integration method, first input the three types of samples into their corresponding base clusterers for training in parallel, and then, according to the weights obtained above, use the weighted voting method to integrate the results of each base clusterer to obtain the final result of population type identification.
2. The method for identifying crowd types based on venue code data and clustering integration according to claim 1, wherein The specific steps of the clustering integration strategy are as follows: The first step: Input the three different sample data sets into their corresponding base clusterers for training respectively; Step 2: After training, output the recognition results of different features for population types respectively. Use to represent the probability that the nth type of feature identifies the population type as i. Step 3: Use the feature weights calculated based on the F-Ratio as the weights of the base clusterers; The fourth step: Perform linear weighted combination on the results of each base clusterer with their corresponding weights, and comprehensively consider the characterization ability of different features for population types to achieve the identification of different populations.
Citation Information
Patent Citations
Crowd type identification method based on mobile phone signaling data
CN110245981A
Risk crowd classification method, device and system, electronic device and storage medium
CN114220555A