A method, system and device for clustering and analyzing a population based on spatio-temporal data

Through the analysis of multi-source spatiotemporal data and the application of density clustering algorithms, the one-sided information caused by single data source analysis is solved, and more accurate and comprehensive crowd clustering analysis is achieved.

CN119513636BActive Publication Date: 2025-05-30CETC BIGDATA RES INST CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510097821.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-30
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

Spatial-temporal data analysis based on a single type of data source can only obtain one-sided information, resulting in the incomplete results of clustering analysis, which affects the accuracy of decisions.

Method used

By obtaining various types of spatiotemporal data, including time data, spatial data, traffic data, environmental data and socio-economic data, and clustering analysis of the population through preset analysis algorithms and density clustering algorithms, a spatiotemporal data matrix is ​​constructed for dimensionality reduction and feature extraction.

Benefits of technology

This method can reflect the characteristics of the population from multiple angles, improve the accuracy and comprehensiveness of clustering analysis, and avoid the one-sided problem based on a single type of data source analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119513636B_ABST
    Figure CN119513636B_ABST
Patent Text Reader

Abstract

The present application discloses a method, system and device for clustering and analyzing a population based on spatio-temporal data. The method of the present application includes: obtaining spatio-temporal data; performing feature processing on the spatio-temporal data to obtain the number of features, and constructing a spatio-temporal data matrix based on the number of features; reducing the dimension of the spatio-temporal data matrix through a preset analysis algorithm to obtain a data matrix; calculating the volume of a preset radius centered on a preset point and the number of points within the radius according to the preset points in the data matrix, and calculating the local density based on the number of points within the radius and the volume; calculating the distance between each point in the data matrix and the nearest target point, and selecting a target neighborhood radius based on the distance and the local density; determining the data dimension based on the number of features, and setting the minimum number of points through the data dimension; marking each point in the data matrix through a density clustering algorithm according to the target neighborhood radius and the minimum number of points to obtain a clustering result; analyzing the features of the clustering result based on multiple clustering clusters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer clustering analysis, and particularly to a method, system and device for clustering and analyzing people based on spatio-temporal data. Background Art

[0002] With the development of cities, spatio-temporal data is of great significance in studying the behavioral characteristics of urban populations. The sources of spatio-temporal data are extensive and diverse, such as taxi GPS trajectory data, user call data, etc. These spatio-temporal data can reflect various aspects of information such as the activity range, travel habits, and social connections of people. Through in-depth mining of spatio-temporal data, urban layout can be better planned, traffic management optimized, and the level of public services improved, etc.

[0003] In the prior art, many studies have proposed methods for analyzing characteristics such as human behavior and human group mobility for different types of spatio-temporal data. Such analysis usually starts from a single type of data, processes the outliers in this type of data, and then analyzes the processed data in combination with algorithm models. For example, from one aspect of human behavior characteristics, the pattern of crowd movement trajectories is analyzed.

[0004] However, data analysis based on a single type of data source can only obtain one-sided information, which will make the results obtained from clustering analysis not comprehensive enough, resulting in deviations in the understanding of human characteristics, and further affecting the accuracy of decision-making. Summary of the Invention

[0005] To solve the above technical problems, this application provides a method, system and device for clustering and analyzing people based on spatio-temporal data.

[0006] The technical solutions provided in this application are described below:

[0007] In the first aspect of this application, a method for clustering and analyzing people based on spatio-temporal data is provided. The method includes:

[0008] Obtain spatio-temporal data, where the spatio-temporal data includes time data, space data, traffic data, environmental data, and socio-economic data;

[0009] Perform feature processing on the spatio-temporal data to obtain the number of features, and construct a spatio-temporal data matrix based on the number of features;

[0010] Reduce the dimension of the spatio-temporal data matrix through a preset analysis algorithm to obtain a data matrix;

[0011] According to the preset points in the data matrix, calculate the volume of a preset radius centered on the preset points and the number of points within the radius, and calculate the local density based on the number of points within the radius and the volume;

[0012] Calculate the distance between each point in the data matrix and the nearest target point, and select the target neighborhood radius based on the distance and the local density;

[0013] Determine the data dimension based on the number of features, and set the minimum number of points through the data dimension;

[0014] Mark each point in the data matrix through a density clustering algorithm according to the target domain radius and the minimum number of points to obtain a clustering result, where the clustering result includes multiple clustering clusters;

[0015] Analyze the characteristics of the clustering result based on the multiple clustering clusters.

[0016] Optionally, the obtaining of the data matrix by reducing the dimension of the spatio-temporal data matrix through a preset analysis algorithm includes:

[0017] Calculate the mean vector through the spatio-temporal data matrix;

[0018] Calculate the covariance matrix according to the spatio-temporal data matrix and the mean vector;

[0019] Perform eigenvalue decomposition on the covariance matrix to obtain eigenvalues and eigenvectors, and construct principal components based on the eigenvectors, where the eigenvalue is the variance explained by each principal component, and the eigenvector is the direction of the principal component;

[0020] Select the target principal component according to the magnitude of the variance, and construct an eigenvector matrix based on the eigenvectors of the target principal component;

[0021] Project the spatio-temporal data onto the eigenvector matrix to obtain the data matrix.

[0022] Optionally, the covariance matrix is calculated by the following formula:

[0023] ;

[0024] Wherein, represents the covariance matrix in the feature and the feature between, represents the number of samples of the spatio-temporal data, represents the th sample in the feature value, represents the th sample in the feature value, represents the th sample of the mean vector, Denote the mean vector of the th sample.

[0025] Optionally, the selecting the target principal component according to the magnitude of the variance and constructing the eigenvector matrix based on the eigenvectors of the target principal component includes:

[0026] Calculating the cumulative variance ratio of each principal component through the variance;

[0027] Selecting the target principal component based on the magnitude of the cumulative variance ratio;

[0028] Constructing the eigenvector matrix based on the eigenvectors of the target principal component.

[0029] Optionally, the selecting the target principal component based on the magnitude of the cumulative variance ratio includes:

[0030] Judging whether the cumulative variance ratio is greater than or equal to 90%;

[0031] If so, determining it as the target principal component.

[0032] Optionally, the marking each point in the data matrix through the density clustering algorithm according to the target domain radius and the minimum number of points to obtain a clustering result, where the clustering result includes multiple clustering clusters includes:

[0033] Step 1: Determine the data matrix as the sample set D=(x1,x2,...,xm), the neighborhood parameter (ϵ, MinPts) and the sample distance metric, where ϵ represents the target domain radius, MinPts represents the minimum number of points, and the sample set is a set composed of multiple TF-IDF vectors;

[0034] Step 2: Calculate the ϵ-neighborhood subsample set Nϵ(xj) of each sample Xj based on the sample distance metric;

[0035] Step 3: Compare the absolute value |Nϵ(xj)| of the ϵ-neighborhood subsample set Nϵ(xj) with the MinPts, and add the sample xj greater than the MinPts to the core object sample set Ω;

[0036] Step 4: When the core object sample set Ω is not empty, randomly select a core object o from the core object sample set Ω and execute the following algorithm:

[0037] Initialize the current cluster core object queue Ωcur={o};

[0038] Initialize the category serial number k=k+1;

[0039] Initialize the current cluster sample set Ck={o};

[0040] Update the unvisited sample set Γ = Γ - {o};

[0041] Step Five: If the current cluster core object queue Ωcur is empty, the current clustering cluster Ck is generated; after generating the clustering cluster Ck, update the cluster partition C = C ∪ {Ck}, and update the core object sample set Ω = Ω - Ck;

[0042] Step Six: If the current cluster core object queue Ωcur is not empty, then execute the following algorithm:

[0043] Take out a core object o' from the current cluster core object queue Ωcur;

[0044] Determine all ϵ-neighborhood subsample sets Nϵ(o') through the neighborhood distance threshold ϵ;

[0045] Let Δ = Nϵ(o') ∩ Γ;

[0046] Update the current cluster sample set Ck = Ck ∪ Δ, and update the unvisited sample set Γ = Γ - Δ;

[0047] Update Ωcur = Ωcur ∪ (Δ ∩ Ω) - {o'};

[0048] Repeat Step Five;

[0049] Step Seven: Output the cluster partition C = {C1, C2,..., Ck} to obtain the clustering result, and the clustering result contains multiple clustering clusters.

[0050] Optionally, after obtaining the spatio-temporal data, where the spatio-temporal data includes time data, space data, traffic data, environmental data, and socio-economic data, before processing the spatio-temporal data to obtain the number of features and constructing a spatio-temporal data matrix based on the number of features, it further includes:

[0051] Perform data preprocessing on the spatio-temporal data, and the data preprocessing includes: data deduplication, missing value processing, and unifying data types.

[0052] Optionally, the local density is calculated by the following formula:

[0053] ;

[0054] Wherein, represents the number of points within the preset radius of the preset point ; represents the volume of the preset radius centered at the preset point ; Indicates the local density.

[0055] The second aspect of the present application provides a system for clustering and analyzing people based on spatio-temporal data, and the system includes:

[0056] An acquisition unit, configured to acquire spatio-temporal data, where the spatio-temporal data includes time data, space data, traffic data, environmental data, and socioeconomic data;

[0057] A construction unit, configured to perform feature processing on the spatio-temporal data to obtain the number of features, and construct a spatio-temporal data matrix based on the number of features;

[0058] A dimensionality reduction unit, configured to reduce the dimensionality of the spatio-temporal data matrix through a preset analysis algorithm to obtain a data matrix;

[0059] A calculation unit, configured to calculate the volume of a preset radius centered on the preset point and the number of points within the radius according to the preset points in the data matrix, and calculate the local density based on the number of points within the radius and the volume;

[0060] A selection unit, configured to calculate the distance between each point in the data matrix and the nearest target point, and select a target neighborhood radius based on the distance and the local density;

[0061] A setting unit, configured to determine the data dimension based on the number of features, and set the minimum number of points through the data dimension;

[0062] A marking unit, configured to mark each point in the data matrix through a density clustering algorithm according to the target domain radius and the minimum number of points to obtain a clustering result, where the clustering result includes multiple clustering clusters;

[0063] An analysis unit, configured to analyze the features of the clustering result based on the multiple clustering clusters.

[0064] The third aspect of the present application provides a device for clustering and analyzing people based on spatio-temporal data, and the device includes:

[0065] A processor, a memory, an input / output unit, and a bus;

[0066] The processor is connected to the memory, the input / output unit, and the bus;

[0067] The memory stores a program, and the processor calls the program to execute the method as described above.

[0068] From the above technical solutions, it can be seen that the present application has the following advantages:

[0069] In the method for clustering and analyzing a population based on spatio-temporal data provided in this application, various types of spatio-temporal data are specifically obtained, including time data, spatial data, traffic data, environmental data, and socioeconomic data. Finally, a preset analysis algorithm and a density clustering algorithm are used to perform clustering analysis on the population. The multi-source spatio-temporal data obtained by this method serves as the data basis for clustering analysis, avoiding the problem that only one-sided information can be obtained by analyzing a single type of data source, and helping to reflect population characteristics from multiple perspectives. For example, time data reflects the activity time pattern, and spatial data determines the geographical location, etc., effectively improving the accuracy of subsequent clustering analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] In order to more clearly illustrate the technical solutions in this application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0071] Figure 1 It is a schematic flowchart of an embodiment of the method for clustering and analyzing a population based on spatio-temporal data provided in this application;

[0072] Figure 2 It is a schematic flowchart of another embodiment of the method for clustering and analyzing a population based on spatio-temporal data provided in this application;

[0073] Figure 3 It is a schematic structural diagram of an embodiment of the system for clustering and analyzing a population based on spatio-temporal data provided in this application;

[0074] Figure 4 It is a schematic structural diagram of an embodiment of the device for clustering and analyzing a population based on spatio-temporal data provided in this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0075] This application provides a method for clustering and analyzing a population based on spatio-temporal data, which can improve the accuracy and comprehensiveness of clustering and analyzing the population. It should be noted that the method for clustering and analyzing a population based on spatio-temporal data in this application is applied to a terminal.

[0076] Spatio-temporal data is a type of data that integrates information on time and space dimensions. It can deeply reflect the development and change laws of things at specific geographical locations over time. In the field of transportation, the driving trajectories of vehicles on different roads at different times are typical spatio-temporal data. Through analysis, it is possible to grasp the peak and trough periods of traffic flow and the spatio-temporal distribution of congested sections, thereby providing a basis for traffic planning and the optimization of intelligent transportation systems; in the environmental field, the changes in meteorological data at different geographical locations over time, such as spatio-temporal data on temperature, humidity, pollutant concentration, etc., help scientists study the trends of climate change and the diffusion laws of environmental pollutants; in the field of urban planning, spatio-temporal data such as the distribution of the population and land use in the city can guide the rational layout and sustainable development of the city.

[0077] Cluster analysis is to automatically group according to the characteristics of the data itself. It is a data analysis method that groups data objects into different clusters, which helps to discover potential structures and laws from complex data sets. In practical applications, enterprises can conduct cluster analysis based on characteristics such as consumers' purchase behaviors, ages, and incomes, so as to divide consumers into different groups for formulating more targeted marketing strategies.

[0078] Please refer to Figure 1 , this application first provides an embodiment of a method for clustering and analyzing people based on spatio-temporal data, and this embodiment includes:

[0079] S101. Obtain spatio-temporal data, where the spatio-temporal data includes time data, space data, traffic data, environmental data, and socio-economic data;

[0080] In this embodiment, to obtain spatio-temporal data, different time data, space data, traffic data, environmental data, and economic data need to be collected from multiple data sources.

[0081] S102. Perform feature processing on the spatio-temporal data to obtain the number of features, and construct a spatio-temporal data matrix based on the number of features;

[0082] In this embodiment, to perform feature processing on the spatio-temporal data to obtain the number of features, first, feature extraction needs to be performed on the time data, space data, traffic data, environmental data, and socio-economic data obtained in step S101. Then, the number of extracted features is statistically analyzed. Based on the statistically analyzed number of features, a spatio-temporal data matrix is constructed with data objects as rows and features as columns. For example, if there are 100 data objects and the number of features is 20, then the constructed spatio-temporal data matrix is a 100-row and 20-column spatio-temporal data matrix.

[0083] In actual application scenarios, for time data, periodic features (such as daily, weekly, and monthly cycles), special time points (such as holidays, morning rush hours), etc. can be extracted; for spatial data, features related to geographical coordinates (such as longitude, latitude, altitude), regional types (such as commercial areas, residential areas), etc. can be extracted; for traffic data, features such as traffic volume, flow direction, traffic congestion level, etc. can be extracted; for environmental data, features such as temperature range, air quality level, etc. can be extracted; for social and economic data, features such as population income levels, consumption tendencies, etc. can be extracted.

[0084] S103. Reduce the dimensionality of the spatio-temporal data matrix through a preset analysis algorithm to obtain a data matrix;

[0085] In this embodiment, first, obtain the constructed spatio-temporal data matrix, and then select a suitable analysis algorithm to reduce the dimensionality of the spatio-temporal data matrix, thereby obtaining a reduced-dimensional data matrix. For example, if the original spatio-temporal data matrix has 20 features, after the dimensionality reduction operation, a data matrix with 5 features can be obtained.

[0086] It should be noted that the dimensionality reduction operation in this embodiment is a technical means of converting high-dimensional data into low-dimensional data. In a spatio-temporal data matrix containing time data, spatial data, traffic data, environmental data, and social and economic data, there are dozens or even hundreds of features to describe each data point. Then, there may be correlations between these features. The purpose of the dimensionality reduction operation is to find a mapping method to map the data points in the original high-dimensional space to a low-dimensional space. For example, the original spatio-temporal data matrix collects a large amount of information about the activities of people in area A through multiple sensors, including features in multiple different time periods, multiple different geographical locations, and many aspects such as traffic, environment, and social economy related to people's activities. Then, the dimensionality reduction operation in this example is to reduce the number of features describing the activities while retaining the main information of people's activities.

[0087] S104. Calculate the volume of a preset radius centered on a preset point in the data matrix and the number of points within the radius, and calculate the local density based on the number of points within the radius and the volume;

[0088] In this embodiment, first, it is necessary to determine a preset point in the data matrix. This preset point is a randomly selected data point, and an initial preset radius is set. Then, calculate the volume of the preset radius centered on this data point, and count the number of points within the preset radius centered on the preset point. Finally, calculate the local density based on the number of points within the obtained radius and the volume.

[0089] The local density can be calculated by the following formula:

[0090] ;

[0091] Among them, represents the preset point the preset radius of the number of points within represents taking the preset point as the center and the preset radius the volume of represents the local density.

[0092] S105. Calculate the distance between each point in the data matrix and the nearest target point, and select the target neighborhood radius based on the distance and the local density;

[0093] In this embodiment, to calculate the distance between each point in the data matrix and the nearest target point, for each point in the data matrix, the entire data matrix needs to be traversed to find its nearest target point. The distance between each point and other points can be compared to determine the nearest target point for each point. After determining the nearest target point, calculate the distance between each point and the nearest target point. When selecting the target neighborhood radius, the calculated local density reflects the density around the data point. When the local density of a point is high, the data points around it are relatively dense. Then, a suitable target neighborhood radius can be determined according to the numerical relationship between its distance from the nearest target point and the local density.

[0094] S106. Determine the data dimension based on the number of features, and set the minimum number of points through the data dimension;

[0095] In this embodiment, the data dimension is determined according to the number of features obtained in step S102. For example, if the number of features obtained after feature processing is d, then the data dimension is d. Then, the minimum number of points is set based on the determined data dimension. Further, the method of setting the minimum number of points is to determine the minimum number of points based on a specific formula related to the data dimension.

[0096] It should be noted that the data dimension refers to the number of features or variables describing the data, which is the degree of freedom or the number of coordinate axes of the data in space. The purpose of setting the minimum number of points through the data dimension in this step is to select a suitable threshold to distinguish noise points and core points.

[0097] The specific formula for setting the minimum number of points can be calculated by the following formula:

[0098] ;

[0099] Among them, represents the minimum number of points, represents the data dimension.

[0100] S107. Mark each point in the data matrix through a density clustering algorithm according to the target domain radius and the minimum number of points, obtaining a clustering result that contains multiple clusters;

[0101] A cluster is a set formed by data points clustering together according to similarity during the clustering analysis process. In the data space, data points with similar characteristics are grouped into the same cluster. For example, when analyzing customer consumption data, customer data with similar consumption habits and similar consumption amounts may form a cluster; in image recognition, pixel points with similar characteristics such as color and texture also form clusters. The data points within a cluster have a high degree of similarity to each other, while the data points between different clusters have significant differences.

[0102] In this embodiment, first, obtain the target neighborhood radius and the minimum number of points determined in step S106. Then, for each point in the data matrix, with this point as the center, determine its neighborhood range according to the target neighborhood radius, count the number of points within the neighborhood range, and mark each point through a density clustering algorithm. For example, if the number of points is greater than or equal to the minimum number of points, then this point is regarded as a core point, and the points within its neighborhood are marked as belonging to the same cluster as this core point. Then continue to check the neighborhood of the newly marked points. If there are also points that meet the minimum number of points requirement within their neighborhoods, then these points are also included in the current cluster. This process continuously iterates to expand the cluster until no new points can be added. For those points that are neither core points nor can be included in the cluster by the core points within their neighborhoods, they are marked as noise points. Mark each point in the data matrix through such a density clustering algorithm, and finally obtain a clustering result that contains multiple clusters and noise points.

[0103] The following provides a specific embodiment of obtaining multiple clusters through a density clustering algorithm:

[0104] Step 1: Determine the data matrix as the sample set D = (x1, x2,..., xm), the neighborhood parameters (ϵ, MinPts), and the sample distance measurement method, where ϵ represents the target domain radius, MinPts represents the minimum number of points, and the sample set is a set composed of multiple TF-IDF vectors;

[0105] Step 2: Based on the sample distance measurement method, calculate the ϵ-neighborhood subsample set Nϵ(xj) of each sample Xj;

[0106] Step 3: Compare the absolute value |Nϵ(xj)| of the ϵ-neighborhood subsample set Nϵ(xj) with MinPts, and add the sample xj greater than MinPts to the core object sample set Ω;

[0107] Step 4: When the core object sample set Ω is not empty, randomly select a core object o from the core object sample set Ω, and execute the following algorithm:

[0108] Initialize the current cluster core object queue Ωcur = {o};

[0109] Initialize the category serial number k = k + 1;

[0110] Initialize the current cluster sample set Ck = {o};

[0111] Update the unvisited sample set Γ = Γ - {o};

[0112] Step 5: If the current cluster core object queue Ωcur is empty, the current clustering cluster Ck is generated; after generating the clustering cluster Ck, update the cluster partition C = C ∪ {Ck}, and update the core object sample set Ω = Ω - Ck;

[0113] Step 6: If the current cluster core object queue Ωcur is not empty, then execute the following algorithm:

[0114] Take out a core object o' from the current cluster core object queue Ωcur;

[0115] Determine all ϵ-neighborhood subsample sets Nϵ(o') through the neighborhood distance threshold ϵ;

[0116] Let Δ = Nϵ(o') ∩ Γ;

[0117] Update the current cluster sample set Ck = Ck ∪ Δ, and update the unvisited sample set Γ = Γ - Δ;

[0118] Update Ωcur = Ωcur ∪ (Δ ∩ Ω) - {o'};

[0119] Repeat Step 5;

[0120] Step 7: Output the cluster partition C = {C1, C2,..., Ck}, obtain the clustering result, and the clustering result contains multiple clustering clusters.

[0121] Different from traditional algorithms such as K-means, the density clustering algorithm can form clusters of arbitrary shapes and adapt to the complex behaviors and distributions of people.

[0122] This algorithm can handle data with large density variations. In human behavior, there may be more behavioral characteristics during some time periods and less during other time periods, and this kind of variation can be well captured by this algorithm. Different from algorithms such as K-means, the density clustering algorithm does not need to pre-specify the number of clusters, avoiding the prior knowledge of the data structure.

[0123] S108. Analyze the characteristics of the clustering result based on multiple clustering clusters.

[0124] In this embodiment, for the obtained multiple clustering clusters, the basic statistical features of each clustering cluster are calculated respectively. The basic statistical features can calculate the central position of the clustering cluster and the number of data points within the clustering cluster, and the purpose is to measure the size of the clustering cluster. After the calculation is completed, the distribution of the data within the clustering cluster is analyzed. The statistical quantities such as variance can be calculated to describe the degree of dispersion of the data in each feature dimension, and then the characteristic differences between different clustering clusters such as the differences in the central positions, the differences in sizes, and the differences in data distributions of different clustering clusters are compared. Further, the relationship between the clustering cluster and the external environment can also be analyzed. For example, in spatio-temporal data, it is analyzed whether a certain clustering cluster is related to a specific geographical area or time period.

[0125] In this embodiment, by obtaining various types of spatio-temporal data, including time data, space data, traffic data, environmental data, and socio-economic data, finally, clustering analysis is performed on the population through a preset analysis algorithm and a density clustering algorithm.

[0126] Further, this embodiment uses the obtained multi-source spatio-temporal data as the data basis for clustering analysis, avoiding the problem that only partial information can be obtained by analyzing based on a single type of data source, and helping to reflect the characteristics of the population from multiple perspectives. For example, time data reflects the activity time pattern, and space data determines the geographical location, etc., effectively improving the accuracy of subsequent clustering analysis; the spatio-temporal data is subjected to feature processing to obtain the number of features and a spatio-temporal data matrix is constructed, making the relationship between multi-source data clearer, avoiding the chaos and disorder of data, and thus improving the efficiency and accuracy of the entire data analysis process, providing an effective data organization form for accurately understanding human behavior and group mobility characteristics; through the dimensionality reduction operation of the preset analysis algorithm, while retaining the main information of the data, the complexity and computational amount of the data are greatly reduced, helping to more accurately analyze characteristics such as human behavior and group mobility, and avoiding analysis deviations caused by overly complex data and excessive noise; by calculating the local density, it is used to mine the distribution characteristics of data points in the data matrix, and the distance from each point to a preset point is calculated, which helps to reflect the spatial relationship between data points and discover the clustering tendency of the data; calculating the distance between each point and the nearest target point helps to further understand the tightness relationship between data points and can reflect the microscopic structure of the data; selecting the target neighborhood radius based on the distance and local density can better adapt to the uneven distribution of the data, avoiding the problem that a fixed radius may lead to an overly large radius in a dense area and including too many irrelevant data points, and helping to improve the accuracy of subsequent data analysis.

[0127] Furthermore, in this embodiment, by determining the data dimension based on the number of features, the structural features of the data can be accurately grasped. The clear data dimension helps for the accurate understanding and processing of the data during subsequent analysis. By setting the minimum number of points according to the data dimension, it is used to judge whether a cluster has sufficient representativeness or effectiveness, which can avoid generating unreasonable clustering results and make the analysis of spatio-temporal data more scientific and reasonable. According to the target domain radius and the minimum number of points, each point in the data matrix is marked through a density clustering algorithm to obtain the clustering result, which helps to discover the spatial and temporal laws hidden in the data and identify anomalies in the data. Analyzing the characteristics of the clustering result based on multiple clusters helps to intuitively understand the aggregation center and aggregation degree of the data, so as to better utilize the data for decision-making services.

[0128] Please refer to Figure 2 , Figure 2 which is another embodiment of a method for clustering and analyzing people based on spatio-temporal data provided by this application. This embodiment includes:

[0129] S201. Obtain spatio-temporal data, where the spatio-temporal data includes time data, spatial data, traffic data, environmental data, and socio-economic data;

[0130] In this embodiment, step S201 is similar to step S101 of the previous embodiment and will not be elaborated here.

[0131] S202. Perform data preprocessing on the spatio-temporal data, where the data preprocessing includes: data deduplication, missing value processing, and unifying data types;

[0132] In this embodiment, for the obtained spatio-temporal data, first traverse each data record in the spatio-temporal dataset, compare the key identification fields or all fields in the data record through the data characteristics, and when identical records are identified, only keep one of them. Then for the fields with missing values, if the number of missing values is small, the mean, median, or mode can be used for filling; if it is time series data, interpolation filling can be performed according to the values of the previous and subsequent time points; if it is categorical data, the category with the highest frequency of occurrence can be used for filling. After completing the processing of duplicate values and missing values, then perform the operation of unifying data types, check the data types of each field in the dataset, determine the unified data type, and unify the data with different types. For example, unify all fields representing dates into the date type and unify the fields representing quantities into the numerical type.

[0133] The deduplication of spatio-temporal data in this embodiment can reduce data redundancy, avoid unnecessary repeated processing of the same data in subsequent analysis, thereby improving analysis efficiency and saving storage space; the processing of missing values in spatio-temporal data ensures data integrity and accuracy; unifying the data types of spatio-temporal data helps to perform correct arithmetic and comparison operations in the data analysis process and can avoid calculation errors or logical confusion caused by different data types.

[0134] S203. Perform feature processing on the spatio-temporal data to obtain the number of features, and construct a spatio-temporal data matrix based on the number of features;

[0135] In this embodiment, step S203 is similar to step S102 in the foregoing embodiment, and will not be elaborated here.

[0136] S204. Calculate a mean vector through the spatio-temporal data matrix;

[0137] In this embodiment, first clarify the structure of the spatio-temporal data matrix, where each row represents a spatio-temporal data point and each column represents a specific feature. Then calculate the mean for each column feature separately. For example, for numerical features, add up all the elements in the column and then divide by the number of elements to obtain the mean of the column feature. After calculating the mean of each column feature, combine these means to obtain the mean vector.

[0138] This embodiment calculates a mean vector through the spatio-temporal data matrix, which can provide the average level information of the spatio-temporal data in each feature dimension, and the calculated mean vector can also be used as a benchmark for comparing data differences in different spatio-temporal regions or different time periods.

[0139] S205. Calculate a covariance matrix according to the spatio-temporal data matrix and the mean vector;

[0140] The covariance matrix is a square matrix used to describe the covariance relationship between multiple random variables. If there are n random variables, the covariance matrix is an n×n matrix, which reflects the degree of linear correlation between variables, and the elements on the diagonal are the variances of each variable itself. In data analysis, the covariance matrix can help understand the relationship between different variables. For example, in multivariate time series analysis or multivariate statistical analysis, it is an important basis for operations such as data feature analysis and principal component analysis.

[0141] In this embodiment, obtain the already constructed spatio-temporal data matrix and the calculated mean vector, and then calculate the covariance matrix through the spatio-temporal data matrix and the mean vector. The covariance matrix can be calculated by the following formula:

[0142] ;

[0143] where, Represents the covariance matrix in the features and features the covariance between, represents the number of samples of spatio-temporal data, represents the th sample in the feature value, represents the th sample in the feature value, represents the th sample's mean vector, represents the th sample's mean vector.

[0144] This embodiment calculates the covariance matrix based on the spatio-temporal data matrix and the mean vector. As an important basis for the characteristics of data distribution, this covariance matrix can reflect the linear correlation between each feature, which helps with data dimensionality reduction and feature extraction.

[0145] S206. Perform eigenvalue decomposition on the covariance matrix to obtain eigenvalues and eigenvectors, and construct principal components based on the eigenvectors, where the eigenvalue is the variance explained by each principal component, and the eigenvector is the direction of the principal component;

[0146] Eigenvalue decomposition is a method of decomposing a matrix into a special form and has extensive applications in many fields such as data analysis, image processing, and physics. For a square matrix, it can find a set of eigenvalues and corresponding eigenvectors. For example, for square matrix A, there exists a non-zero vector x and a number λ such that Ax = λx, where λ is the eigenvalue and x is the eigenvector. Through eigenvalue decomposition, the square matrix A can be expressed as the product of the matrix composed of its eigenvectors, the diagonal matrix composed of eigenvalues, and the inverse of the eigenvector matrix.

[0147] In this embodiment, it is necessary to obtain the covariance matrix calculated in step S205. First, perform eigenvalue decomposition on the covariance matrix, usually using relevant algorithms in linear algebra for decomposition. After decomposition, eigenvalues and eigenvectors can be obtained, where the eigenvalue is the variance explained by each principal component, the eigenvector is the direction of the principal component, and there is a positive correlation between the eigenvalue and the eigenvector.

[0148] This embodiment can extract the principal components through eigenvalue decomposition, which can reduce the dimension of the data while retaining most of the data information. The eigenvalue represents the variance explained by each principal component, which can intuitively understand the contribution degree of each principal component to the overall variation of the data. The eigenvector represents the direction of the principal component, which helps to understand the structural relationship of the data in different dimensions.

[0149] S207. Calculate the cumulative variance ratio of each principal component through variance.

[0150] In this embodiment, obtain the eigenvalues corresponding to each principal component obtained in step S206, that is, the variance explained by each principal component, and calculate the cumulative variance ratio of each principal component through this variance. Here, the cumulative variance ratio of each principal component can be calculated through the following steps. Suppose there are principal components, and their corresponding eigenvalues are , respectively. For each principal component , starting from the first principal component, calculate sequentially according to the following formula to obtain the cumulative variance ratio of each principal component:

[0151] ;

[0152] Among them, represents the total variance, represents the sum of variances from the first principal component to the th principal component, represents the total number of principal components, represents the variance corresponding to the th principal component, represents the sum of the th variances, represents the cumulative variance ratio.

[0153] By calculating the cumulative variance ratio through variance in this embodiment, it can intuitively reflect the cumulative contribution degree of each principal component to the overall variance, which helps to determine how many principal components to retain to retain sufficient information while reducing the data dimension, providing a basis for subsequent analysis.

[0154] S208. Judge whether the cumulative variance ratio is greater than or equal to 90%.

[0155] In this embodiment, it is necessary to first obtain the cumulative variance ratio of each principal component calculated in step S207, and sequentially judge the obtained cumulative variance ratio with 90%, and select the principal components that explain most of the variance, that is, select the principal components with a cumulative variance ratio greater than or equal to 90%. This judgment process needs to traverse sequentially. Starting from the first principal component, as the number of principal components increases, continuously check whether the cumulative variance ratio under the current number of principal components is greater than or equal to 90%. For example, if the cumulative variance ratio of the first principal component is less than 90%, then continue to check the cumulative variance ratio of the next principal component, and so on, until the cumulative variance ratios of all principal components are checked. If it is determined that the cumulative variance ratio is greater than or equal to 90%, then execute step S209.

[0156] This embodiment can determine how many principal components can be used to represent the original data while retaining most of the data information by judging whether the cumulative variance ratio is greater than or equal to 90%. This helps to retain data information that has a significant impact on the analysis results to the greatest extent while reducing the data dimension, avoid losing key information due to excessive dimensionality reduction or increasing unnecessary computational complexity due to retaining too many dimensions, and provide an effective data simplification basis for subsequent analysis.

[0157] S209, determining as the target principal component;

[0158] In this embodiment, when it is determined that the cumulative variance ratio is greater than or equal to 90%, the corresponding principal component is determined as the target principal component. For example, in the previous process of calculating the cumulative variance ratio, when k principal components are calculated, their cumulative variance ratio is greater than or equal to 90%, then the k principal components are determined as the target principal components.

[0159] This embodiment determines the target principal component when the cumulative variance ratio is greater than or equal to 90%, which is conducive to accurately determining the key components of the data, simplifying the analysis process, reducing unnecessary calculations and data processing complexity, and improving the accuracy and reliability of subsequent analysis results.

[0160] S210, constructing an eigenvector matrix based on the eigenvectors of the target principal components;

[0161] In this embodiment, after the target principal component is determined, since each principal component has a corresponding eigenvector, the eigenvector corresponding to the target principal component is extracted, and the corresponding eigenvectors are arranged in a preset order to construct a eigenvector matrix. For example, the eigenvector is a two-dimensional vector, and the target principal component is The corresponding eigenvectors are , then we can construct a The eigenvector matrix of .

[0162] This embodiment constructs an eigenvector matrix based on the eigenvector of the target principal component to simplify the complex relationships in the original data. Each sample point in the original data can be represented in the principal component space by multiplying it with the eigenvector matrix. This step helps to explore the structural relationships within the data and can more clearly identify the key features in the target principal component.

[0163] S211, projecting the spatiotemporal data onto the eigenvector matrix to obtain a data matrix;

[0164] In this embodiment, it is necessary to first obtain the constructed feature vector matrix and the preprocessed spatio-temporal data. This spatio-temporal data contains a large number of features. The data matrix obtained by projecting onto the feature vector matrix has its dimension reduced from the original q-dimensional number of original features to the r-dimensional number of target principal components. During the projection process, the most critical information in the spatio-temporal data is retained, and some redundant information is removed. The result after projection redistributes the spatio-temporal data in a new space, making the relationships that were not easily noticeable in the original space become prominent in the new space.

[0165] This embodiment projects the spatio-temporal data onto the feature vector matrix to obtain a data matrix, achieving data dimensionality reduction and feature extraction, which is beneficial to highlighting the main structure and relationships of the data.

[0166] S212. Calculate the distance from each point in the data matrix to the preset point and the volume of a preset radius centered on the preset point according to the preset point in the data matrix, and calculate the local density based on the distance and volume;

[0167] S213. Calculate the distance between each point in the data matrix and the nearest target point, and select the target neighborhood radius based on the local density;

[0168] S214. Determine the data dimension based on the number of features, and set the minimum number of points through the data dimension;

[0169] S215. Mark each point in the data matrix through the density clustering algorithm according to the target domain radius and the minimum number of points to obtain a clustering result, and the clustering result contains multiple clustering clusters;

[0170] S216. Analyze the features of the clustering result based on multiple clustering clusters.

[0171] In this embodiment, steps S212 to S216 are similar to steps S104 to S108 in the foregoing embodiment, and will not be elaborated here.

[0172] The following will elaborate on the system for clustering and analyzing people based on spatio-temporal data provided by this application. Please refer to Figure 3 , Figure 3 This is another embodiment of the system for clustering and analyzing people based on spatio-temporal data provided by this application. The system includes:

[0173] An acquisition unit 301, configured to acquire spatio-temporal data, where the spatio-temporal data includes time data, space data, traffic data, environmental data, and socio-economic data;

[0174] A construction unit 302, configured to perform feature processing on the spatio-temporal data to obtain the number of features, and construct a spatio-temporal data matrix based on the number of features;

[0175] The dimensionality reduction unit 303 is used to reduce the dimensionality of the spatio-temporal data matrix through a preset analysis algorithm to obtain a data matrix;

[0176] The calculation unit 304 is used to calculate the volume of a preset radius centered on a preset point and the number of points within the radius according to the preset points in the data matrix, and calculate the local density based on the number of points within the radius and the volume;

[0177] The selection unit 305 is used to calculate the distance between each point in the data matrix and the nearest target point, and select the target neighborhood radius based on the distance and the local density;

[0178] The setting unit 306 is used to determine the data dimension based on the number of features, and set the minimum number of points through the data dimension;

[0179] The marking unit 307 is used to mark each point in the data matrix through a density clustering algorithm according to the target domain radius and the minimum number of points to obtain a clustering result, and the clustering result includes multiple clustering clusters;

[0180] The analysis unit 308 is used to analyze the characteristics of the clustering result based on multiple clustering clusters.

[0181] Optionally, the dimensionality reduction unit 303 is further used for:

[0182] Calculating a mean vector through the spatio-temporal data matrix;

[0183] Calculating a covariance matrix according to the spatio-temporal data matrix and the mean vector;

[0184] Performing eigenvalue decomposition on the covariance matrix to obtain eigenvalues and eigenvectors, and constructing principal components based on the eigenvectors, where the eigenvalues are the variances explained by each principal component, and the eigenvectors are the directions of the principal components;

[0185] Selecting target principal components according to the magnitudes of the variances, and constructing an eigenvector matrix based on the eigenvectors of the target principal components;

[0186] Projecting the spatio-temporal data onto the eigenvector matrix to obtain a data matrix.

[0187] Optionally, the covariance matrix is calculated by the following formula:

[0188] ;

[0189] Where represents the covariance matrix in the feature and the feature between, represents the number of samples of the spatio-temporal data, represents the th sample in the feature The value of represents the -th sample in the feature value, represents the mean vector of the -th sample, which represents the mean vector of the

[0190] Optionally, the dimensionality reduction unit 303 is further configured to:

[0191] calculate the cumulative variance ratio of each principal component through variance;

[0192] select the target principal component based on the magnitude of the cumulative variance ratio;

[0193] construct a feature vector matrix based on the eigenvectors of the target principal component.

[0194] Optionally, the dimensionality reduction unit 303 is further configured to:

[0195] judge whether the cumulative variance ratio is greater than or equal to 90%;

[0196] If so, determine it as the target principal component.

[0197] Optionally, the marking unit 307 is further configured to:

[0198] Step 1: Determine the data matrix as the sample set D = (x1, x2,..., xm), the neighborhood parameters (ϵ, MinPts), and the sample distance metric, where ϵ represents the target neighborhood radius, MinPts represents the minimum number of points, and the sample set is a set composed of multiple TF-IDF vectors;

[0199] Step 2: Calculate the ϵ-neighborhood subsample set Nϵ(xj) of each sample Xj based on the sample distance metric;

[0200] Step 3: Compare the absolute value |Nϵ(xj)| of the ϵ-neighborhood subsample set Nϵ(xj) with MinPts, and add the sample xj greater than MinPts to the core object sample set Ω;

[0201] Step 4: When the core object sample set Ω is not empty, randomly select a core object o from the core object sample set Ω and execute the following algorithm:

[0202] Initialize the current cluster core object queue Ωcur = {o};

[0203] Initialize the class serial number k = k + 1;

[0204] Initialize the current cluster sample set Ck = {o};

[0205] Update the unvisited sample set Γ = Γ - {o};

[0206] Step Five: If the current cluster core object queue Ωcur is empty, the current clustering cluster Ck is generated; after generating the clustering cluster Ck, update the cluster partition C = C ∪ {Ck}, and update the core object sample set Ω = Ω - Ck;

[0207] Step Six: If the current cluster core object queue Ωcur is not empty, then execute the following algorithm:

[0208] Take out a core object o' from the current cluster core object queue Ωcur;

[0209] Determine all ϵ-neighborhood subsample sets Nϵ(o') through the neighborhood distance threshold ϵ;

[0210] Let Δ = Nϵ(o') ∩ Γ;

[0211] Update the current cluster sample set Ck = Ck ∪ Δ, and update the unvisited sample set Γ = Γ - Δ;

[0212] Update Ωcur = Ωcur ∪ (Δ ∩ Ω) - {o'};

[0213] Repeat Step Five;

[0214] Step Seven: Output the cluster partition C = {C1, C2,..., Ck}, and obtain the clustering result, where the clustering result contains multiple clustering clusters.

[0215] Optionally, it further includes a processing unit 309, which is specifically used for:

[0216] Perform data preprocessing on the spatio-temporal data, and the data preprocessing includes: data deduplication, missing value processing, and unifying data types.

[0217] Optionally, the local density is calculated by the following formula:

[0218] ;

[0219] Among them, represents the preset point the preset radius of the number of points within represents the volume of the preset radius centered on the preset point , represents the local density.

[0220] This application also provides a device for clustering and analyzing people based on spatio-temporal data. Please refer to Figure 4 , Figure 4An embodiment of the apparatus for clustering analysis of people based on spatio-temporal data provided by this application, the apparatus includes:

[0221] A processor 401, a memory 402, an input / output unit 403, and a bus 404;

[0222] The processor 401 is connected to the memory 402, the input / output unit 403, and the bus 404;

[0223] The memory 402 stores a program, and the processor 401 calls the program to execute any of the above methods.

[0224] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described systems, apparatuses, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0225] In several embodiments provided by this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of the apparatuses or units can be in electrical, mechanical, or other forms.

[0226] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0227] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0228] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, read-only memory), random access memories (RAM, random access memory), magnetic disks, or optical discs.

Claims

1. A method for clustering analysis of a population based on spatiotemporal data, characterized in that: The method comprises: Acquiring spatiotemporal data, the spatiotemporal data including time data, space data, traffic data, environmental data, and socio-economic data; Performing feature processing on the spatiotemporal data to obtain a feature quantity, and constructing a spatiotemporal data matrix based on the feature quantity; The spatial-temporal data matrix is ​​reduced in dimension by a preset analysis algorithm to obtain a data matrix; According to the preset point in the data matrix, the volume of a preset radius centered at the preset point and the number of points within the radius are calculated, and local density is calculated based on the number of points within the radius and the volume; Calculating the distance between each point in the data matrix and the nearest target point, and selecting a target neighborhood radius based on the distance and the local density; Determining a data dimension based on the number of features, and setting a minimum number of points by the data dimension; According to the target neighborhood radius and the minimum number of points, each point in the data matrix is ​​marked by a density clustering algorithm to obtain a clustering result, wherein the target neighborhood radius is used to define the radius of the neighborhood range in the density clustering algorithm, and the clustering result includes multiple clusters; The characteristics of the clustering result are analyzed based on the multiple clustering clusters.

2. The method according to claim 1, characterized in that The step of reducing the dimension of the spatiotemporal data matrix by a preset analysis algorithm to obtain a data matrix comprises: Obtaining a mean vector by calculating the spatiotemporal data matrix; Calculate a covariance matrix based on the spatiotemporal data matrix and the mean vector; Performing eigenvalue decomposition on the covariance matrix to obtain eigenvalues ​​and eigenvectors, and constructing principal components based on the eigenvectors, wherein the eigenvalues ​​are the variances explained by each principal component, and the eigenvectors are the directions of the principal components; Selecting a target principal component according to the size of the variance, and constructing an eigenvector matrix based on the eigenvector of the target principal component; The spatiotemporal data is projected onto the eigenvector matrix to obtain a data matrix.

3. The method according to claim 2, characterized in that The covariance matrix is ​​calculated as follows: ; in, Denotes the covariance matrix Medium Features and Features The covariance between represents the number of samples of the spatiotemporal data, Indicates Samples in the feature The value of Indicates Samples in the feature The value of Indicates The mean vector of samples, Indicates The mean vector of samples.

4. The method according to claim 2, characterized in that: The step of selecting a target principal component according to the size of the variance and constructing a eigenvector matrix based on the eigenvector of the target principal component includes: Calculating the cumulative variance ratio of each principal component through the variance; Selecting a target principal component based on the size of the cumulative variance ratio; An eigenvector matrix is ​​constructed based on the eigenvectors of the target principal components.

5. The method according to claim 4, characterized in that The selecting of the target principal component based on the size of the cumulative variance ratio includes: Determine whether the cumulative variance ratio is greater than or equal to 90%; If so, it is determined to be the target principal component.

6. The method according to claim 1, characterized in that According to the target neighborhood radius and the minimum number of points, each point in the data matrix is ​​marked by a density clustering algorithm to obtain a clustering result, wherein the target neighborhood radius is used to define the radius of the neighborhood range in the density clustering algorithm, and the clustering result includes multiple clusters including: Step 1: Determine that the data matrix is ​​a sample set D=(x1,x2,...,xm), a neighborhood parameter (ϵ,MinPts) and a sample distance metric, where ϵ represents the target neighborhood radius, MinPts represents the minimum number of points, and the sample set is a set consisting of multiple TF-IDF vectors; Step 2: Based on the sample distance measurement method, calculate the ϵ-neighborhood subsample set Nϵ(xj) of each sample Xj; Step 3: Compare the absolute value |Nϵ(xj)| of the ϵ-neighborhood subsample set Nϵ(xj) with the MinPts, and add the sample xj greater than the MinPts to the core object sample set Ω; Step 4: When the core object sample set Ω is not empty, randomly select a core object o from the core object sample set Ω and execute the following algorithm: Initialize the current cluster core object queue Ωcur={o}; Initialize category number k=k+1; Initialize the current cluster sample set Ck={o}; Update the unvisited sample set Γ=Γ-{o}; Step 5: If the current cluster core object queue Ωcur is empty, the current cluster Ck is generated; after the cluster Ck is generated, the cluster partition C=C∪{Ck} is updated, and the core object sample set Ω=Ω-Ck is updated; Step 6: If the current cluster core object queue Ωcur is not empty, execute the following algorithm: Take out a core object o' from the current cluster core object queue Ωcur; Determine all ϵ-neighborhood subsample sets Nϵ(o') through the neighborhood distance threshold ϵ; Let Δ = Nϵ(o')∩Γ; Update the current cluster sample set Ck=Ck∪Δ, and update the unvisited sample set Γ=Γ-Δ; Update Ωcur = Ωcur ∪ (Δ ∩ Ω) - {o'}; Repeat step 5; Step 7: Output cluster partition C={C1, C2, ..., Ck} to obtain a clustering result, which includes multiple clusters.

7. The method according to claim 1, characterized in that After acquiring the spatiotemporal data, the spatiotemporal data including time data, space data, traffic data, environmental data and socio-economic data, and before processing the spatiotemporal data to obtain feature quantities and constructing a spatiotemporal data matrix based on the feature quantities, the method further includes: The spatiotemporal data is preprocessed, and the data preprocessing includes: data deduplication, missing value processing, and data type unification.

8. The method according to any one of claims 1 to 6, wherein the local density is calculated by the following formula: ; in, Indicates the preset point The preset radius The points within Indicates that the preset point The preset radius of the center The volume of represents the local density.

9. A system for clustering analysis of populations based on spatiotemporal data, characterized in that: include: An acquisition unit, used to acquire spatiotemporal data, wherein the spatiotemporal data includes time data, space data, traffic data, environmental data, and socio-economic data; A construction unit, used for performing feature processing on the spatiotemporal data to obtain feature quantities, and constructing a spatiotemporal data matrix based on the feature quantities; A dimension reduction unit, used for reducing the dimension of the spatiotemporal data matrix to obtain a data matrix by using a preset analysis algorithm; a calculation unit, configured to calculate, according to a preset point in the data matrix, a volume of a preset radius centered on the preset point and the number of points within the radius, and calculate a local density based on the number of points within the radius and the volume; a selection unit, configured to calculate the distance between each point in the data matrix and the nearest target point, and select a target neighborhood radius based on the distance and the local density; A setting unit, configured to determine a data dimension based on the number of features, and set a minimum number of points through the data dimension; A marking unit, used for marking each point in the data matrix according to the target neighborhood radius and the minimum number of points by a density clustering algorithm to obtain a clustering result, wherein the target neighborhood radius is used to define the radius of the neighborhood range in the density clustering algorithm, and the clustering result includes a plurality of clusters; An analyzing unit is used to analyze the characteristics of the clustering result based on the multiple clustering clusters.

10. A device for clustering analysis of a population based on spatiotemporal data, characterized in that: The device comprises: Processor, memory, input-output unit, and bus; The processor is connected to the memory, the input and output unit, and the bus; The memory stores a program, and the processor calls the program to execute the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Multi-dimensional data processing analysis method, device and system and storage medium

    CN118211038A

  • Artificial intelligence data aggregation method based on big data

    CN118520020A