A data distributed clustering method and device based on local direction centrality

By constructing the preferred search K-means tree global index in the clustering algorithm and using data sampling and Hilbert curve partitioning methods, the problems of low computational efficiency and poor partitioning effect of the clustering algorithm are solved, and efficient distributed computing and good clustering accuracy are achieved.

CN115658809BActive Publication Date: 2025-06-24SHANGHAI YUNTU DIGITAL MIRROR SOFTWARE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211265216.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-17
Publication Date
2025-06-24
Estimated Expiration
2042-10-17

AI Technical Summary

Technical Problem

In the prior art, the computing efficiency of the clustering algorithm is not high and the partitioning effect is poor, making it difficult to achieve efficient distributed computing when processing large-scale data.

Method used

The distributed clustering method of data based on local directional centering is adopted, and the process of nearest neighbor search and data partitioning is optimized by constructing a preferred search K-means tree global index and combining data sampling and Hilbert curve partitioning methods, thereby improving the parallel computing efficiency of the algorithm.

Benefits of technology

While maintaining clustering accuracy, the parallel acceleration effect and scalability of the algorithm are significantly improved, and the computability of massive data and the utilization of distributed system computing resources are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115658809B_ABST
    Figure CN115658809B_ABST
Patent Text Reader

Abstract

The present invention discloses a data distributed clustering method and device based on local direction centrality. The method includes the following steps: S1. Submit the parameters required for the algorithm task in a distributed cluster environment and read the data to be clustered; S2. Construct a priority search K-means tree global index based on the complete data and share the index variables to each working node in the cluster; S3. Divide the complete data by combining data sampling and Hilbert curve partitioning method; S4. Parallelly execute CDC local clustering on each working node; S5. Perform inter-region cluster merging according to the maximum reachable distance of the local clusters to generate complete clusters; S6. Output the clustering result to the distributed file system. The method of the present invention performs distributed optimization and acceleration on the CDC clustering algorithm from two perspectives of algorithm process optimization and parallel processing optimization, aiming to improve the computing efficiency of the CDC algorithm and provide a feasible optimization scheme for the application of this algorithm in massive data mining and machine learning tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data mining, and in particular to a data distributed clustering method and device based on local direction centrality. Background Art

[0002] In recent years, a large number of clustering algorithm studies have proposed effective solutions for problems such as arbitrary-shaped cluster recognition, outlier detection, and high-dimensional data processing. However, the density heterogeneity and weak connectivity of data distribution are still common and unresolved problems in the application scenarios of clustering analysis. Since the internal points of a cluster tend to be surrounded by their neighbor points in all directions, while the boundary points only have neighbor points within a certain direction range, the internal points and boundary points can be divided according to this difference in the distribution of neighboring directions. Accordingly, the local direction centrality clustering algorithm CDC measures the direction uniformity of the K-nearest neighbor (KNN) distribution of data points by establishing a local direction centrality metric (DCM), and realizes the division of internal points and boundary points in a density-independent manner; at the same time, using boundary points to constrain the connection of internal points can avoid cross-cluster connectivity and realize the separation of weakly connected clusters, providing an effective solution to the above problems. The accuracy of this algorithm has been verified on artificial and real datasets. However, the nearest neighbor search has a time complexity of O(n 2 ), and as the data scale increases, the computing efficiency decreases significantly, and even the situation where a single machine cannot calculate occurs, making it impossible to cope with the exponentially growing data scale. To address the above problems, in addition to improving the process to reduce the time complexity of the algorithm itself, the distributed computing efficiency of the clustering algorithm can also be improved from the perspective of parallel computing.

[0003] Parallelization has become a hot topic in the performance optimization of current clustering algorithms. Commonly used distributed computing frameworks include Hadoop, Spark, and Flink, etc. Among them, Spark is a new generation of big data parallel processing platform, which has the advantages of being simple to use, rich in functions, and automatically fault-tolerant. Compared with the classic big data parallel processing platform Hadoop, Spark's memory-based data management makes it more suitable for clustering algorithms that require multiple rounds of iteration. Existing research has proposed many parallelization schemes for Spark-based clustering algorithms. These studies have improved the efficiency of big data clustering to a certain extent by designing the algorithm into three stages: data partitioning, distributed local clustering, and global merging; however, Spark's default partitioning strategy easily leads to unbalanced data load in partitions due to ignoring the spatial proximity of clusters. When the data in a partition is skewed, it will cause an unbalanced workload of nodes in the Shuffle stage, that is, the amount of data processed by each node in the cluster varies greatly and the execution time is inconsistent, thus reducing the utilization rate of cluster resources and the computing efficiency of distributed algorithms.

[0004] It can be seen from this that there are technical problems of low computing efficiency and poor partitioning effect in the prior art. Summary of the Invention

[0005] The present invention provides a data distributed clustering method and device based on local direction centrality clustering, which is used to solve or at least partially solve the technical problems of low computing efficiency and poor partitioning effect in the prior art.

[0006] To solve the above technical problems, the first aspect of the present invention provides a data distributed clustering method based on local direction centrality, including:

[0007] S1: Receive the parameters required for the clustering task, including environment parameters, clustering algorithm parameters, partitioning parameters, and nearest neighbor search parameters, configure and register a serializer, and read the complete data to be clustered from the distributed file system;

[0008] S2: Construct a priority search K-means tree global index based on the read complete data to be clustered, and share the global index to each worker node through the master node of the distributed cluster;

[0009] S3: Partition the complete data to be clustered by combining data sampling and Hilbert curve partitioning method, and obtain the corresponding partition ID. Send the partition data corresponding to the partition ID to the corresponding worker node through the master node of the distributed cluster;

[0010] S4: Each worker node of the distributed cluster executes the CDC local clustering algorithm in parallel, specifically including: the worker node performs k-nearest neighbor search on the partition data respectively through the shared global index and calculates the DCM value, and divides the internal points and boundary points according to the relationship between the DCM value and the DCM threshold, and then merges the internal points based on the reachable distance from the internal points to the boundary points. The merged internal points are grouped into the same internal point cluster, mark the internal point cluster ID, search for the nearest internal point to the boundary point and mark the boundary point cluster ID to obtain the local cluster, where the DCM value is the angular variance formed by the data point and its k neighboring points in the two-dimensional space;

[0011] S5: The master node of the distributed cluster merges the local clusters between partitions according to the maximum reachable distance of the local clusters to generate a complete cluster as the clustering result;

[0012] S6: Output the clustering result to the distributed file system.

[0013] In one implementation, step S1 includes:

[0014] S1.1: The distributed cluster receives the parameters required for the clustering task. Among them, the environmental parameters include the file path, the clustering algorithm parameters include the number of neighbors and the proportion of boundary points, the partitioning parameters include the partitioning type, the partitioning sampling rate ratio, and the number of partitions, and the nearest neighbor search parameters include the index type parameter, the construction parameter, and the search parameter;

[0015] S1.2: Register the serializer for the geometric type object and the index;

[0016] S1.3: Read the complete data to be clustered in the distributed file system according to the file path and perform projection conversion.

[0017] In one implementation, step S2 includes:

[0018] S2.1: Initialize the index structure according to the index type parameter and the construction parameter. Among them, the construction parameter includes the branching factor branch, the maximum number of iterations I of K-means max and the initial centroid selection method C alg ;

[0019] S2.2: Calculate the centroid of the complete data to be clustered and construct the root node of the index tree;

[0020] S2.3: Select branch initial partition centroids according to C alg and divide the data into the nearest partitions;

[0021] S2.4: Update the partition centroids and re-divide the data until the partition centroids remain unchanged or the update reaches I max ;

[0022] S2.5: Construct nodes according to the partition centroids and add them to the child node set of the parent node;

[0023] S2.6: Repeat steps S2.3 to S2.5 until the number of data in the partition is less than branch to obtain the constructed global index of the priority search K-means tree, and represent it with a variable;

[0024] S2.7: Distribute the global index variable of the priority search K-means tree to each worker node through the master node of the distributed cluster.

[0025] In one implementation, step S3 includes:

[0026] S3.1: Sample the complete data to be clustered according to the partitioning sampling rate ratio;

[0027] S3.2: Calculate the Hilbert coding values of the sampled data and sort the sampled points according to the value size;

[0028] S3.3: Uniformly divide the sampling points into a number of intervals corresponding to the number of partitions, record the division positions as partitions.

[0029] S3.4: Expand the sampling points to form a rectangular partition range and generate partition IDs for all data.

[0030] S3.5: Distribute the corresponding partition data to each working node in the cluster according to the partition IDs.

[0031] In one implementation, step S4 includes:

[0032] S4.1: Perform k-nearest neighbor search on the data to be clustered in each partition based on the global index structure of the priority search K-means tree to obtain k neighbor points of the data points.

[0033] S4.2: Calculate the angular variance DCM formed by each data point and its k neighbor points in the two-dimensional space.

[0034]

[0035] where k is the number of neighbor points of the data point, (θ1, θ2, …, θ k ) are the angles formed by the data point and its neighbor points, and satisfy

[0036] S4.3: Merge the DCM value results of all data points in the partition, sort them, and then calculate the threshold T according to the boundary point ratio parameter Ratio DCM , if DCM is less than T DCM then mark the data point as an internal point, otherwise mark it as a boundary point;

[0037] S4.4: Calculate the minimum distance between the internal point p i and all boundary points q m as the reachable distance r i of the internal point p i , that is, r i = min(d(p i , q m )); where p i is the internal point, q m is the boundary point, and d(p i , q m ) is the distance between the internal point p i and the boundary point q m ;

[0038] S4.5: Merge the internal points and mark the internal point cluster IDs according to the connection rule that the distance between two internal points is not greater than the sum of the reachable distances of the two points, that is, d(p i , p j ) ≤ ri +r j ; where r i , r j are the reachable distances of the internal points p i , p j respectively, and d(p i , p j ) is the distance between the internal points p i and p j ;

[0039] S4.6: Search for the internal point closest to the boundary point, and use the cluster ID where the searched internal point is located as the boundary point cluster ID to obtain the local cluster.

[0040] In one implementation, step S5 includes:

[0041] Sort the reachable distances of each internal point in the local cluster C α to obtain the maximum reachable distance R α = max(r i );

[0042] Perform interval-based cluster merging according to the connection rule that the distance between two clusters is not greater than the sum of the reachable distances of the two clusters, that is, D(C a , C β ) ≤ R α + R β , and update the cluster ID to generate a complete cluster, where C a , C β are two different local clusters, R α , R β are the maximum reachable distances of the local clusters C a , C β respectively, and D(C a , C β ) is the distance between the internal points with the maximum reachable distance in C a , C β .

[0043] In one implementation, after step S6, the method further includes: Outputting the clustering evaluation result and the calculation time result to the distributed file system.

[0044] Based on the same inventive concept, the second aspect of the present invention provides a data distributed clustering device based on local direction centrality, including:

[0045] An initialization module, configured to receive the parameters required for the clustering task, including environment parameters, clustering algorithm parameters, partitioning parameters, and nearest neighbor search parameters, configure and register a serializer, and read the complete data to be clustered from the distributed file system;

[0046] A global index construction module, which is used to construct a priority search K-means tree global index based on the read complete data to be clustered, and share the global index to each worker node through the master node of the distributed cluster;

[0047] A data partitioning module, which is used to partition the complete data to be clustered by combining data sampling and Hilbert curve partitioning method, and obtain the corresponding partition ID, and send the partition data corresponding to the partition ID to the corresponding worker node through the master node of the distributed cluster;

[0048] A local clustering module, which is used to parallelly execute the CDC local clustering algorithm through each worker node of the distributed cluster. Specifically, it includes: performing k-nearest neighbor search on the partition data respectively by using the global index shared by the master node and calculating the DCM value, and dividing the internal points and boundary points according to the relationship between the DCM value and the DCM threshold, then merging the internal points based on the reachable distance from the internal points to the boundary points, and the merged internal points are classified into the same internal point cluster, marking the internal point cluster ID, searching for the nearest internal point to the boundary point and marking the boundary point cluster ID to obtain local clusters, where the DCM value is the angular variance formed by the data point and its k neighboring points in the two-dimensional space;

[0049] A global merging module, which is used to merge the local clusters between partitions according to the maximum reachable distance of the local clusters through the master node of the distributed cluster to generate complete clusters as the clustering results;

[0050] A result output module, which is used to output the clustering results to the distributed file system.

[0051] Based on the same inventive concept, the third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed, the method described in the first aspect is implemented.

[0052] Based on the same inventive concept, the fourth aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in the first aspect is implemented.

[0053] Compared with the prior art, the advantages and beneficial technical effects of the present invention are as follows:

[0054] The present invention proposes a data distributed clustering method based on local directional centrality, which realizes data clustering in a distributed cluster environment by using a local directional centrality clustering algorithm. This method improves the efficiency of nearest neighbor search in local clustering based on a priority search K-means tree index, and optimizes the data partitioning in the parallel processing flow based on data sampling and Hilbert curve partitioning method. Among them, the priority search K-means tree index converts the distance calculation of point pairs into the neighborhood query of nodes to accelerate the search for nearest neighbor points, which can narrow the query range to improve the efficiency of local clustering; in addition, the partitioning method combining data sampling and Hilbert curve takes into account both the partitioning efficiency and the spatial proximity of data distribution. It not only reduces the data volume of partitioning calculation through data sampling to improve the partitioning speed, but also constructs an equilibrium partition by combining the Hilbert curve with good spatial aggregation characteristics to improve the performance of the parallel algorithm. The method of the present invention can be applied to massive two-dimensional point data, and has good parallel acceleration effect and scalability on the premise of maintaining clustering accuracy, so as to improve the computability of the algorithm facing massive data and the utilization rate of computing resources in the distributed system, and provide a feasible optimization scheme for various big data mining and machine learning applications of the CDC clustering algorithm. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0056] Figure 1 It is a schematic diagram of an artificial data set provided by an embodiment of the present invention;

[0057] Figure 2 It is a comparison chart of the clustering accuracy between the clustering method of the present invention and the existing serial clustering method;

[0058] Figure 3 It is a comparison chart of the execution time between the clustering method of the present invention and the existing serial clustering method;

[0059] Figure 4 It is a flowchart of a distributed local directional centrality clustering algorithm based on Spark in an embodiment of the present invention;

[0060] Figure 5 It is a schematic diagram of an algorithm for constructing a priority search K-means tree in an embodiment of the present invention;

[0061] Figure 6 It is a flowchart of a data partitioning algorithm in an embodiment of the present invention;

[0062] Figure 7 It is a schematic diagram of the local clustering algorithm in an embodiment of the present invention;

[0063] Figure 8 It is a schematic diagram of the query-priority search K-means tree algorithm in an embodiment of the present invention;

[0064] Figure 9 It is a schematic diagram of the partition global merging algorithm in an embodiment of the present invention;

[0065] Figure 10 It is a flowchart of the data distributed clustering method based on local direction centrality provided by an embodiment of the present invention;

[0066] Figure 11 It is a framework diagram of the data distributed clustering device based on local direction centrality provided by an embodiment of the present invention;

[0067] Figure 12 It is a schematic diagram of the structure of the computer-readable storage medium provided by an embodiment of the present invention;

[0068] Figure 13 It is a schematic diagram of the structure of the computer device provided by an embodiment of the present invention;

[0069] Figure 14 It is a schematic diagram of the computer distributed architecture provided by an embodiment of the present invention. Detailed implementation manners

[0070] The technical problem to be solved by the present invention is how to accelerate the nearest neighbor search in the algorithm process to achieve a more efficient data point division and design a fast and balanced data partitioning method to improve the overall performance of the parallel distributed system and the efficiency of the CDC clustering algorithm in view of the low computing efficiency of the CDC clustering algorithm when processing massive data and the data skew problem of the Spark default partitioning strategy.

[0071] Specifically, the method of the present invention takes the CDC clustering algorithm as the research object and the distributed computing framework as the technical support. In view of the computational intensity faced by the CDC clustering algorithm in the massive data scenario and the data skew problem of the Spark default partitioning, a priority search K-means tree index is constructed to accelerate the nearest neighbor search and reduce the computational complexity of the algorithm; and the space partitioning is optimized by combining data sampling and the Hilbert curve partitioning method to solve problems such as low partitioning efficiency, partition data skew, and high cross-node communication cost. On this basis, a two-dimensional parallel CDC clustering algorithm based on Spark is designed and implemented to meet the requirements of the CDC clustering algorithm for processing massive data application scenarios.

[0072] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0073] Embodiment 1

[0074] An embodiment of the present invention provides a data distributed clustering method based on local direction centrality, including:

[0075] S1: Receive the parameters required for the clustering task, including environment parameters, clustering algorithm parameters, partitioning parameters, and nearest neighbor search parameters, configure and register a serializer, and read the complete data to be clustered from the distributed file system;

[0076] S2: Construct a priority search K-means tree global index based on the read complete data to be clustered, and share the global index to each worker node through the master node of the distributed cluster;

[0077] S3: Partition the complete data to be clustered by combining data sampling and Hilbert curve partitioning method, and obtain the corresponding partition ID. Send the partition data corresponding to the partition ID to the corresponding worker node through the master node of the distributed cluster;

[0078] S4: Each worker node of the distributed cluster executes the CDC local clustering algorithm in parallel, specifically including: the worker node performs k-nearest neighbor search on the partition data respectively through the shared global index and calculates the DCM value, and divides the internal points and boundary points according to the relationship between the DCM value and the DCM threshold, and then merges the internal points based on the reachable distance from the internal points to the boundary points. The merged internal points are grouped into the same internal point cluster, mark the internal point cluster ID, search for the nearest internal point to the boundary point and mark the boundary point cluster ID to obtain the local cluster, where the DCM value is the angular variance formed by the data point and its k neighboring points in the two-dimensional space;

[0079] S5: The master node of the distributed cluster merges the local clusters between partitions according to the maximum reachable distance of the local clusters to generate a complete cluster as the clustering result;

[0080] S6: Output the clustering result to the distributed file system.

[0081] Specifically, by constructing a priority search K-means tree global index in step S2, the nearest neighbor search step in the CDC local clustering algorithm in step S4 can be accelerated, thereby improving the calculation efficiency.

[0082] Step S3 is the data partitioning phase. First, random sampling is used to reduce the amount of data for partitioning calculation. Then, the data is evenly partitioned by calculating the Hilbert coding values of the sampling points. Finally, all the data is partitioned according to the sampling point partitions. After partitioning, the partitioned data is distributed to each working node in the cluster.

[0083] Step S4 is the local clustering phase. CDC includes three steps: nearest neighbor search, data partitioning, and internal point connection.

[0084] Step S5 is the global merging phase. According to the maximum reachable distance of the local clusters, the local clusters in each partition are merged, and the cluster IDs are updated simultaneously to generate complete clusters.

[0085] Step S6 is the result output. The clustering result is output to the distributed file system.

[0086] This solution is a distributed optimization and acceleration solution proposed for the high time complexity of the CDC clustering algorithm and its difficulty in processing large-scale data. The solution aims to achieve good parallel acceleration effects and scalability while maintaining the clustering accuracy of the algorithm, with the expectation of improving the computability of the algorithm for massive data and the utilization rate of computing resources in the distributed system, that is, the computability problem of the CDC clustering algorithm in the massive data scenario and the data skew problem of the native data partitioning scheme of Apache Spark.

[0087] In one implementation, step S1 includes:

[0088] S1.1: The distributed cluster receives the parameters required for the clustering task. Among them, the environment parameters include the file path, the clustering algorithm parameters include the number of neighbors and the boundary point ratio, the partitioning parameters include the partitioning type, the partitioning sampling rate ratio, and the number of partitions, and the nearest neighbor search parameters include the index type parameter, the construction parameter, and the search parameter;

[0089] S1.2: Register the serializer for the geometric type object and the index;

[0090] S1.3: Read the complete data to be clustered in the distributed file system according to the file path and perform projection transformation.

[0091] In one implementation, step S2 includes:

[0092] S2.1: Initialize the index structure according to the index type parameter and the construction parameter. Among them, the construction parameter includes the branching factor branch, the maximum number of iterations I of K-means max and the initial centroid selection method C alg ;

[0093] S2.2: Calculate the centroid of the complete data to be clustered and construct the root node of the index tree;

[0094] S2.3: According to C alg Select branch initial partition centroids and partition the data into the nearest partitions;

[0095] S2.4: Update the partition centroids and repartition the data until the partition centroids remain unchanged or the update reaches I max ;

[0096] S2.5: Construct nodes according to the partition centroids and add them to the set of child nodes of the parent node;

[0097] S2.6: Repeat steps S2.3 to S2.5 until the number of data in the partition is less than branch, obtain the constructed global index of the priority search K-means tree, and represent it with a variable;

[0098] S2.7: Distribute the global index variable of the priority search K-means tree to each worker node through the master node of the distributed cluster.

[0099] In one implementation, step S3 includes:

[0100] S3.1: Sample the complete data to be clustered according to the partition sampling rate ratio;

[0101] S3.2: Calculate the Hilbert coding values of the sampled data and sort the sampled points according to the value size;

[0102] S3.3: Uniformly divide the sampled points into the corresponding number of intervals as the number of partitions, record the division positions as partitions;

[0103] S3.4: Expand the sampled points to form a rectangular partition range and generate the partition IDs of all the data;

[0104] S3.5: Distribute the corresponding partition data to each worker node of the cluster according to the partition IDs.

[0105] In one implementation, step S4 includes:

[0106] S4.1: Perform k-nearest neighbor search on the data to be clustered in each partition based on the global index structure of the priority search K-means tree to obtain k neighbor points of the data points;

[0107] S4.2: Calculate the angular variance DCM formed by each data point and its k neighbor points in the two-dimensional space;

[0108]

[0109] Among them, k is the number of neighbor points of the data point, and (θ1, θ2, …, θ k ) are the angles formed by the data point and its neighbor points, and satisfy

[0110] S4.3: Combine the DCM value results of all data points in the partition and sort them, and then calculate the threshold T according to the boundary point ratio parameter Ratio DCM . If the DCM is less than T DCM , then mark the data point as an internal point, otherwise mark it as a boundary point;

[0111] S4.4: Calculate the minimum distance between the internal point p i and all boundary points q m as the reachable distance r i of the internal point p i , that is, r i = min(d(p i , q m )); where p i is an internal point, q m is a boundary point, and d(p i , q m ) is the distance between the internal point p i and the boundary point q m ;

[0112] S4.5: Merge the internal points according to the connection rule that the distance between two internal points is not greater than the sum of the reachable distances of the two points and mark the internal point cluster ID, that is, d(p i , p j ) ≤ r i + r j ; where r i , r j are the reachable distances of the internal points p i , p j respectively, and d(p i , p j ) is the distance between the internal point p i and the internal point p j ;

[0113] S4.6: Search for the internal point closest to the boundary point, and use the cluster ID of the internal point where the search result is located as the boundary point cluster ID to obtain the local cluster.

[0114] Specifically, the reachable distance is the minimum distance between the internal point and all boundary points, or the distance from the internal point to its closest boundary point. The internal points p i , p j are different internal points.

[0115] In one implementation, step S5 includes:

[0116] Sort the reachable distances of each interior point in the local cluster C α to obtain the maximum reachable distance R α = max(r i );

[0117] Perform interval-based cluster merging according to the maximum reachable distance and the connection rule that the distance between two clusters is not greater than the sum of the reachable distances of the two clusters, that is, D(C a , C β ) ≤ R α + R β , and update the cluster ID to generate a complete cluster, where C a , C β are two different local clusters, R α , R β are the maximum reachable distances of the local clusters C a , C β respectively, and D(C a , C β ) is the distance between the interior points with the maximum reachable distance in C a , C β .

[0118] Among them, the reachable distance between two clusters refers to the reachable distance of each interior point in the local cluster, and the maximum reachable distance is used to measure the boundary that the local cluster can reach, with the same principle as that of the interior point.

[0119] In one implementation, after step S6, the method further includes: outputting the clustering evaluation result and the calculation time result to a distributed file system.

[0120] Please refer to Figure 10 , for the detailed flowchart of the clustering method provided by the embodiments of the present invention.

[0121] The present invention discloses a distributed optimization and acceleration method for a local direction centrality clustering algorithm (Clustering by Local Direction Centrality, CDC) based on priority search K-means tree nearest neighbor search and data sampling and Hilbert curve data partitioning, including: S1. Submitting the required parameters of the algorithm task in a distributed cluster environment, configuring and registering a serializer, and reading the data to be clustered from a distributed file system; S2. Constructing a global index of a priority search K-means tree based on the complete data and sharing the index variable to each working node in the cluster; S3. Dividing the complete data by combining data sampling and Hilbert curve partitioning method and distributing the partitioned data to each working node in the cluster; S4. Parallelly executing CDC local clustering on each working node, including three steps of nearest neighbor search, data partitioning, and internal point connection, that is, first performing K-nearest neighbor search on the data to be clustered in each partition based on the global index and calculating the local direction centrality metric value (Direction Centrality Metric, DCM) for each data point, then dividing the internal points and boundary points according to the set DCM threshold parameter, and finally merging other internal points based on the reachable distance from the internal points to the boundary points and marking the internal point cluster ID, searching for the nearest internal point of the boundary point and marking the boundary point cluster ID to obtain local clusters; S5. Merging the clusters between partitions according to the maximum reachable distance of the local clusters, and updating the cluster ID at the same time to generate complete clusters; S6. Outputting the clustering result to the distributed file system. Aiming at the computability problem faced by the CDC clustering algorithm in the scenario of massive data and the data skew problem of the native data partitioning scheme of Apache Spark, the method of the present invention performs distributed optimization and acceleration on the CDC clustering algorithm from two perspectives of algorithm process optimization and parallel processing optimization. In terms of algorithm process optimization, the nearest neighbor search is optimized by constructing a priority search K-means tree index, reducing the computational complexity of the CDC clustering algorithm; in terms of parallel processing optimization, a spatial partition is established for the sampled data by using the Hilbert curve, solving the problems of low efficiency in constructing the default partition based on the complete data, data skew, and high communication cost of cross-node data transmission. The method of the present invention aims to improve the computational efficiency of the CDC algorithm and provide a feasible optimization scheme for the application of this algorithm in massive data mining and machine learning tasks.

[0122] To more clearly illustrate the beneficial effects of the technical solutions disclosed by the present invention, 6 two-dimensional artificial data sets with different data distributions are selected in the specific embodiments (such as Figure 1Experiments were carried out as shown in Figure 2 to calculate the clustering evaluation indexes of the serial and parallel optimization algorithms to verify the clustering accuracy of the present invention. The range of the experimental algorithm parameter k was set to 5 - 50, with a value interval of 5; the range of Ratio was set to 0.05 - 0.4, with a value interval of 0.05; the number of partitions of the parallel algorithm was set to 4. After completing the clustering calculations for all parameter combinations, the clustering accuracy was calculated using the clustering results with the best performance of the serial and parallel algorithm clustering evaluation indexes respectively. The experimental results show that the present invention maintains good clustering accuracy compared with the serial CDC as

[0123] shown. The POI data of national interest points of Amap was selected, and the data set was divided into different scale levels of 100,000 - 1,000,000 for experiments to verify the parallel acceleration effect of the present invention. The range of the clustering algorithm parameter k was set to 30 - 50; the value interval of Ratio was 10, and the range was set to 0.05 - 0.3, with a value interval of 0.05. The nearest neighbor search parameter K was set to 32, and Imax was set to 11. The partition sampling rate was set to 1%, and the number of partitions was set to 16, 32, 64, 128. To ensure the reliability of the experimental results, the execution time was taken as the average of the execution times of different parameter combinations of the algorithm. The experimental results show that the present invention has a significant performance improvement compared with the serial CDC algorithm. As the scale of the data set increases, the execution time of the present invention increases approximately linearly, gradually increasing from 14.40 seconds at a data scale of 100,000 to 1091.96 seconds at a scale of 1,000,000 as Figure 3 shown, and the growth rate significantly slows down.

[0124] Figure 2 and Figure 3 also compared the original single - machine algorithm with the parallel acceleration algorithm proposed in this application from the perspectives of clustering accuracy and execution time on real - world data sets, verifying the feasibility and efficiency of the method of this application, and showing significant progress.

[0125] Next, taking Apache Spark as an example, the implementation process will be described. The test single - machine configuration is a 4 - core 8 - thread 3.40GHz CPU and 16G of memory, and the operating system is Windows. The calculation process is as Figure 4 shown.

[0126] The present invention improves the computability of the algorithm for large - scale data through the distributed optimization and acceleration method of the local direction centrality clustering algorithm, while improving the utilization rate of the computing resources of the distributed system by the algorithm and the application potential of the distributed system for spatial data, and assisting various data mining applications.

[0127] Next, the algorithm process of the present invention will be elaborated in detail in combination with the accompanying drawings of the present invention. The specific steps are as follows:

[0128] After submitting the environmental parameters, clustering algorithm parameters, partitioning parameters, and nearest neighbor search parameters to the master node, read the data to be clustered stored in the distributed file system HDFS;

[0129] Build a global index structure of a priority search K-means tree for the data to be clustered on the master node as Figure 5 shown, and broadcast it to the Spark cluster;

[0130] Combine data sampling and Hilbert curve partitioning method to construct a spatial partition for the sampling points as Figure 6 shown, and distribute the data to the worker nodes according to the partition;

[0131] Parallelly execute the CDC local clustering algorithm on each worker node as Figure 7 shown, including steps of nearest neighbor search, data partitioning, and connecting internal points. That is, within each partition, first perform KNN nearest neighbor search on the data to be clustered on the global index structure as Figure 8 shown and calculate the DCM value, then divide the internal points and boundary points according to the threshold of DCM, and finally merge the internal points based on the reachable distance from the internal points to the boundary points and mark the internal point cluster ID, search for the nearest internal point of the boundary point and mark the boundary point cluster ID to obtain local clusters;

[0132] On the master node, perform inter-partition cluster merging according to the maximum reachable distance of the local clusters as Figure 9 shown, and update the cluster ID at the same time to generate complete clusters;

[0133] Output the clustering result to HDFS.

[0134] Embodiment 2

[0135] Based on the same inventive concept, this embodiment provides a data distributed clustering device based on local direction centrality, including:

[0136] Initialization module 1, used to receive the parameters required for the clustering task, including environmental parameters, clustering algorithm parameters, partitioning parameters, and nearest neighbor search parameters, configure and register a serializer, and read the complete data to be clustered from the distributed file system;

[0137] Global index construction module 2, used to construct a global index of a priority search K-means tree based on the complete data to be clustered, and share the global index to each worker node through the master node of the distributed cluster;

[0138] Data partitioning module 3, used to partition the complete data to be clustered by combining data sampling and Hilbert curve partitioning method, and obtain the corresponding partition ID, and send the partition data corresponding to the partition ID to the corresponding worker node through the master node of the distributed cluster;

[0139] The local clustering module 4 is used to parallelly execute the CDC local clustering algorithm through each working node of the distributed cluster, specifically including: the working nodes perform k-nearest neighbor search on the partitioned data respectively through the shared global index and calculate the DCM value, and divide the internal points and boundary points according to the relationship between the DCM value and the DCM threshold, then merge the internal points based on the reachable distance from the internal points to the boundary points, the merged internal points are classified into the same internal point cluster, mark the internal point cluster ID, search for the nearest internal point to the boundary point and mark the boundary point cluster ID to obtain the local cluster, where the DCM value is the angular variance formed by the data point and its k neighboring points in the two-dimensional space;

[0140] The global merging module 5 is used to merge the local clusters between partitions according to the maximum reachable distance of the local clusters through the master node of the distributed cluster to generate a complete cluster as the clustering result;

[0141] The result output module 6 is used to output the clustering result to the distributed file system.

[0142] Since the device introduced in the second embodiment of the present invention is the device adopted for implementing the data distributed clustering method based on local directional centrality in the first embodiment of the present invention, as Figure 11 shown, therefore, based on the method introduced in the first embodiment of the present invention, those skilled in the art can understand the specific structure and deformation of the device, so it will not be elaborated here. Any device adopted by the method in the first embodiment of the present invention belongs to the scope protected by the present invention.

[0143] Embodiment III

[0144] Based on the same inventive concept, please refer to Figure 12 , the present invention also provides a computer-readable storage medium 300, on which a computer program 311 is stored, and when the program is executed, it implements the method described in Embodiment I.

[0145] Since the computer-readable storage medium introduced in the third embodiment of the present invention is the computer-readable storage medium adopted for implementing the data distributed clustering method based on local directional centrality in the first embodiment of the present invention, therefore, based on the method introduced in the first embodiment of the present invention, those skilled in the art can understand the specific structure and deformation of the computer-readable storage medium, so it will not be elaborated here. Any computer-readable storage medium adopted by the method in the first embodiment of the present invention belongs to the scope protected by the present invention.

[0146] Embodiment IV

[0147] Based on the same inventive concept, the present application also provides a computer device, as Figure 13As shown, it includes a memory 401, a processor 402, and a computer program 403 stored on the memory and executable on the processor. When the processor executes the above program, it implements the method in the first embodiment.

[0148] In the specific implementation process, the implementation framework in the computer-readable storage medium or computer device of the present invention is a distributed architecture, specifically as Figure 14 shown. This distributed architecture is a master-slave distributed structure including a distributed file system, a master node, and worker nodes.

[0149] Since the computer device introduced in the fourth embodiment of the present invention is the computer device used to implement the data distributed clustering method based on local direction centrality in the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, those skilled in the art can understand the specific structure and variations of this computer device, so it will not be elaborated here. Any computer device used in the method of the first embodiment of the present invention belongs to the scope protected by the present invention.

[0150] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0151] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0152] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concepts. Therefore, the appended claims are intended to be interpreted to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.

[0153] Obviously, those skilled in the art can make various modifications and variations to the embodiments of the present invention without departing from the spirit and scope of the embodiments of the present invention. Thus, if these modifications and variations of the embodiments of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A data distributed clustering method based on local direction centrality, characterized in that, Including: S1: Receive the parameters required for the clustering task, including environment parameters, clustering algorithm parameters, partitioning parameters, and nearest neighbor search parameters, configure and register a serializer, and read the complete data to be clustered from the distributed file system; S2: Build a priority search K-means tree global index based on the read complete data to be clustered, and share the global index with each worker node through the master node of the distributed cluster; S3: Partition the complete data to be clustered by combining data sampling and Hilbert curve partitioning method, and obtain the corresponding partition ID. Send the partition data corresponding to the partition ID to the corresponding worker node through the master node of the distributed cluster; S4: Each worker node in the distributed cluster executes the CDC local clustering algorithm in parallel, specifically including: the worker node performs k-nearest neighbor search on the partition data respectively through the shared global index and calculates the DCM value, and divides the internal points and boundary points according to the relationship between the DCM value and the DCM threshold. Then, based on the reachable distance from the internal points to the boundary points, the internal points are merged. The merged internal points are grouped into the same internal point cluster, mark the internal point cluster ID, search for the nearest internal point to the boundary point and mark the boundary point cluster ID to obtain the local cluster, where the DCM value is the angular variance formed by the data point and its k neighboring points in the two-dimensional space; S5: The master node of the distributed cluster merges the local clusters between partitions according to the maximum reachable distance of the local clusters to generate the complete cluster as the clustering result; S6: Output the clustering result to the distributed file system; Among them, step S5 includes: For local clusters Sort the reachable distances of each internal point in ; Cluster merging between intervals is performed according to the connection rule that the maximum reachable distance and the distance between two clusters are not greater than the sum of the reachable distances of the two clusters, that is , and the cluster ID is updated to generate a complete cluster, where are two different local clusters, are respectively the local clusters of the maximum reachable distance, is the distance between the internal points that obtain the maximum reachable distance in 2. The data distributed clustering method based on local direction centrality according to claim 1, wherein Step S1 includes: S1.1: The distributed cluster receives the parameters required for the clustering task. Among them, the environment parameters include the file path, the clustering algorithm parameters include the number of neighbors and the boundary point ratio, the partitioning parameters include the partitioning type, the partitioning sampling rate ratio, and the number of partitions, and the nearest neighbor search parameters include the index type parameter, the construction parameter, and the search parameter; S1.2: Register the serializer for the geometric type object and the index; S1.3: Read the complete data to be clustered in the distributed file system according to the file path and perform projection conversion.

3. The data distributed clustering method based on local direction centrality according to claim 2, wherein Step S2 includes: S2.1: Initialize the index structure according to the index type parameter and the construction parameters, where the construction parameters include the branching factor branch, the maximum number of iterations I of K-means max and the initial centroid selection method C alg ; S2.2: Calculate the centroid of the complete data to be clustered and build the root node of the index tree; S2.3: According to C alg Select the centroid of branch initial partitions and partition the data into the nearest partitions; S2.4: Update the partition centroids and re-partition the data until the partition centroids remain unchanged or the update reaches I max ; S2.5: Build nodes according to the partition centroid and add them to the child node set of the parent node; S2.6: Repeat steps S2.3 to S2.5 until the number of data in the partition is less than branch to obtain the constructed priority search K-means tree global index, and represent it with a variable; S2.7: Distribute the priority search K-means tree global index variable to each worker node through the master node of the distributed cluster.

4. The data distributed clustering method based on local direction centrality according to claim 2, characterized in that Step S3 includes: S3.1: Perform data sampling on the complete data to be clustered according to the partitioning sampling rate ratio; S3.2: Calculate the Hilbert coding value of the sampled data and sort the sampled points according to the value size; S3.3: Evenly divide the sampled points into the corresponding number of intervals as the number of partitions, record the division position as the partition; S3.4: Expand from the sampling points to form a rectangular partition range, and generate partition IDs for all the data. S3.5: Distribute the corresponding partition data to each working node in the cluster according to the partition IDs.

5. The data distributed clustering method based on local direction centrality according to claim 2, wherein Step S4 includes: S4.1: Perform k-nearest neighbor search on the data to be clustered in each partition based on the global index structure of the priority search K-means tree to obtain k neighbor points of the data points. S4.2: Calculate the angular variance DCM formed by each data point and its k neighbor points in the two-dimensional space. where k is the number of neighbor points of the data point, is the angle formed by the data point and its neighbor points, and satisfies ; S4.3: Merge the DCM value results of all data points in the partition and sort them, and then calculate the threshold according to the boundary point ratio parameter If the DCM is less than then mark the data point as an internal point, otherwise mark it as a boundary point; calculate the threshold S4.4: Calculate internal points and all boundary points to obtain the reachable distance of the internal points by taking the minimum distance between them , that is ; where is the internal point is the boundary point is the distance between the internal point and the boundary point . S4.5: Merge the internal points according to the connection rule that the distance between two internal points is not greater than the sum of the reachable distances of the two points, and mark the internal point cluster ID, that is ; where are the reachable distances of internal point and internal point respectively, is the distance between internal point and internal point . S4.6: Search for the internal point closest to the boundary point, and use the cluster ID of the internal point where the searched internal point is located as the boundary point cluster ID to obtain the local clusters.

6. The data distributed clustering method based on local direction centrality according to claim 1, characterized in that, After step S6, the method further includes: Outputting the clustering evaluation result and the calculation time result to the distributed file system.

7. A data distributed clustering device based on local direction centrality, characterized in that It includes: An initialization module, configured to receive the parameters required for the clustering task, including environment parameters, clustering algorithm parameters, partition parameters, and nearest neighbor search parameters, configure and register a serializer, and read the complete data to be clustered from the distributed file system. A global index construction module, configured to construct a global index of the priority search K-means tree based on the read complete data to be clustered, and share the global index to each working node through the master node of the distributed cluster. A data partitioning module, configured to partition the complete data to be clustered by combining data sampling and the Hilbert curve partitioning method, and obtain the corresponding partition IDs, and send the partition data corresponding to the partition IDs to the corresponding working nodes through the master node of the distributed cluster. A local clustering module, configured to execute the CDC local clustering algorithm in parallel through each working node of the distributed cluster. Specifically, the working node performs k-nearest neighbor search on the partition data respectively through the shared global index and calculates the DCM value, divides the internal points and boundary points according to the relationship between the DCM value and the DCM threshold, then merges the internal points based on the reachable distance from the internal points to the boundary points, and the merged internal points are grouped into the same internal point cluster, mark the internal point cluster ID, search for the internal point closest to the boundary point and mark the boundary point cluster ID to obtain the local clusters, where the DCM value is the angular variance formed by the data point and its k neighboring points in the two-dimensional space. A global merging module, configured to merge the local clusters between partitions according to the maximum reachable distance of the local clusters through the master node of the distributed cluster to generate complete clusters as the clustering result. A result output module, configured to output the clustering result to the distributed file system. The global merging module is specifically configured to: For local clusters Sort the reachable distances of each internal point in the local clusters to obtain the maximum reachable distance ; Cluster merging between intervals is performed according to the connection rule that the maximum reachable distance and the distance between two clusters are not greater than the sum of the reachable distances of the two clusters, that is , and update the cluster ID to generate a complete cluster, where are two different local clusters are respectively the local clusters of the maximum reachable distance is the distance between the internal points that obtain the maximum reachable distance in 8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed, it implements the method described in any one of claims 1 to 6.

9. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Big data space-oriented data local density clustering method

    CN111652305A

  • Data clustering method and device, electronic equipment and readable storage medium

    CN114219023A