A multi-GPU distributed parallel K-Means clustering method
By using the Ring-Allreduce framework to optimize the K-Means clustering algorithm in a multi-GPU environment, the problems of increased computing costs and high data communication costs in the existing multi-GPU strategy are solved, and efficient distributed parallel computing for multi-GPUs is achieved.
Patent Information
- Application Number
- CN202210859003.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-20
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-07-20
AI Technical Summary
When existing multi-GPU strategies process large data volumes, the computing cost increases, and data parallelism strategies may not necessarily achieve optimal computing efficiency, which has the problem of high cost of data communication between GPUs.
The multi-GPU distributed parallel K-Means clustering method based on the Ring-Allreduce framework is adopted to optimize data communication between GPUs through the Ring Allreduce algorithm, reduce the number of computing resource contention, and significantly reduce the cost of communication between GPUs.
It realizes efficient calculation of K-means algorithm in multi-GPU environment, minimizes computing resource contention, significantly reduces computing costs, and solves the problem of linear increase in computing costs when the data volume is large.
Smart Images

Figure CN115114031B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data mining, and in particular to a multi-GPU distributed parallel K-Means clustering method based on a Ring-Allreduce framework. Background Art
[0002] K-Means clustering algorithm is one of the most commonly used clustering algorithms in fields such as data mining. This algorithm is suitable for processing data sets with a large number of samples. The use of parallel computing is a common method to improve the efficiency of machine learning methods. Its main principle is to allocate different computing resources to data operations that have the same and independent computing methods, so as to better improve processor performance and algorithm efficiency. The optimization efficiency of parallel computing is related to the hardware performance of computing processing and parallel strategy. The parallel strategy is how many threads, blocks, grids and other GPU resource strategies are allocated to data operations, and the multi-GPU parallel strategy is a parallel computing method that can handle larger amounts of data.
[0003] The multi-GPU strategy currently used in the K-Means clustering algorithm usually uses the traditional data parallel strategy. The data parallel distributed strategy divides the clustering data set into multiple subsets and then distributes them to different GPUs. Each GPU runs the K-Means clustering algorithm separately. The data parallel strategy used by the K-Means clustering algorithm can achieve certain computational efficiency optimization. However, in the calculation of large data volumes in industrial software and other applications, the data parallel strategy may not achieve the optimal computational efficiency. There are several problems: (1) As the number of GPUs increases, the computational cost increases. (2) As the amount of computational data increases, the computational cost increases.
[0004] Liao Liefa and others from Jiangxi University of Science and Technology proposed a parallel K-means clustering algorithm based on Spark and ASPSO [1]. The data set is roughly divided by the partition function, and the network partitioning strategy PCCV is used to calculate the data network correlation coefficient and then divide the data network to obtain network units. The SPFG strategy is used to cover the local area of the data points, update the sample points in the data set, and obtain the number of local clusters. The ASPSO strategy is used to calculate the adaptive parameters and obtain the local cluster centroid. The CRNN strategy is used to calculate the cluster radius of each cluster, and the similarity is judged according to the cluster similarity function. The Spark parallel computing framework is combined to merge the clusters with large similarity, and then the clustering results are output to improve the efficiency of clustering calculation.
[0005] Since the optimization technology based on Spark and ASPSO parallelization is still a data parallel type of optimization method, the data parallel strategy may not achieve the optimal efficiency improvement for the K-Means clustering algorithm. The computing cost of using this strategy will increase with the application of increasing data volume, because as the number of GPUs in the computing environment increases, the data communication cost between GPUs will increase, which will increase the computing cost of the entire algorithm. Summary of the invention
[0006] In view of the defects of the prior art, the present invention provides a multi-GPU distributed parallel K-Means clustering method.
[0007] In order to achieve the above invention object, the technical solution adopted by the present invention is as follows:
[0008] A multi-GPU distributed parallel K-Means clustering method includes the following steps:
[0009] S1: Divide the remote sensing image into the expected number of clusters, that is, set the number of clusters k, select k initial cluster centers Q = {Q1...Qk}, set the numbers of k GPUs involved in the calculation, that is, number them GPU1 to GPUk, and pass the k cluster center data and the entire data set into k data sets respectively.
[0010] Among them, there is an i-th (1<=i<=k) GPU, 1<=i<=k, which will transmit the data set and the i-th numbered cluster center Qi from the CPU. The GPU allocates independent and different threads or blocks of parallel computing resources to each data.
[0011] S2: Calculate the distance Distance(Qi,i) between the data of the i-th thread of the i-th GPU and the i-th cluster center on the GPU.
[0012] S3: GPU (i+1) calculates the distance Distance (Qi+1, i+1) from the cluster center Qi+1 to the i+1th thread data.
[0013] S4: Pass Distance(Qi,i) from GPUi to GPU(i+1), and compare the size of Distance(Qi+1,i+1) with that of GPU(i+1). The smaller distance will be recorded on GPU(i+1).
[0014] S5: Steps from S2 to S4 are repeatedly executed for the data at the i-th position until the minimum distance record Distance at the same position of all GPUs is the same.
[0015] S6: Steps from S2 to S5 are repeatedly executed for the data at the i+kth position until all the data on the GPU find the nearest cluster center.
[0016] S7: Inter-GPU traversal distance record.
[0017] S8: Calculate the mean of all data in cluster i, calculate the new cluster center and update Qi.
[0018] S9: Repeat S1 to S8 until the cluster center does not change. Finally, a cluster with 5 points as the cluster center is obtained, and the data set is divided into 5 classes. End.
[0019] Preferably, the calculation used in S2 is Euclidean, which is calculated as follows:
[0020]
[0021] Furthermore, S7 is specifically as follows: GPUi traverses the data distance records to check whether there is data Distance(Qi,i+nk)(1<=i<=k)&&(n∈N) closer to the cluster center Qi; if so, the data point i+nk is assigned to cluster i.
[0022] Compared with the prior art, the advantages of the present invention are:
[0023] (1) It can reduce the data communication cost between GPUs in a multi-GPU environment.
[0024] (2) It is implemented based on the Ring-Allreduce framework, with high computational efficiency and easy implementation.
[0025] (3) This minimizes the amount of computing resource contention and significantly reduces the computational cost of the K-means algorithm in a multi-GPU environment, solving the problem of linear increase in the computational cost of the K-means algorithm when the amount of data is large. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 is a Ring-Allreduce architecture diagram of an embodiment of the present invention;
[0027] Figure 2 is a data flow diagram between GPUs after the optimization algorithm of an embodiment of the present invention;
[0028] Figure 3 is a schematic diagram of three-dimensional remote sensing image data according to an embodiment of the present invention;
[0029] Figure 4 is a three-dimensional remote sensing image clustering diagram after cluster analysis in an embodiment of the present invention;
[0030] Figure 5 It is a flow chart of the K-Means clustering method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0031] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and examples.
[0032] K-Means Clustering Algorithm
[0033] K-mean algorithm is also known as k-means algorithm. The k in K-means algorithm means that the cluster is divided into k clusters, and means means taking the mean of the data values in each cluster as the center of the cluster, or the centroid, that is, using the centroid of each class to describe the cluster. The algorithm idea is roughly that for a given sample set, K-means divides the sample set into K clusters according to the distance between samples. Let the points in the cluster be as closely connected as possible, and let the distance between clusters be as large as possible. The specific steps are as follows:
[0034] 1 Select the K value. The choice of K is usually determined according to demand.
[0035] 2. Randomly select K samples from the samples as the initial cluster centers.
[0036] 3. Calculate the distance E between the data sample and each cluster center. Usually, the distance is set as the Euclidean distance. Record the cluster center of the nearest neighbor of the data and divide the data into the cluster where the nearest neighbor cluster center is located. Euclidean distance calculation:
[0037]
[0038] 4. Calculate the average of the data contained in all clusters and recalculate the new cluster center
[0039] 5 Repeat steps 3 and 4 until the cluster center no longer changes. The operation ends and the cluster formed is the final clustering result.
[0040] Ring Allreduce Algorithm
[0041] The communication cost of the GPU communication strategy modified by the Ring Allreduce algorithm is constant and independent of the number of GPUs in the system, and is determined only by the slowest connection between the GPUs in the system; if only bandwidth is considered as a factor in the communication cost (and latency is ignored), then Ring Allreduce is an optimal communication algorithm. Figure 2 As shown in the figure, the principle of the Ring Allreduce algorithm is to arrange the GPUs in a logical ring. Each GPU should have a left neighbor and a right neighbor; it will only send data to its right neighbor and receive data from its left neighbor. The algorithm is divided into two steps, scatter-reduce and all-gather:
[0042] 1In the scatter-reduce step, the GPUs will exchange data so that each GPU will eventually get a part of the final result. When the data is divided into N blocks, each GPU has a block of the same size to receive the data, and then each GPU performs N-1 iterations. The GPU only receives data from one GPU, the data in the GPU block is the same as the received data, and calculates the new data. At the same time, the GPU will only send the data to another GPU, thus completing a data iteration, and continue to iterate the calculation of different blocks within the GPU to calculate the sum until all values are calculated in all blocks.
[0043] 2 In the all-gather step, the GPUs will exchange these blocks so that all GPUs will eventually get the complete final result. At this time, each GPU has a block that stores all GPU data in the block and updates it to the received final calculated value. At the same time, the GPU sends the block of data storing the final calculated value to the other GPU, completing a data iteration, and continuing until all blocks of each GPU are stored with the final calculated value.
[0044] In common multi-GPU parallel computing, the communication cost increases linearly with the number of GPUs. Each N GPU in the Ring-AllReduce architecture will send and receive values N-1 times for scatter-reduce and N-1 times for all-gather. Each GPU will send a K / N value, where K is the amount of data that needs to be Ring-Allreduced on each GPU. When the inter-GPU data communication bandwidth is B, the total amount of data transferred to each GPU DataTransferred is calculated as follows, and the time T to transfer this data is calculated as follows:
[0045]
[0046] T = DataTransferred / B
[0047] It can be seen that the transmission overhead of Ring-Allreduce communication basically eliminates the impact of the number of working nodes and does not increase with the increase of the number of GPUs N. The overall communication speed is only affected by the adjacent GPUs in the slowest connection ring.
[0048] Optimizing K-Means Clustering Algorithm through Ring Allreduce Framework
[0049] The present invention uses the Ring-Allreduce principle to optimize the K-Means algorithm, that is, the communication speed of the K-Means algorithm after the optimization of the present invention is only limited by the slowest (lowest bandwidth) connection between adjacent GPUs in the ring. Given the correct selection of the neighbor position of each GPU, the algorithm of the present invention is the fastest algorithm with optimal bandwidth, which minimizes the amount of computing resource contention and significantly reduces the communication cost between GPUs.
[0050] The data flow of the K-mean algorithm is optimized by the method of the present invention. Figure 2 As shown, the data stream is transmitted between multiple GPUs, and a single GPU only obtains the data stream from an adjacent GPU, performs only one distance operation, and transmits the operation result to another adjacent GPU. The present invention proposes a K-mean algorithm optimized based on the Ring-All reduce principle. The implementation case of the present invention is a cluster analysis of three-dimensional satellite remote sensing image data. The two-dimensional pixel display of the three-dimensional satellite remote sensing image data set is as follows Figure 3 As shown in the figure, the cluster image after the cluster analysis of the three-dimensional image data is completed is as follows Figure 4 The existing optimized K-means method has not been designed for parallel computing of three-dimensional data in a multi-GPU environment, while the K-means algorithm optimized by the present invention is suitable for cluster analysis of three-dimensional data, and can fully utilize the computing performance of a multi-GPU environment (the memory usage rate of each GPU reaches more than 95%) without affecting the cluster analysis results.
[0051] The optimized K-means algorithm flow chart of the invention is as follows: Figure 5 As shown, the specific steps of the present invention will be described below:
[0052] (1) According to the expected number of clusters of the remote sensing image, the image is divided into five categories, such as water, building, land, woodland, and grassland. That is, the number of clusters k is set to 5, and five initial cluster centers Q = {Q1…Q5} are selected. The five GPUs involved in the calculation are numbered from GPU1 to GPU5, and the five cluster center data and the entire data set are respectively transferred to the five data sets.
[0053] Among them, the i-th (1<=i<=5) GPU (the activities of all GPUs are the same as those of the i-th GPU) will receive the data set and the i-th cluster center Qi from the CPU, and the GPU will allocate independent and different threads or blocks of parallel computing resources to each data.
[0054] (2) The distance between the data of the ith thread of the ith GPU and the center of the ith cluster on the GPU is calculated. Distance(Qi,i). This embodiment uses the Euclidean method to calculate the distance. The calculation is as follows:
[0055]
[0056] (3) GPU (i+1) calculates the distance Distance (Qi+1, i+1) from the cluster center Qi+1 to the i+1th thread data.
[0057] (4) Distance(Qi,i) is transferred from GPUi to GPU(i+1) and compared with Distance(Qi+1,i+1) of GPU(i+1). The smaller distance is recorded on GPU(i+1).
[0058] (5) Steps (2) to (4) are repeated for the data at the i-th (1<=i<=5) position until the minimum distance record Distance at the same position of all GPUs is the same.
[0059] (6) Steps (2) to (5) are repeated for the data at position i+5 (1<=i<=5) until all data on the GPU find the nearest cluster center.
[0060] (7) GPUs traverse distance records. For example, GPUi traverses data distance records to check whether there is data Distance(Qi,i+n5)(1<=i<=5)&&(n∈N) that is closer to the cluster center Qi. If so, the data point i+n5 is assigned to cluster i.
[0061] (8) Calculate the mean of all data in cluster i, calculate the new cluster center and update Qi.
[0062] (9) Repeat steps (1) to (8) until the cluster center no longer changes. Finally, we get a cluster with 5 points as the cluster center, and divide the data set into 5 classes. End.
[0063] The communication speed of the method of the present invention is only limited by the slowest (lowest bandwidth) connection between adjacent GPUs in the ring. Given the correct selection of the neighbor position of each GPU, the algorithm of the present invention is the fastest algorithm with optimal bandwidth, which minimizes the amount of computing resource contention and significantly reduces the communication cost between GPUs.
[0064] The method according to the present invention described above can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk or magneto-optical disk), or as a computer code that is originally stored in a remote recording medium or a non-temporary machine-readable medium downloaded through a network and will be stored in a local recording medium, so that the method described herein can be stored in such software processing on a recording medium using a general-purpose computer, a special-purpose processor or programmable or special-purpose hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component (e.g., RAM, ROM, flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by a computer, a processor or hardware, the processing method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the processing shown here, the execution of the code converts the general-purpose computer into a special-purpose computer for executing the processing shown here.
[0065] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the implementation methods of the present invention, and should be understood that the protection scope of the present invention is not limited to such special statements and embodiments. Those skilled in the art can make various other specific variations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed by the present invention, and these variations and combinations are still within the protection scope of the present invention.
Claims
1. A multi-GPU distributed parallel K-Means clustering method, characterized in that: The following steps are involved: S1: According to the expected number of clusters of the remote sensing image, the number of clusters k is set, k initial cluster centers Q = {Q1…Qk} are selected, the k GPU numbers involved in the calculation are set, i.e., numbered from GPU1 to GPUk, and the k cluster center data and the entire data set are respectively transferred into the k data sets; The i-th (1<=i<=k) GPU, 1<=i<=k, will receive the data set and the i-th numbered cluster center Qi from the CPU, and the GPU will allocate independent and different threads or blocks of parallel computing resources to each data; S2: Calculate the distance Distance(Qi,i) between the data of the i-th thread of the i-th GPU and the i-th cluster center on the GPU; S3: GPU (i+1) calculates the distance Distance (Qi+1, i+1) from the cluster center Qi+1 to the i+1th thread data; S4: Transfer Distance(Qi,i) from GPUi to GPU(i+1), compare it with Distance(Qi+1,i+1) of GPU(i+1), and the smaller distance will be recorded on GPU(i+1); S5: Steps from S2 to S4 are repeatedly executed for the data at the i-th position until the minimum distance record Distance at the same position of all GPUs is the same; S6: Steps from S2 to S5 are repeatedly executed for the data at the i+kth position until all the data on the GPU find the nearest cluster center; S7: Inter-GPU traversal distance record; S8: Calculate the mean of all data in cluster i, calculate the new cluster center and update Qi; S9: Repeat S1 to S8 until the cluster center no longer changes; finally, a cluster with 5 points as the cluster center is obtained, and the data set is divided into 5 classes, and then it ends.
2. The multi-GPU distributed parallel K-Means clustering method according to claim 1, characterized in that: The calculation used in S2 is Euclid, and the calculation is as follows:
3. The multi-GPU distributed parallel K-Means clustering method according to claim 1, characterized in that: S7 is as follows: GPUi traverses the data distance records to check whether there is data Distance(Qi,i+nk)&&(n∈N) that is closer to the cluster center Qi. If so, the data point i+nk is assigned to cluster i.
Citation Information
Patent Citations
High-performance parallel implementation method of K-means algorithm on domestic Sunway 26010 multi-core processor
CN108509270A
Parallel K-means optimization method based on Spark and ASPSO
CN113128617A