Optimization Method and Device for High-Performance Linpack Benchmark Program Based on Heterogeneous Platforms
By adopting the pipelined processing of OpenMP multi-threading and GPU programming model on a heterogeneous platform, the workload of CPU and GPU is reasonably allocated, and the problem of performance gap between CPU and GPU in heterogeneous systems is solved, and efficient HPL benchmark testing is achieved.
Patent Information
- Application Number
- CN202211725671.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-12-30
AI Technical Summary
In heterogeneous systems, the performance gap between CPU and GPU is too large. The HPL benchmark test under the existing heterogeneous computing platform has a communication bottleneck and cannot reflect the optimal performance of the system.
Using OpenMP-based multi-threaded operation method, the computing tasks of CPU and GPU are processed in a pipeline, and the panel decomposition and trailing matrix updates in HPL are used to reasonably allocate the workload of CPU and GPU, and make full use of the parallel computing capabilities of GPU.
It improves the computing efficiency of heterogeneous platforms, reduces the idle time of CPU and GPU, and accurately reflects the optimal performance of heterogeneous systems.
Smart Images

Figure CN115827251B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of high-performance computing, and particularly to an optimization method and device for a high-performance Linpack benchmark program based on a heterogeneous platform. Background Art
[0002] In recent years, computational science using supercomputers as tools has penetrated into all levels of scientific research and engineering design, such as petaflop-level computing, artificial intelligence, etc. A single computer cannot meet the computing requirements of high-performance computing (HPC) applications. Therefore, researchers have begun to apply heterogeneous architectures, which have become an effective method for parallelizing many practical problems and are considered the mainstream architecture of future computing platforms. Especially with the emergence of Graphics Processing Units (GPUs), heterogeneous computing has entered the era of parallel computing, aiming to utilize the computing potential of different types of computing resources to process various workloads. Due to the fact that heterogeneity has gradually become the mainstream, many software developers and hardware manufacturers have started to optimize programs to maximize the utilization and development of the computing capabilities of these platforms. Most of the previous work was based on a CPU-centric design model, that is, storing the original data in the host memory and transferring the data required for matrix operations to the GPU memory through the PCI Express (PCI-E) bus. However, the amount of data required for parallel tasks is too large, and transmitting such a large amount of data always makes the PCI-E bandwidth a performance bottleneck, which leads to an increase in operating overhead and low program efficiency.
[0003] In a heterogeneous system, once more GPU cards are inserted into each node, the performance gap between the CPU and the GPU will rapidly expand. How to reasonably allocate the workload between the CPU and the GPU is a key issue for HPL in a heterogeneous system. Existing HPL benchmarks under heterogeneous computing platforms have bottlenecks in many aspects such as algorithms and communications, and cannot reflect the optimal performance of the system. Therefore, it is very necessary to optimize its performance under a heterogeneous platform. Summary of the Invention
[0004] To solve the technical problems existing in the prior art, the present invention provides an optimization method and device for a high-performance Linpack benchmark program based on a heterogeneous platform. The multi-threaded running method based on OpenMP can give full play to the multi-core performance of the CPU side. The panel decomposition and trailing matrix update in HPL are pipelined using the Stream in the GPU programming model, ensuring that both the CPU and the GPU are fully loaded during the program operation, reducing the idle time of the CPU and the GPU, and accelerating the computing efficiency.
[0005] The first object of the present invention is to provide an optimization method for a high-performance Linpack benchmark program based on a heterogeneous platform.
[0006] The second object of the present invention is to provide a computer device.
[0007] The first object of the present invention can be achieved by adopting the following technical solutions:
[0008] An optimization method for a high-performance Linpack benchmark program based on a heterogeneous platform, the method comprising:
[0009] S1. Initialize the running parameters of HPL according to a configuration file, and use a random number generation algorithm to generate a matrix of a set scale, where the memory space occupied by the matrix completely covers the size of the GPU video memory participating in the calculation;
[0010] S2. Execute the panel decomposition process on the CPU in a multi-threaded parallelization manner using OpenMP to make the utilization rate of all cores of the CPU reach 100%;
[0011] S3. Use a GPU-Aware based Ring broadcast algorithm to broadcast the decomposed panel data and matrix row exchange information to the remaining column processes in the same row;
[0012] S4. After the GPU receives the matrix row exchange information, according to the principle of merged memory access of the GPU and the characteristics of the cache, call a custom kernel to complete the row exchange operation on the matrix stored in the video memory;
[0013] S5. After completing the row exchange operation in the GPU, the GPU uses the received panel matrix data to update the remaining matrix data, calls the DTRSM() function to update the upper triangular matrix and the optimized GEMM function to update the trailing matrix;
[0014] S6. Recursively execute steps S2 to S5 until the entire matrix decomposition is completed, call a timing function to calculate the total time for the matrix to complete the decomposition, and at the same time verify the result to obtain the floating-point operation performance of the heterogeneous platform.
[0015] In a preferred technical solution, the step S1 includes: writing the running parameters of HPL into the configuration file before running HPL, where the running parameters include the size N of the generated matrix, the size NB of the matrix after partitioning, the total number of row processes P, the total number of column processes Q, and the broadcast algorithm; allocate a matrix of size N*(N + 1) according to the size N of the matrix and distribute it to the GPU video memory in each node, and then call the random number generation algorithm of the GPU to initialize the data of the matrix.
[0016] Specifically, the step S5 specifically includes: using the lookahead algorithm with a depth of 1 in HPL, pre-decomposing a panel panel0 of the matrix on the CPU, and secondarily splitting the matrix data outside the panel into a panel panel1 and a matrix M1;
[0017] The panel panel1 updates the upper triangular matrix and the trailing matrix according to the already decomposed panel0;
[0018] The update of the matrix M1 and the decomposition of the panel panel1 on the CPU are calculated in parallel through the Stream in the GPU programming model.
[0019] The second object of the present invention can be achieved by adopting the following technical solutions:
[0020] A computer device includes a processor and a memory for storing the executable program of the processor. When the processor executes the program stored in the memory, it implements the above-mentioned optimization method for the high-performance Linpack benchmark program based on a heterogeneous platform.
[0021] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0022] 1. The present invention provides an optimization method and device for a high-performance Linpack benchmark program based on a heterogeneous platform, and solves the problem of too large a performance gap between the CPU and the GPU caused by inserting more GPUs into each node in a heterogeneous system with multiple GPUs and multiple nodes by reasonably allocating the workload between the CPU and the GPU.
[0023] 2. The present invention uses a multi-threaded running method based on OpenMP to give full play to the multi-core performance of the CPU side.
[0024] 3. The present invention makes full use of the access characteristics of the GPU and efficiently performs access operations such as row swapping in parallel on the GPU.
[0025] 4. The present invention pipelines the computing tasks on the CPU and the GPU to reduce the idle waiting time caused by the large computing power gap between the CPU and the GPU. Description of the Drawings
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on the structures shown in these drawings without creative efforts.
[0027] Figure 1 It is a flowchart of the optimization method for the high-performance Linpack benchmark program based on a heterogeneous platform in an embodiment of the present invention;
[0028] Figure 2 It is a data flow diagram of the optimization method for the high-performance Linpack benchmark program based on a heterogeneous platform in an embodiment of the present invention. Specific embodiments
[0029] Next, the technical solution of the present invention will be further described in detail in conjunction with the drawings and embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. The implementation manners of the present invention are not limited thereto. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0030] The high-performance Linpack (HPL) benchmark program obtains the floating-point computing performance of a computing system by solving a dense non-linear equation system Ax = b, and is used as one of the important reference criteria for the TOP500 ranking. HPL mainly calculates an n-order linear equation system through Gaussian elimination based on partial pivoting, and this process requires 2 / 3n 3 + 2n 2 + O(n) double-precision floating-point additions and multiplications. Since matrix multiplication occupies a large amount of running time in the HPL benchmark, and the graphics processing unit GPU in a heterogeneous system is good at processing a large number of parallel tasks, such as matrix multiplication.
[0031] The Message Passing Interface (MPI) is a standardized portable message passing interface for running on various parallel computing architectures. It defines a series of high-level interfaces, making the software development process simpler and hiding the complexity of communication between processes.
[0032] Through the open-source software stack ROCm that supports graphics processing unit computing, the present invention uses the Heterogeneous compute Interface for Portability (HIP) to accelerate a large number of parallel computing tasks such as matrix multiplication and matrix row swapping. HIP is a programming model made by AMD and provides a series of toolkits to develop code running on the GPU. This process usually requires multiple device-side execution streams to achieve peak utilization, and the streams run according to the first-in, first-out (FIFO) principle when the GPU operates.
[0033] Embodiment 1:
[0034] The present invention optimizes the problems in HPL under heterogeneous platforms, such as high communication latency between the CPU and GPU, unreasonable task scheduling, and low efficiency of the memory access kernel. After optimization, the HPL first performs parallel panel decomposition. After broadcasting the decomposed panels, it calls the row exchange kernel function optimized for the characteristics of GPU memory access and matrix multiplication (GEMM) to complete the update of the trailing matrix. Finally, the above process is recursively performed to complete the overall calculation. In addition, considering the high communication latency between the CPU and GPU in HPL, the present invention uses the stream in the GPU programming model to pipeline the panel decomposition and trailing matrix update in HPL, ensuring that both the CPU and GPU are fully loaded during the program execution, reducing the idle time of the CPU and GPU, and further accelerating the calculation efficiency.
[0035] As Figure 1 shown in the flowchart of the optimization method for the high-performance Linpack benchmark program based on a heterogeneous platform. The optimization method for the high-performance Linpack benchmark program based on a heterogeneous platform includes the following steps:
[0036] Step 1: Initialize the running parameters of HPL according to the configuration file, and use a random number generation algorithm to generate a matrix of a specific scale, where the memory space occupied by the matrix completely covers the size of the GPU video memory participating in the calculation;
[0037] Specifically, the configuration file contains all the running parameters required for HPL operation. Before running HPL, all the running parameters required for HPL operation are written into the configuration file. The running parameters include the size N of the generated matrix, the size NB of the matrix after partitioning, the total number of row processes P, the total number of column processes Q, and the broadcast algorithm. After specifying N, a matrix of size N*(N + 1) can be allocated and distributed to the GPU video memory in each node. After completing the allocation of the video memory space, the random number generation algorithm of the GPU is called to initialize the data of the matrix. The steps for generating random numbers include: initializing the random number generator using the rocrand_create_generator() function, then setting a fixed random seed by calling rocrand_set_seed(), and finally calling rocrand_generate_uniform_double() to generate double-precision floating-point random numbers that conform to a uniform distribution in the video memory.
[0038] Step 2: Perform the panel decomposition process on the CPU in a parallelized manner using OpenMP multi-threading, so that the utilization rate of all cores of the CPU reaches 100%;
[0039] Specifically, the process of performing panel decomposition on the CPU in the form of OpenMP multi-thread parallelization includes invoking the OpenMP multi-thread library to parallelize the decomposition calculation of the panel, and setting the number of OpenMP threads to core / P0, where core is equal to the number of CPU cores in a node, and P0 is the number of row processes of the current node. Therefore, the utilization rate of all CPU cores in the node can reach 100%, maximizing the computing efficiency on the CPU side.
[0040] Step 3: Use the GPU-Aware based Ring broadcast algorithm to broadcast the decomposed panel data and matrix row exchange information to the remaining column processes in the same row;
[0041] Specifically, after the decomposition of the panel is completed, the panel data is broadcast to the column processes in the same row process. The process of broadcasting the panel data specifically includes: setting the process with the panel data as the root process, and the root process sends the panel data to an adjacent single column process; after receiving the data, the adjacent single column process further forwards it to another adjacent process, and finally broadcasts the panel data in a "ring" manner. This broadcast method can effectively reduce the data communication volume of the root process.
[0042] Preferably, the optimized broadcast algorithm is implemented based on GPU-Aware MPI. Since each process is bound to a GPU, it is possible to achieve point-to-point data transmission from GPU to other GPUs without the need to transfer through the CPU. The steps of using the GPU-Aware based Ring broadcast algorithm to broadcast the decomposed panel data and matrix row exchange information to the remaining column processes in the same row specifically include:
[0043] At one end, the GPU uses MPI_Send() to send data, and at the other end, the GPU uses MPI_Recv() and MPI_Iprobe() to receive data. The data buffers in MPI_Send() and MPI_Recv() both use the video memory address in the GPU.
[0044] Step 4: After the GPU receives the matrix row exchange information, according to the principle of combined memory access of the GPU and the characteristics of the cache, call the custom kernel to complete the row exchange operation on the matrix stored in the video memory;
[0045] Specifically, when the GPU receives the broadcast row exchange information, the row data is stored in the programmable cache LDS (Local Data Share). After all the data of a row of the matrix is read into the LDS, synchronize all GPU threads, and write the data in adjacent storage spaces into the video memory of the GPU to complete the row exchange operation.
[0046] In this implementation, when the GPU receives the broadcast row exchange information, it can use the programmable cache LDS in the GPU hardware structure to implement an efficient row exchange operation. LDS is a cache in the GPU that provides a larger read / write bandwidth than global memory. Since LDS is shared by different threads in the same thread block in the GPU, it provides a cooperation mechanism for the threads and enables synchronization operations for different threads. The GPU also has the characteristic of coalesced memory access, that is, when adjacent threads write data to adjacent global memory spaces, the clock cycles spent on access can be reduced. Therefore, during the row exchange process of HPL, the single-row data of the matrix is stored in the LDS. After all the data in this row is read into the LDS, the __synchtreads() statement is called to synchronize all the GPU threads in the same thread block. Then, by taking advantage of the coalesced memory access characteristic of the GPU again, the data in adjacent storage spaces is written from the LDS to the video memory of the GPU, thus efficiently completing the row exchange operation.
[0047] Step Five: After completing the row exchange operation in the GPU, the GPU updates the remaining matrix data using the received panel matrix data, including calling the DTRSM() function to update the upper triangular matrix and the optimized GEMM (General Purpose Matrix Multiplication) function to update the trailing matrix;
[0048] Specifically, in HPL, a look ahead algorithm with a depth of 1 is used, that is, on the CPU, a panel panel0 of the matrix is pre-decomposed according to Step Two, and at the same time, the matrix data outside the panel is secondarily sliced into a panel panel1 and a matrix M1.
[0049] In the look ahead algorithm, the panel panel1 calls the DTRSM() function to update the upper triangular matrix and the GEMM() function to update the trailing matrix according to the already decomposed panel0.
[0050] After completing the update of the panel panel1, the update of the matrix M1 and the decomposition of the panel panel1 on the CPU can be parallelized through the Stream in the GPU programming model. At this time, the CPU and the GPU perform parallel computing asynchronously, and the computing efficiency is maximized.
[0051] As Figure 2 shown, the data flow chart of the optimization method for the high-performance Linpack benchmark program based on the heterogeneous platform describes the data processing flow of the above steps. Assume that the LU decomposition of part of the data has been completed. For the data composed of and The panel composition performs panel decomposition corresponding to the data processing process in Step 2. Row exchange corresponds to the data processing process in Step 4. After completing panel decomposition and row exchange, for call the DTRSM() function and for A″ i Execute GEMM() corresponding to the data processing process in Step 5.
[0052] Step 6: Recursively execute Steps 2 to 5 until the entire matrix decomposition is completed. Call the timing function to calculate the total time for the matrix to complete decomposition, and at the same time verify the result, that is, obtain the floating-point operation performance of the heterogeneous platform.
[0053] Specifically, when the recursion completes the calculations in Steps 2 to 5, the original matrix in the video memory has been decomposed into a lower triangular matrix L and an upper triangular matrix U, and the data of the original matrix is overwritten. After the decomposition is completed, execute HPL_pdtrsv() to solve the system of equations to obtain the result of the system of equations. The time from Step 2 to solving the system of equations is the index for HPL to evaluate the floating-point performance of a system. Verify the result of the system of equations, that is, obtain the floating-point operation performance of the heterogeneous platform.
[0054] In summary, the present invention can provide an efficient HPL operation scheme for a heterogeneous system based on CPU + GPU. By using the optimized method of OpenMP multi-threading to perform panel decomposition, the computing power of the CPU can be fully utilized; use the GPU-Aware based Ring broadcast algorithm to broadcast the decomposed panel, reducing the overhead of data transfer on the host memory; make full use of the access characteristics of the GPU to perform row exchange operations efficiently in parallel on the GPU; pipeline the computing tasks on the CPU and GPU to reduce the idle waiting time caused by the large computing power gap between the CPU and GPU. The HPL benchmark test under the optimized heterogeneous computing platform can solve the bottlenecks in many aspects such as algorithms and communications to a certain extent, and more accurately reflect the optimal performance of the heterogeneous system.
[0055] Example 2:
[0056] This embodiment provides a computer device, which can be a server, a computer, etc. It includes a processor, a memory, an input device, a display, and a network interface connected through a system bus. The processor is used to provide computing and control capabilities. The memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. When the processor executes the computer program stored in the memory, it implements the optimized method for the high-performance Linpack benchmark test program based on the heterogeneous platform in the above-mentioned Example 1, as follows:
[0057] S1. Initialize the running parameters of HPL according to the configuration file, and use a random number generation algorithm to generate a matrix of a set scale, where the memory space occupied by the matrix completely covers the size of the GPU video memory participating in the calculation;
[0058] S2. Execute the panel decomposition process on the CPU in a parallelized manner with OpenMP multi-threading, so that the utilization rate of all cores of the CPU reaches 100%;
[0059] S3. Use the GPU-Aware based Ring broadcast algorithm to broadcast the decomposed panel data and matrix row exchange information to the remaining column processes in the same row;
[0060] S4. After the GPU receives the matrix row exchange information, according to the principle of combined memory access of the GPU and the characteristics of the cache, call a custom kernel to complete the row exchange operation on the matrix stored in the video memory;
[0061] S5. After completing the row exchange operation in the GPU, the GPU uses the received panel matrix data to update the remaining matrix data, calls the DTRSM() function to update the upper triangular matrix and the optimized GEMM function to update the trailing matrix;
[0062] S6. Recursively execute steps S2 to S5 until the entire matrix decomposition is completed, call the timing function to calculate the total time for the matrix to complete the decomposition, and at the same time verify the result to obtain the floating-point operation performance of the heterogeneous platform.
[0063] Specifically, step S1 includes: writing the running parameters of HPL into the configuration file before running HPL, where the running parameters include the size N of the generated matrix, the size NB of the matrix after block division, the total number of row processes P, the total number of column processes Q, and the broadcast algorithm; allocate a matrix of size N*(N + 1) according to the size N of the matrix and distribute it to the GPU video memory in each node, and then call the random number generation algorithm of the GPU to initialize the data of the matrix.
[0064] The specific content of step S5 includes: using the lookahead algorithm with a depth of 1 in HPL to pre-decompose a panel panel0 of the matrix on the CPU, and perform secondary segmentation on the matrix data outside the panel, segmenting it into panel1 and matrix M1;
[0065] Panel panel1 updates the upper triangular matrix and the trailing matrix according to the already decomposed panel0;
[0066] Parallelize the update of matrix M1 and the decomposition of panel1 on the CPU through the stream in the GPU programming model.
[0067] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A method for optimizing a high-performance Linpack benchmark program based on a heterogeneous platform, characterized in that It includes the following steps: S1. Initialize the running parameters of HPL according to the configuration file, and use a random number generation algorithm to generate a matrix of a set scale, where the memory space occupied by the matrix completely covers the size of the GPU video memory participating in the calculation; S2. Execute the panel decomposition process on the CPU in a parallelized manner with OpenMP multi-threading, so that the utilization rate of all cores of the CPU reaches 100%; S3. Use the GPU-Aware based Ring broadcast algorithm to broadcast the decomposed panel data and matrix row exchange information to the remaining column processes in the same row; S4. After the GPU receives the matrix row exchange information, call a custom kernel to complete the row exchange operation on the matrix stored in the video memory; The step S4 includes: when the GPU receives the broadcast row exchange information, store the row data in the programmable high-speed cache LDS. When all the data of a row of the matrix is read into the LDS, synchronize all GPU threads, and write the data in the adjacent storage space into the video memory of the GPU to complete the row exchange operation; S5. After completing the row exchange operation in the GPU, the GPU uses the received panel matrix data to update the remaining matrix data, and calls the DTRSM() function to update the upper triangular matrix and the GEMM function to update the trailing matrix; S6. Recursively execute steps S2 to S5 until the entire matrix decomposition is completed. Call the timing function to calculate the total time for the matrix to complete the decomposition, and at the same time verify the result to obtain the floating-point operation performance of the heterogeneous platform.
2. The high-performance Linpack benchmark program optimization method based on a heterogeneous platform according to claim 1, wherein The step S1 includes: write the running parameters of HPL into the configuration file before running HPL. The running parameters include the size N of the generated matrix, the size NB of the matrix after block division, the total number P of row processes, the total number Q of column processes, and the broadcast algorithm; allocate a matrix of size N*(N + 1) according to the size N of the matrix and distribute it to the GPU video memory in each node, and then call the random number generation algorithm of the GPU to initialize the data of the matrix.
3. The high-performance Linpack benchmark program optimization method based on a heterogeneous platform according to claim 1, characterized in that The step S2 includes: call the OpenMP multi-threading library to perform parallelized calculation of the decomposition of the panel, and set the number of OpenMP threads to core / P0, where core is the number of cores of the CPU in a single node and P0 is the number of row processes in the current node.
4. The method for optimizing the high-performance Linpack benchmark program based on a heterogeneous platform according to claim 1, wherein The step S3 includes: use MPI_Send() to send data on one end of the GPU, and use MPI_Recv() and MPI_Iprobe() to receive data on the other end of the GPU. The data buffers in MPI_Send() and MPI_Recv() both use the video memory address in the GPU.
5. The high-performance Linpack benchmark program optimization method based on a heterogeneous platform according to claim 1, characterized in that The step S5 specifically includes: use the look ahead algorithm with a depth of 1 in HPL to pre-decompose a panel panel0 of the matrix on the CPU, and perform secondary segmentation on the matrix data outside the panel, segmenting it into panel1 and matrix M1; The panel panel1 updates the upper triangular matrix and the trailing matrix according to the already decomposed panel0; Updating the matrix M1 through streams in the GPU programming model and performing parallelized decomposition calculations on the panel panel1 on the CPU.
6. The optimization method for the high-performance Linpack benchmark program based on a heterogeneous platform according to claim 5, wherein step S6 specifically includes: After the recursive calculation of steps two to five is completed, the original matrix in the video memory has been decomposed into a lower triangular matrix L and an upper triangular matrix U, and the data of the original matrix is overwritten. After the decomposition, HPL_pdtrsv() is executed to solve the system of equations to obtain the result of the system of equations; the result of the system of equations is verified, and based on the time of the result of the system of equations obtained by the solution, the floating-point operation performance of the heterogeneous platform is obtained.
7. A computer device, comprising a processor and a memory for storing processor-executable programs, characterized in that, When the processor executes the program stored in the memory, it implements the optimization method for the high-performance Linpack benchmark program based on a heterogeneous platform according to any one of claims 1-6.
Citation Information
Patent Citations
Method for achieving large-scale high-performance Linpack testing benchmark for GPDSP
CN104615516A
System and method for re-factorizing a square matrix into lower and upper triangular matrices on a parallel processor
US20140196043A1