Computer system and performance test method, device and equipment thereof and medium
Patent Information
- Application Number
- CN202311438966.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-01
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-11-01
Smart Images

Figure CN117370090B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of this disclosure relate to a performance testing method for a computer system, a performance testing apparatus for a computer system, a computer system, an electronic device, and a computer-readable storage medium. Background Technology
[0002] With the development of processors and high-performance networks, the scale of high-performance computing (HPC) applications running in various massively parallel processing (MPP) environments has also grown rapidly. These applications typically involve ultra-large-scale inter-process communication via high-speed interconnect networks to collaboratively solve a large problem.
[0003] The HPL (High Performance Linpack) benchmark is a set of benchmark programs used to measure the actual peak computing performance of high-performance computers / clusters. HPL evaluates the floating-point computing capabilities of high-performance computing clusters by testing the solution of a system of dense linear algebraic equations of degree N in one variable using Gaussian elimination. Summary of the Invention
[0004] At least one embodiment of this disclosure provides a performance testing method for a computer system, wherein the computer system runs M processes, each process including a main thread and a broadcast sub-thread. The method includes: the M processes performing multiple decomposition and update operations on a target matrix, wherein the target matrix includes multi-row, multi-column blocks arranged in an array, the multi-row, multi-column blocks corresponding to the M processes, each decomposition and update operation including: determining a target block from the multi-row, multi-column blocks of the target matrix; the main threads of N processes corresponding to the columns where the target block is located performing a decomposition operation on the target block to obtain the decomposition matrix and row exchange information for this operation; the broadcast sub-threads of the M processes performing inter-process broadcast operations on the decomposition matrix and the row exchange information; and the main thread of each process performing a row exchange operation and a matrix update operation based on the decomposition matrix and the row exchange information; wherein M is an integer greater than 1, and N is a positive integer less than M.
[0005] For example, in the performance testing method provided in at least one example of the above embodiments of this disclosure, the broadcast sub-threads of the M processes perform inter-process broadcast operations for the decomposition matrix and the row swap information, including: the broadcast sub-threads of the N processes broadcasting the row swap information of the current operation to the broadcast sub-threads of other processes among the M processes, so that the other processes start executing the row swap operation; after broadcasting the row swap information of the current operation, the broadcast sub-threads of the N processes broadcast the decomposition matrix of the current operation to the broadcast sub-threads of the other processes.
[0006] For example, in the performance testing method provided in at least one example of the above embodiments of this disclosure, the main thread of each process performs a row swap operation and a matrix update operation based on the decomposition matrix and the row swap information, including: in response to obtaining the row swap information of the current operation, the main thread of each process starts to perform a row swap operation on the target matrix based on the row swap information of the current operation to obtain the swap matrix of the current operation; after obtaining the decomposition matrix of the current operation, the main thread of each process performs a matrix update operation on the target matrix based on the decomposition matrix of the current operation and the swap matrix of the current operation.
[0007] For example, in the performance testing method provided in at least one example of the above embodiments of this disclosure, the multi-row, multi-column block is divided into M groups, the same number as the M processes. Each group forms a block matrix, and each process corresponds to one block matrix. The block matrix corresponding to each process includes a first part and a second part. The M processes perform multiple decomposition and update operations on the target matrix, including: during the execution of a row exchange operation on the first part based on the row exchange information of the current operation, each process performs a matrix update operation on the second part based on the decomposition matrix obtained in the previous operation; during the execution of a matrix update operation on the first part based on the decomposition matrix of the current operation, each process performs a row exchange operation on the second part based on the row exchange information of the current operation.
[0008] For example, in the performance testing method provided in at least one example of the above embodiments of this disclosure, the main thread of each process performs row swapping operations and matrix update operations based on the decomposition matrix and the row swapping information, including: after obtaining the row swapping information of the current operation, the main thread of each process performs a row swapping operation on the first part based on the row swapping information of the current operation to obtain a first swapping matrix for the current operation; after obtaining the decomposition matrix of the current operation, the main thread of each process performs a matrix update operation on the first part based on the decomposition matrix of the current operation and the first swapping matrix of the current operation; the main thread of each process performs a row swapping operation on the second part based on the row swapping information of the current operation to obtain a second swapping matrix for the current operation; and the main thread of each process performs a matrix update operation on the second part based on the decomposition matrix of the current operation and the second swapping matrix of the current operation.
[0009] For example, in a performance testing method provided by at least one example of the above embodiments of this disclosure, each process, while performing a row swap operation for the first part based on the row swap information of the current operation, performs a matrix update operation for the second part based on the decomposition matrix obtained from the previous operation. This includes: the main thread of each process, while performing a row swap operation for the first part based on the row swap information of the current operation, performs a matrix update operation for the second part based on the decomposition matrix and the second swap matrix obtained from the previous operation. During the matrix update operation for the first part based on the decomposition matrix of the current operation, each process performs a row swap operation for the second part based on the row swap information of the current operation. This includes: the main thread of each process, while performing a matrix update operation for the first part based on the decomposition matrix and the first swap matrix of the current operation, performs a row swap operation for the second part based on the row swap information of the current operation.
[0010] For example, in the performance testing method provided in at least one example of the above embodiments of this disclosure, the multiple processes perform multiple decomposition and update operations on the target matrix, which further includes: during the matrix update operation on the second part based on the decomposition matrix of the current operation and the second exchange matrix of the current operation, each process, after obtaining the row exchange information of the next operation of the current operation, performs a row exchange operation on the first part based on the row exchange information of the next operation.
[0011] For example, in the performance testing method provided in at least one example of the above embodiments of this disclosure, for each process corresponding to a block matrix, the first part is at least one column in the block matrix, and the second part is the remaining columns in the block matrix other than the at least one column.
[0012] For example, in the performance testing method provided in at least one example of the above embodiments of this disclosure, the M process arrays are arranged. The broadcast sub-threads of the M processes perform inter-process broadcast operations for the decomposition matrix and the row exchange information, including: multiple processes located in the same process row broadcasting the row exchange information and the decomposition matrix in the process row in a point-to-point manner.
[0013] For example, in the performance testing method provided by at least one example of the above embodiments of this disclosure, the multi-row, multi-column block includes P diagonal blocks located on the diagonal of the target matrix; the M processes perform multiple decomposition and update operations on the target matrix, including: taking the P diagonal blocks as P target blocks respectively, and the M processes sequentially performing the decomposition and update operations P times based on the P target blocks; where P is an integer greater than 1.
[0014] For example, in the performance testing method provided in at least one example of the above embodiments of this disclosure, the target matrix is stored in row-based storage format.
[0015] At least one embodiment of this disclosure provides a performance testing apparatus for a computer system, wherein the computer system runs M processes, each process including a main thread and a broadcast sub-thread. The apparatus includes a decomposition and update module, configured to cause the M processes to perform multiple decomposition and update operations on a target matrix, wherein the target matrix includes multi-row, multi-column blocks arranged in an array, the multi-row, multi-column blocks corresponding to the M processes. The decomposition and update module includes a determination sub-module, a decomposition sub-module, a broadcast sub-module, and an update sub-module. The determination sub-module is configured to perform multiple decomposition and update operations on the target matrix. The target block is determined from the multi-row, multi-column blocks of the matrix; the decomposition submodule is configured to enable the main threads of N processes corresponding to the column where the target block is located to perform the decomposition operation on the target block, and obtain the decomposition matrix and row exchange information for this operation; the broadcast submodule is configured to enable the broadcast subthreads of the M processes to perform inter-process broadcast operations on the decomposition matrix and the row exchange information; the update submodule is configured to enable the main thread of each process to perform row exchange operation and matrix update operation based on the decomposition matrix and the row exchange information; where M is an integer greater than 1 and N is a positive integer less than M.
[0016] For example, in the performance testing apparatus provided in at least one example of the above embodiments of this disclosure, the broadcast submodule is configured to: cause the broadcast subthreads of the N processes to broadcast the row swap information of the current operation to the broadcast subthreads of other processes among the M processes, so that the other processes start to execute the row swap operation; after broadcasting the row swap information of the current operation, cause the broadcast subthreads of the N processes to broadcast the decomposition matrix of the current operation to the broadcast subthreads of the other processes.
[0017] For example, in a performance testing apparatus provided in at least one example of the above embodiments of this disclosure, the update submodule is configured to: in response to obtaining row swap information of the current operation, cause the main thread of each process to start performing a row swap operation on the target matrix based on the row swap information of the current operation to obtain the swap matrix of the current operation; after obtaining the decomposition matrix of the current operation, cause the main thread of each process to perform a matrix update operation on the target matrix based on the decomposition matrix of the current operation and the swap matrix of the current operation.
[0018] At least one embodiment of this disclosure provides a computer system including M process modules. The M process modules are configured to perform multiple decomposition and update operations on a target matrix. The target matrix comprises multi-row, multi-column blocks arranged in an array. Each multi-row, multi-column block corresponds to one of the M processes. Each process module includes a main thread unit and a broadcast sub-thread unit. In each decomposition and update operation, the main thread units of the N process modules corresponding to the columns of the target block are configured to perform a decomposition operation on the target block to obtain the decomposition matrix and row exchange information for this operation. The broadcast sub-thread units of the M process modules are configured to perform inter-process broadcast operations on the decomposition matrix and the row exchange information. The main thread unit of each process module is configured to perform row exchange operations and matrix update operations based on the decomposition matrix and the row exchange information. Here, M is an integer greater than 1, and N is a positive integer less than M.
[0019] At least one embodiment of this disclosure provides an electronic device, including a processor; a memory storing one or more computer program modules; wherein the one or more computer program modules are configured to be executed by the processor to implement the performance testing method provided in any embodiment of this disclosure.
[0020] At least one embodiment of this disclosure provides a computer-readable storage medium storing non-transitory computer-readable instructions, which, when executed by a computer, can implement the performance testing method provided in any embodiment of this disclosure. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings of the embodiments will be briefly described below. Obviously, the drawings described below only relate to some embodiments of this disclosure and are not intended to limit this disclosure.
[0022] Figure 1 A schematic diagram of matrix decomposition and updating is shown;
[0023] Figure 2 A schematic diagram illustrating the correspondence between matrix partitioning and processes is shown.
[0024] Figure 3 A schematic diagram of a process execution flow is shown;
[0025] Figure 4 A flowchart of a performance testing method provided by at least one embodiment of the present disclosure is shown;
[0026] Figure 5 A schematic diagram of broadcast line exchange information and broadcast decomposition matrix provided in at least one embodiment of the present disclosure is shown;
[0027] Figure 6 A schematic diagram illustrating the execution flow of the first and second parts provided in at least one embodiment of this disclosure is shown;
[0028] Figure 7 A schematic block diagram of a performance testing apparatus for a computer system provided in at least one embodiment of the present disclosure is shown.
[0029] Figure 8 A schematic block diagram of an electronic device provided in at least one embodiment of the present disclosure is shown;
[0030] Figure 9 A schematic block diagram of another electronic device provided in at least one embodiment of the present disclosure is shown; and
[0031] Figure 10 A schematic diagram of a computer-readable storage medium provided in at least one embodiment of the present disclosure is shown. Detailed Implementation
[0032] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.
[0033] Unless otherwise defined, the technical or scientific terms used in this disclosure shall have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms “first,” “second,” and similar terms used in this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms “an,” “a,” or “the,” and similar terms do not indicate a quantity limitation, but rather indicate the presence of at least one. The terms “including,” “comprising,” or “containing,” and similar terms mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. The terms “connected,” “linked,” or similar terms are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. The terms “upper,” “lower,” “left,” and “right,” etc., are used only to indicate relative positional relationships, and these relative positional relationships may change accordingly when the absolute position of the described objects changes.
[0034] To better understand the embodiments of this disclosure, some terms appearing in the embodiments of this disclosure are explained below:
[0035] Supercomputer: A large-scale, high-performance computer system that uses high-performance network connections.
[0036] HPL: The main standard for ranking the performance of international supercomputers, using Gaussian elimination with column pivoting to solve large-scale dense linear equation systems.
[0037] Heterogeneous: A computer system containing processors with different architectures, specifically referring to a computer system that uses CPU and GPU for collaborative computing in this disclosure.
[0038] Decomposition (pfact): A step in HPL that performs LU decomposition (triangular decomposition) on the NB column matrix to generate L2 / L1 / IPIV data for subsequent broadcasting.
[0039] Broadcast (bcast): A step in HPL used to broadcast L2 / L1 / IPIV data to all processes in the same row. The broadcast algorithm can be custom.
[0040] Row swap: A step in HPL that uses IPIV to swap data between a subset of rows of a process's local matrix.
[0041] Matrix update: A step in HPL that performs dtrsm (triangular matrix solving) and dgemm (matrix multiplication) on a matrix that has undergone row swapping. In this embodiment, it can also be simply referred to as matrix multiplication / update.
[0042] The HPL test program is a major standard for ranking the performance of international supercomputers. Supercomputers are large-scale, high-performance computer systems using high-performance network connections. The HPL test program uses Gaussian elimination with column pivoting to solve large-scale dense linear equation systems. This program places high demands on the computation, memory access, and networking capabilities of supercomputer systems. The main goal of the program is to find the solution to an N-dimensional linear equation system Ax = b. By selecting local column pivots, the N×(N+1) coefficient matrix [A|b] is decomposed into LU decomposition (triangular decomposition) into the following form: [A|b] = [[LU]|y]. Since the transformations made by the lower triangular matrix L are gradually applied to b during the decomposition process, the final solution x of the equation system can be transformed into solving the linear equation system Ux = y with the upper triangular matrix U as the coefficient matrix.
[0043] Figure 1 A schematic diagram of matrix decomposition and updating is shown.
[0044] like Figure 1 As shown, the large matrix A is divided into multiple blocks in units of NB. These blocks are cyclically distributed in both row and column directions throughout the solution process. In the LU decomposition, each step progresses from the top left to the bottom right in units of NB. A certain step may include the following steps:
[0045] a) The processes that possess the current NB column will collaboratively execute the decomposition to obtain U11, L11, L21 and row exchange information IPIV (not shown in the figure);
[0046] b) Each process broadcasts its own line exchange information IPIV, U11, L11, L21 in the line direction;
[0047] c) After receiving the broadcast data, other column processes perform row swapping operations within their respective column ranges based on the row swapping information IPIV.
[0048] d) All processes perform matrix updates: U12 = L11 -1 U12; New TR = Old TR - L21U12 (TR is...) Figure 1 (trailing matrix in the bottom right corner).
[0049] For example, in some embodiments, L1 can be used to represent... Figure 1 The diagonal block (L11+U11) in the diagram is represented by U, where U represents U12 and L2 represents L21.
[0050] Figure 2 A schematic diagram illustrating the correspondence between matrix partitioning and processes is shown.
[0051] like Figure 2As shown, a large matrix A is divided into 8×8 blocks of size NB (blocks A11 to A88), each block being NB×NB in size. For example, there are 4 processes (processes 0 to 3) participating in the solution process. The blocks of size NB are cyclically distributed in the row and column directions among the 4 processes to obtain the correspondence between processes and blocks. Each process corresponds to a block matrix.
[0052] For example, multiple update operations are performed on the large matrix A according to the diagonal blocks A11 to A88 in the order from top left to bottom right. The following example illustrates this. Figure 2 Let's take the diagonal block A11 as an example. The columns (A11~A81) containing diagonal block A11 are distributed in processes 0 and 2. Processes 0 and 2 perform the LU decomposition operation. After the LU decomposition, the lower triangular part of diagonal block A11 is... Figure 1 L11 in the middle, together with other blocks A21 to A81 in the same column as diagonal block A11, forms Figure 1 In the decomposition process, to ensure stability, the diagonal element must be the largest in its column. Therefore, during decomposition, the row containing the diagonal element needs to be swapped with the row containing the largest element in the corresponding column. This results in an IPIV array, which represents the row swap information during the decomposition process. IPIV(i) = j indicates that the i-th row and the j-th row were swapped. The decomposition process does not change the block division of other columns. After the decomposition is complete, both process 0 and process 2 obtain the decomposition matrix L11 / L21 and the row swap information IPIV.
[0053] Since the decomposition matrix L11 / L21 is part of the block matrix corresponding to process 0 and process 2, and is not present in other processes, but is needed in subsequent matrix updates, the decomposition matrix L11 / L21 and the row exchange information IPIV need to be sent to all processes in the row direction. For example, process 0 sends it to process 1, and process 2 sends it to process 3.
[0054] After the broadcast, each process performs a row swap operation. Although A11-A81 have already been swapped in processes 0 and 2, each process has four blocks in the column direction. The blocks to the right of block A11 need to be row swapped. Taking processes 0 and 2 as an example, processes 0 and 2 use the row swap information IPIV to perform the row swap operation, generating a swap matrix U of 1NB rows × 3NB columns. This swap matrix U is the result of swapping the three blocks A13 / A15 / A17 to the right of block A11 using the row swap information IPIV. This swap matrix U, as part of process 0, also needs to be broadcast to process 2 in the same column, because process 2 needs to use it to perform matrix update operations. Similarly, processes 1 and 3 perform row swapping operations using row swapping information IPIV, generating another swapping matrix U of size 1NB rows × 4NB columns. This swapping matrix U is the result of swapping the four blocks A12 / A14 / A16 / A18 in the same row as block A11 using row swapping information IPIV. After obtaining swapping matrix U, process 1 broadcasts it to process 3 in the same column, so that process 3 can use swapping matrix U to perform a matrix update operation. For ease of description, the process of process 0 and process 1 broadcasting the swapping matrix to processes 2 and 3 respectively is combined into the row swapping operation.
[0055] After the row swap operation, each process can perform the matrix update operation. Taking process 0 as an example, it first executes U=L11 on U. -1 *U, then execute on process 0:
[0056] (A33,A35,A37)=(A33,A35,A37)-A31*U;
[0057] (A53,A55,A57)=(A53,A55,A57)-A51*U;
[0058] (A73,A75,A77)=(A73,A75,A77)-A71*U.
[0059] This completes the matrix update operation for diagonal block A11, and the solution for the first NB rows and NB columns of the entire matrix A is finished. The above operation is then performed again for the next diagonal block A22, and so on, completing the entire solution process after 8 iterations.
[0060] Figure 3 A schematic diagram of a process execution flow is shown.
[0061] like Figure 3As shown in the figure, each process determines its own coordinate in the process matrix (S101), and determines whether the process is located in the j-th column (S102), wherein the processes in the j-th column are a column of processes that need to perform a decomposition operation. If the process is located in the j-th column, it performs decomposition calculation with other processes also located in the j-th column (S103), obtains a decomposition matrix and row exchange information, and initiates row-direction broadcasting based on the decomposition matrix and row exchange information (S104). If the process is not located in the j-th column, it waits for the processes in the j-th column to complete the decomposition operation and perform direction broadcasting, then receives the information broadcast in the row direction (S105). After obtaining the decomposition matrix and row exchange information, a column-direction row exchange operation can be performed (S106), and after the row exchange operation, a matrix update operation can be performed (S107). k represents the position of the currently calculated block in the global matrix, NB represents the minimum unit of data division, k += NB represents continuing the decomposition and update operations of the next block, k < N indicates that the decomposition of all blocks has not been completed yet. When the condition k < N is satisfied, the process returns to step S102 to enter the next round of iteration until the LU decomposition process of the entire matrix is completed, and exits the loop (S108). For example, the matrix multiplication operation in the matrix update operation can be invoked to be executed by a GPU (Graphics Processing Unit).
[0062] In the above process, most of the running time is concentrated in the matrix update process. With the rapid improvement of GPU computing performance, the time of matrix operation is significantly reduced, while the improvement of network communication bandwidth is far less than that of computing performance, which makes communication (broadcasting process) become the key factor affecting the HPL performance of high-performance computing systems.
[0063] At least one embodiment of the present disclosure provides a performance testing method for a computer system, a performance testing apparatus for a computer system, a computer system, an electronic device and a computer-readable storage medium. The performance testing method comprises: performing multiple decomposition and update operations on a target matrix by M processes, wherein the target matrix comprises multiple rows and multiple columns of blocks arranged in an array, and the multiple rows and multiple columns of blocks have a corresponding relationship with the M processes; each decomposition and update operation comprises: determining a target block from the multiple rows and multiple columns of blocks of the target matrix; having main threads of N processes corresponding to the column where the target block is located among the M processes perform a decomposition operation on the target block to obtain a decomposition matrix and row exchange information of the current operation; having broadcast subthreads of the M processes perform an inter-process broadcast operation on the decomposition matrix and the row exchange information; having the main thread of each process perform a row exchange operation and a matrix update operation based on the decomposition matrix and the row exchange information; wherein M is an integer greater than 1, and N is a positive integer less than M.
[0064] This performance testing method starts a main thread and a broadcast sub-thread in each process participating in the computation. The broadcast sub-thread performs inter-process broadcast operations, while the main thread performs operations such as decomposition and row swapping. Since the broadcast operations are executed by a separate sub-thread, the broadcast operations can be executed in parallel with other operations (such as row swapping), reducing the sequential relationship between the broadcast operations and other operations in the process flow, reducing the time wasted waiting for the broadcast operations to be executed, and reducing the impact of communication on the performance of the high-performance computing system (HPL).
[0065] The performance testing method provided in at least one embodiment of this disclosure is used in a computer system, which runs M processes, each process including a main thread and a broadcast sub-thread.
[0066] For example, in some embodiments, the computer system may be a computer cluster consisting of multiple computers, such as a high-performance computer cluster, where M processes are distributed and run across multiple computers. In other embodiments, the computer system may be a single computer.
[0067] Figure 4 A flowchart of a performance testing method provided by at least one embodiment of the present disclosure is shown.
[0068] like Figure 4 As shown, the method includes: multiple decomposition and update operations on a target matrix performed by M processes. The target matrix comprises multi-row, multi-column blocks arranged in an array, and each multi-row, multi-column block corresponds to one of the M processes. Each decomposition and update operation includes steps S210 to S240.
[0069] Step S210: Determine the target block from the multi-row, multi-column blocks of the target matrix.
[0070] Step S220: The main threads of the N processes corresponding to the column where the target block is located among the M processes execute the decomposition operation for the target block, and obtain the decomposition matrix and row exchange information for this operation.
[0071] Step S230: The broadcast sub-threads of M processes perform inter-process broadcast operations for decomposing the matrix and exchanging row information.
[0072] Step S240: The main thread of each process performs row swapping and matrix update operations based on the decomposed matrix and row swapping information.
[0073] For example, M is an integer greater than 1, and N is a positive integer less than M.
[0074] For example, in the following embodiments, Figure 2The matrix A shown is an example of the target matrix, and processes 0 to 3 are used as examples of M processes. Matrix A includes 64 blocks arranged in an array, and these 34 blocks are examples of the aforementioned multi-row, multi-column blocks. However, this disclosure is not limited thereto. In other embodiments, the target matrix may have other sizes and / or other numbers of blocks, and more or fewer processes may participate in the performance testing process.
[0075] For example, a multi-row, multi-column block is divided into M groups, the same number as the M processes. Each group forms a block matrix, and each process corresponds to a block matrix. For example, the 64 blocks of matrix A can be divided into 4 groups, each corresponding to processes 0 to 3, with each group forming a block matrix. For example, multiple blocks of matrix A can be cyclically distributed in both rows and columns among processes 0 to 3. Processes 0 to 3 can also be arranged in matrix form, such as forming a 2×2 process matrix. Multiple blocks in the first row of matrix A can be cyclically distributed among processes 0 and 1 in the first row of the process matrix; multiple blocks in the second row of matrix A can be cyclically distributed among processes 2 and 3 in the second row of the process matrix; multiple blocks in the third row of matrix A can be cyclically distributed among processes 0 and 1 in the first row of the process matrix; multiple blocks in the fourth row of matrix A can be cyclically distributed among processes 2 and 3 in the second row of the process matrix, and so on, until all blocks of matrix A are allocated.
[0076] For example, the multi-row, multi-column blocks of the target matrix include P diagonal blocks located on the diagonal of the target matrix. Here, P is an integer greater than 1. For example, the eight diagonal blocks A11 to A88 in matrix A can be examples of these P diagonal blocks.
[0077] For example, multiple decomposition and update operations on a target matrix performed by M processes may include: taking P diagonal blocks as P target blocks, and having M processes perform P decomposition and update operations sequentially based on the P target blocks.
[0078] For example, in step S210, for the first decomposition and update operation, the first diagonal block A11 of matrix A can be determined as the target block, and steps S210 to S240 are performed on diagonal block A11; for the second decomposition and update operation, the second diagonal block A22 can be determined as the target block, and steps S210 to S240 are performed on diagonal block A22; and so on, until the last decomposition and update operation, where the last diagonal block A88 of matrix A can be determined as the target block, and steps S210 to S240 are performed on diagonal block A88, completing the entire decomposition process. For example, in the following embodiments, diagonal block A11 is used as the target block for illustration. In some of the following embodiments, one decomposition and update operation can also be referred to as one iteration.
[0079] For example, each of the M processes running in a computer system can include at least one main thread and one broadcast sub-thread for broadcast operations. Figure 2 The process shown includes a main thread 00 and a broadcast sub-thread 01; process 1 includes a main thread 10 and a broadcast sub-thread 11; process 2 includes a main thread 20 and a broadcast sub-thread 21; and process 3 includes a main thread 30 and a broadcast sub-thread 31. For example, the main thread can be used to perform decomposition operations, row swapping operations, and matrix update operations, while the broadcast sub-threads can be used to perform broadcast operations. In some embodiments, the main thread can also be divided into multiple sub-threads as needed to perform decomposition operations, row swapping operations, and matrix update operations respectively.
[0080] For example, in step S220, determining the N processes corresponding to the target block can be achieved by using multiple processes distributed in the column where the target block A11 is located as the N processes. Figure 2 As shown, when the target block is a diagonal block A11, target block A11 is located in the first column of matrix A. The first column of matrix A is distributed in processes 0 and 2. Therefore, processes 0 and 2 can be N processes corresponding to target block A11. After determining the N processes corresponding to the target block, the main thread of these N processes can be used to perform the decomposition operation on the target block, obtaining the decomposition matrix L11 / L21 and the row exchange information IPIV for this iteration.
[0081] For example, in step S230, the decomposition matrix L11 / L21 and row exchange information IPIV are broadcast between processes in each row of the process matrix. The broadcast operation can be performed by the broadcast sub-threads of each process. For example, after the main threads of process 0 and process 2 complete the decomposition operation on the target block A11, the main threads 00 of process 0 and 20 of process 2 obtain the decomposition matrix L11 / L21 and row exchange information IPIV. The main thread 00 of process 0 can send the decomposition matrix L11 / L21 and row exchange information IPIV to the broadcast sub-thread 01 of process 0, so that the broadcast sub-thread 01 of process 0 can send the decomposition matrix L11 / L21 and row exchange information IPIV to the broadcast sub-thread 11 of process 1 in the same row. Similarly, the main thread 20 of process 2 can send the decomposition matrix L11 / L21 and row exchange information IPIV to the broadcast sub-thread 21 of process 2, so that the broadcast sub-thread 21 of process 2 can send the decomposition matrix L11 / L21 and row exchange information IPIV to the broadcast sub-thread 21 of process 2 in the same row. After receiving the decomposition matrix L11 / L21 and the row exchange information IPIV, the broadcast sub-thread 11 of process 1 and the broadcast sub-thread 31 of process 3 can send the decomposition matrix L11 / L21 and the row exchange information IPIV to their respective main threads, so that the main threads can perform subsequent operations based on the decomposition matrix L11 / L21 and the row exchange information IPIV.
[0082] For example, in step S240, processes 0 to 3 perform row swapping and matrix update operations based on the decomposition matrix L11 / L21 and the row swapping information IPIV.
[0083] For example, in one iteration, steps S230 and S240 can be executed sequentially, that is, step S230 is executed first and then step S240 begins. Alternatively, in one iteration, steps S230 and S240 can be executed intermittently or at least partially synchronously. For example, after broadcasting part of the information in the decomposition matrix and row exchange information in step S230, the process in step S240 can perform a row exchange operation based on the received part of the information. During or after the row exchange operation, step S230 can broadcast another part of the information in the decomposition matrix and row exchange information, and step S240 can perform a matrix update operation based on the received other part of the information.
[0084] In the performance testing method of this embodiment, a main thread and a broadcast sub-thread are started in each process participating in the computation. The broadcast sub-thread is used to perform inter-process broadcast operations, while the main thread performs operations such as decomposition and row swapping. Since the broadcast operation is executed by a separate sub-thread, the broadcast operation can be executed in parallel with other operations, reducing the sequential relationship between the broadcast operation and other operations in the process, reducing the time wasted waiting for the broadcast operation to be executed, and reducing the impact of communication on the performance of the high-performance computing system (HPL).
[0085] For example, in some embodiments, some operations in adjacent iterations can be executed synchronously. Taking the first and second iterations as examples, when the decomposition operation of the first iteration is completed and the broadcast operation begins, the decomposition operation of the second iteration can begin. When the broadcast operation of the first iteration is completed and the row swapping operation begins, the broadcast operation of the second iteration can begin, achieving parallelism between the row swapping operation of the first iteration and the broadcast operation of the second iteration. If the broadcast operation and the row swapping operation in each process are executed in the same thread, then the broadcast information of the second iteration can only be detected intermittently in the middle of the row swapping operation of the first iteration. There is still a sequential relationship between the row swapping operation and the broadcast operation, and the parallelism between the row swapping operation and the broadcast operation cannot be achieved, affecting the synchronization characteristics of the row swapping process and slowing down the overall row swapping process. If the embodiments of this disclosure are adopted, the broadcast operation is executed in a separate sub-thread. During the execution of the row swapping operation of the first iteration in the main thread, the broadcast sub-thread can execute the broadcast operation of the second iteration, achieving parallelism between the row swapping operation of the first iteration and the broadcast operation of the second iteration.
[0086] For example, for an iteration, step S230 can be executed before step S240 begins. That is, the decomposition matrix and row swap information for this iteration are broadcast before the row swap operation begins. Each process's row swap operation only uses the relatively small row swap information IIV and does not use the decomposition matrix. The decomposition matrix is used for matrix update operations. In other words, each process can begin the row swap operation immediately after receiving the row swap information IIV. Waiting for both the decomposition matrix and the row swap information to be received before starting the row swap operation would delay the row swap operation and slow down the overall process.
[0087] Therefore, in some embodiments of this disclosure, step S230 is further divided into two operations: broadcasting the matrix decomposition and broadcasting row exchange information. Step S240 includes two operations: performing the row exchange operation and performing the matrix update operation. In one iteration, based on the use of a separate broadcast sub-thread to perform the broadcast operation, some operations in steps S230 and S240 can be executed synchronously (i.e., in parallel).
[0088] For example, in step S230, the broadcast sub-threads of N processes broadcast the row swap information of this operation to the broadcast sub-threads of other processes in M processes, so that the other processes can start executing the row swap operation; after broadcasting the row swap information of this operation, the broadcast sub-threads of N processes broadcast the decomposition matrix of this operation to the broadcast sub-threads of other processes.
[0089] For example, in step S240, in response to obtaining the row swap information of this operation, the main thread of each process starts to perform a row swap operation on the target matrix based on the row swap information of this operation to obtain the swap matrix of this operation; after obtaining the decomposition matrix of this operation, the main thread of each process performs a matrix update operation on the target matrix based on the decomposition matrix and the swap matrix of this operation.
[0090] Figure 5 A schematic diagram of broadcast line exchange information and broadcast decomposition matrix provided in at least one embodiment of the present disclosure is shown.
[0091] like Figure 5 As shown, taking the broadcast from process 0 to process 1 as an example, after the main thread 00 in process 0 completes the decomposition operation, it obtains the decomposition matrix L11 / L21 and the row swapping information IPIV. Process 0 first uses its broadcast sub-thread 01 to send the row swapping information IPIV to the broadcast sub-thread 11 of process 1. After receiving the row swapping information IPIV, sub-thread 11 sends it to the main thread 10. The main thread 10 can then begin performing the row swapping operation based on the row swapping information IPIV. During the row swapping operation performed by the main thread 10, sub-thread 01 can send the decomposition matrix L11 / L21 to sub-thread 11, and sub-thread 11 sends the decomposition matrix L11 / L21 to the main thread 10. After the main thread 10 completes the row swapping operation, obtains the swapping matrix U, and acquires the decomposition matrix L11 / L21, it can begin performing the matrix update operation based on the swapping matrix U and the decomposition matrix L11 / L21.
[0092] For example, each of the N processes sends the row exchange information and the decomposition matrix to the other processes in the same row twice, sending the row exchange information first and then the decomposition matrix. This allows each process to start executing the row exchange operation earlier, speeding up the process. Furthermore, using a separate broadcast sub-thread to perform the broadcast operation allows the row exchange operation and the broadcasting of the decomposition matrix in each iteration to be executed in parallel, maximizing the computational hiding of the communication process and further reducing the impact of communication on the overall process. The message length of the row exchange information IPIV is much smaller than the total length of the row exchange information and the decomposition matrix (the ratio of IPIV message length to the total length is, for example, less than 1 / 1000). Therefore, broadcasting the row exchange information first allows the row exchange operation to start much earlier, achieving complete parallelism with the broadcasting of the decomposition matrix, maximizing network bandwidth utilization, and increasing the computational hiding rate of communication.
[0093] For example, the block matrix corresponding to each process may include a first part and a second part. For example, the first part is at least one column in the block matrix, and the second part is the remaining columns in the block matrix other than the at least one column.
[0094] For example, in some embodiments, the number of columns in the first part and the second part can be the same. In the case where the block matrix includes 2K columns (K is a positive integer), the first part can be the first K columns of those 2K columns, and the second part can be the last K columns of those 2K columns. Figure 2 As shown, the first part of the block matrix corresponding to process 0 can be the two columns containing blocks A11 and A13, and the second part can be the two columns containing blocks A15 and A17. The first part of the block matrix corresponding to process 1 can be the two columns containing blocks A12 and A14, and the second part can be the two columns containing blocks A16 and A18. In some other embodiments, the number of columns in the first part and the second part can be different.
[0095] For example, each process can perform a row swap operation for the first part (i.e., the current round of row swap operation) based on the row swap information of this operation, and perform a matrix update operation for the second part (i.e., the previous round of matrix update operation) based on the decomposition matrix obtained in the previous operation; each process can also perform a row swap operation for the second part (i.e., the current round of row swap operation) based on the row swap information of this operation, while performing a matrix update operation for the first part (i.e., the current round of matrix update operation) based on the decomposition matrix of this operation.
[0096] For example, during the matrix update operation for the second part based on the decomposed matrix of this operation (i.e., the current round of matrix update operation), each process can perform a row swap operation for the first part (i.e., the next round of row swap operation) based on the row swap information of the next operation after obtaining the row swap information of the next operation.
[0097] It should be noted that, in the embodiments of this disclosure, "this operation" refers to the current decomposition and update operation, which in some embodiments is also referred to as the current iteration or the current round of iteration. Correspondingly, "the previous operation" refers to the previous decomposition and update operation, which in some embodiments is also referred to as the previous iteration or the previous round of iteration. "The next operation" refers to the next decomposition and update operation, which in some embodiments is also referred to as the next iteration or the next round of iteration.
[0098] For example, taking process 0 as an example, during the first round of matrix update operations for the first part (e.g., the two columns containing blocks A11 and A13) (at which point the first round of row swap operations for the first part has already been completed), if a second round of row swap information is received, process 0 can begin executing the second round of row swap operations for the second part (e.g., the two columns containing blocks A15 and A17). After the second round of row swap operations for the second part is completed, the second round of matrix update operations for the second part can begin. During the execution of the second round of matrix update operations for the second part, if a third round of row swap information is received, process 0 can begin executing the third round of row swap operations for the first part. The same logic applies to other processes (processes 1-3). By dividing the block matrix corresponding to each process into a first part and a second part, and having each part execute the matrix update operations of the previous round and the row swap operations of the next round in adjacent rounds respectively, parallelism between the matrix update operations of the previous round and the row swap operations of the next round in adjacent rounds is achieved, enabling cross-iteration operation synchronization and resulting in higher processing efficiency.
[0099] For example, for each process, after obtaining the row swap information for this operation, the main thread of each process performs a row swap operation for the first part based on the row swap information for this operation to obtain the first swap matrix for this operation; after obtaining the decomposition matrix for this operation, the main thread of each process can perform a matrix update operation for the first part based on the decomposition matrix and the first swap matrix for this operation; the main thread of each process performs a row swap operation for the second part based on the row swap information for this operation to obtain the second swap matrix for this operation; the main thread of each process performs a matrix update operation for the second part based on the decomposition matrix and the second swap matrix for this operation.
[0100] For example, each process can perform row swapping operations on the first part and the second part respectively, to obtain a first swapping matrix (denoted by symbol U1) corresponding to the first part and a second swapping matrix (denoted by symbol U2) corresponding to the second part.
[0101] For example, a matrix update operation may include performing a first matrix update operation (triangular matrix solving, dtrsm) on the commutation matrix U and a second matrix update operation (matrix multiplication, dgemm) using the decomposition matrix L2 and the commutation matrix U. Each process may perform the matrix update operation for the first part and the second part separately. The matrix update operation for the first part may include performing a first matrix update operation on the first commutation matrix U1 and performing a second matrix update operation using the decomposition matrix L2 and the first commutation matrix U1. The matrix update operation for the second part may include performing a first matrix update operation on the second commutation matrix U2 and performing a second matrix update operation using the decomposition matrix L2 and the second commutation matrix U2.
[0102] For example, during the row swap operation for the first part based on the row swap information of this operation, the main thread of each process can perform a matrix update operation for the second part based on the decomposition matrix obtained in the previous operation and the second swap matrix U2 obtained in the previous operation.
[0103] For example, the main thread of each process can perform a matrix update operation for the first part based on the decomposed matrix and the first swap matrix U1 of this operation, and perform a row swap operation for the second part based on the row swap information of this operation.
[0104] For example, during the matrix update operation for the second part based on the decomposed matrix and the second swap matrix U2 of this operation, the main thread of each process can, after obtaining the row swap information of the next operation, perform a row swap operation for the first part based on the row swap information of the next operation.
[0105] Figure 6 A schematic diagram of the execution flow of the first and second parts provided in at least one embodiment of this disclosure is shown.
[0106] like Figure 6As shown, row swapping operations can be performed by the CPU (Central Processing Unit), while matrix update operations can be performed by the GPU. For each process, after receiving the first round of broadcast row swapping information, the CPU can first execute the first round of row swapping operations for the first part (step S301), and simultaneously receive the first round of broadcast decomposition matrix (after receiving the decomposition matrix, it copies it to the GPU). After the GPU determines that the first round of row swapping operations for the first part is complete and receives the first round of decomposition matrix, the GPU can begin executing the first round of matrix update operations for the first part (step S401). During the GPU's execution of the first round of matrix update operations for the first part, the CPU can execute the first round of row swapping operations for the second part (step S302). After the GPU determines that the first round of row swapping operations for the second part is complete, it can execute the first round of matrix update operations for the second part (step S402). While the GPU is executing the first round of matrix update operations for the second part, after the CPU determines that the first round of matrix update operations for the first part is complete and receives the second round of broadcast row swapping information, the CPU can execute the second round of row swapping operations for the first part (step S303) and simultaneously receive the second round of broadcast decomposition matrix. After the GPU determines that the second round of row swapping operations for the first part is complete and receives the second round of decomposition matrix, the GPU can begin executing the second round of matrix update operations for the first part (step S403). While the GPU is executing the first round of matrix update operations for the first part, the CPU can execute the second round of row swapping operations for the second part (step S304). This process continues until the final round of matrix update operations for both the first and second parts is completed.
[0107] For example, when entering the first update phase, the row swapping operation of the first part has been completed, and the matrix update operation of the first part and the row swapping operation of the second part can be started directly, realizing the parallelism of row swapping and matrix update operations to the greatest extent.
[0108] For example, in some embodiments of this disclosure, the matrix update operation and the row swapping operation are executed in the same thread, that is, multiple threads are not used to execute the row swapping operation and the matrix update operation separately. Since the matrix update operation is a relatively simple function that calls the GPU interface and can be executed asynchronously, the call only needs to be initiated where the dependency relationship is satisfied, and no additional thread is required. The main thread performs the row swapping operation after initiating the asynchronous matrix multiplication. These two parts of the operation do not need to be implemented using additional threads. Figure 6 The lines in the diagram represent data dependencies.
[0109] For example, M processes are arranged in an array, and multiple processes located in the same process line exchange information and decompose matrices in a point-to-point manner through broadcasting within the process line.
[0110] For example, when each process row includes multiple processes, the processes participating in the decomposition operation send the obtained decomposition matrix and row exchange information to other processes in the same row in a point-to-point manner. The decomposition matrix and row exchange information can be sent in two separate transmissions. Point-to-point transmission can be understood as one process sending data to another; that is, data is sent one-to-one between multiple processes in the same process row. For example, the first process row includes process 0, process 1, and process 2. After process 0 performs the decomposition operation and obtains the decomposition matrix and row exchange information, it can first send the row exchange information to process 1. Process 1 receives the row exchange information, stores it, and can forward it to process 2. Alternatively, process 0 can first send the row exchange information to process 1 and then send it to process 2. The decomposition matrix is also sent point-to-point within a process row. In this mode, broadcasting can be implemented asynchronously. The running states of processes in different columns are different. Through reasonable process arrangement, bandwidth contention between different processes on the same node can be avoided to some extent.
[0111] For example, if the target matrix is stored in column storage format (matrix data is stored continuously in the column direction), then the data in a row of the target matrix is discrete. Row swapping requires swapping two rows, and the swapped data is not continuous. This will cause row swapping (the process of copying multiple rows of data) to involve a large number of discrete memory accesses. Even if GPU is used for acceleration, it will result in a large waste of memory access bandwidth and low memory access efficiency.
[0112] For example, in at least one embodiment of this disclosure, the target matrix is stored in a row-oriented format, that is, the matrix data of the target matrix is stored continuously in the row direction. The matrix is stored in a row-oriented format, so the data is continuous when the rows are swapped. Using the GPU to accelerate data copying can achieve very good results and improve the bandwidth utilization of row swapping.
[0113] For example, when there are many processes in the row direction, the broadcast delay is one of the main factors affecting HPL performance. The embodiments of this disclosure provide conditions for the parallelism of broadcast operations with other operations by using a separate sub-thread to perform broadcast operations, which can significantly improve broadcast reception efficiency. The test data in Table 1 below verifies this.
[0114] Table 1
[0115]
[0116] For example, Table 1 shows the test results of the solution provided by at least one embodiment of this disclosure and a certain prior art on 1-256 nodes on a high-performance heterogeneous cluster. It can be seen that the solution provided by at least one embodiment of this disclosure has better scalability than the prior art. The efficiency of the two algorithms is basically the same on a single node. However, as the number of nodes increases, the efficiency of the solution provided by at least one embodiment of this disclosure decreases very little, while the efficiency of the prior art decreases significantly with the increase of the number of nodes.
[0117] For example, at least one embodiment of this disclosure uses a separate sub-thread to perform the broadcast operation, enabling parallel execution of broadcasting and other operations such as row swapping. For instance, the broadcast operation in the next iteration and the row swapping operation in the current iteration may run concurrently. However, without using a broadcast sub-thread, the HPL implementation intersperses broadcast detection during row swapping, which affects the synchronization characteristics of the row swapping process and slows down the overall row swapping process. At least one embodiment of this disclosure executes the broadcast separately using a sub-thread, avoiding dependencies in the communication process and improving the response speed of broadcast transmission and reception.
[0118] For example, at least one embodiment of this disclosure splits the broadcast data for transmission. Since the row swapping process only requires the IPIV data in the broadcast, this allows the row swapping operation to begin execution as early as possible. Furthermore, after obtaining the IPIV data for the next iteration, broadcasting the IPIV data first allows the next round of row swapping operations to begin execution as early as possible, thus also facilitating the parallel execution of the current matrix update operation and the next round of row swapping operations.
[0119] For example, at least one embodiment of this disclosure divides the block matrix corresponding to each process into two blocks, which is equivalent to dividing the target matrix into two blocks. Compared with the scheme of dividing the target matrix into more than two blocks, the scale of the single matrix update operation and row swap operation of at least one embodiment of this disclosure is increased. On the one hand, it reduces the number of GPU function calls, and on the other hand, it improves the matrix update efficiency and row swap bandwidth utilization.
[0120] For example, at least one embodiment of this disclosure uses a row-first storage matrix, which can improve the data continuity of the row switching process, further reduce the time required for row switching, and significantly improve bandwidth utilization.
[0121] For example, the thread communication technology disclosed in at least one embodiment is not only applicable to HPL, but also to HPC (High Performance Computing) applications with multi-directional communication. Matrix storage format optimization can also be used for HPC application optimization.
[0122] Figure 7A schematic block diagram of a performance testing apparatus 500 for a computer system provided in at least one embodiment of the present disclosure is shown.
[0123] For example, such as Figure 7 As shown, the performance testing apparatus 500 includes a decomposition and update module, which comprises a determination submodule 510, a decomposition submodule 520, a broadcast submodule 530, and an update submodule 540. These components are interconnected via a bus system and / or other forms of connection mechanisms (not shown). For example, these modules can be implemented as hardware (e.g., circuit) modules, software modules, or any combination of both, as is the case in the following embodiments, and will not be repeated here. For example, these units can be implemented using a central processing unit (CPU), a graphics processing unit (GPU), a tensor processor (TPU), a field-programmable gate array (FPGA), or other forms of processing units with data processing capabilities and / or instruction execution capabilities, along with corresponding computer instructions. It should be noted that... Figure 7 The components and structure of the performance testing apparatus 700 shown are merely exemplary and not limiting; the performance testing apparatus 700 may also have other components and structures as needed.
[0124] The decomposition and update module is configured to enable M processes to perform multiple decomposition and update operations on the target matrix, where the target matrix consists of multi-row, multi-column blocks arranged in an array, and the multi-row, multi-column blocks correspond to the M processes.
[0125] The determination submodule 510 is configured to determine the target block from the multi-row, multi-column blocks of the target matrix. For example, the determination submodule 510 can execute... Figure 4 Step S210 is described.
[0126] The decomposition submodule 520 is configured to have the main threads of N processes out of M processes, corresponding to the columns containing the target block, execute the decomposition operation on the target block, obtaining the decomposition matrix and row swapping information for this operation. For example, the decomposition submodule 520 can execute... Figure 4 Step S220 is described.
[0127] Broadcast submodule 520 is configured to enable broadcast subthreads of M processes to perform inter-process broadcast operations for matrix decomposition and row-level information exchange. For example, broadcast submodule 520 can execute... Figure 4 Step S230 is described.
[0128] The update submodule 540 is configured so that the main thread of each process performs row swapping and matrix update operations based on the decomposed matrix and row swapping information. For example, the update submodule 540 can execute... Figure 4 Step S240 is described.
[0129] For example, M is an integer greater than 1, and N is a positive integer less than M.
[0130] For example, the determining submodule 510, the decomposition submodule 520, the broadcast submodule 530, and the updating submodule 540 can be hardware, software, firmware, or any feasible combination thereof. For example, the determining submodule 510, the decomposition submodule 520, the broadcast submodule 530, and the updating submodule 540 can be dedicated or general-purpose circuits, chips, or devices, or they can be a combination of a processor and memory. The embodiments of this disclosure do not limit the specific implementation of the above-mentioned units.
[0131] For example, the determining submodule 510, decomposition submodule 520, broadcast submodule 530, and updating submodule 540 may include code and programs stored in memory; the processor may execute the code and programs to implement some or all of the functions of the determining submodule 510, decomposition submodule 520, broadcast submodule 530, and updating submodule 540 as described above. For example, the determining submodule 510, decomposition submodule 520, broadcast submodule 530, and updating submodule 540 may be dedicated hardware devices used to implement some or all of the functions of the determining submodule 510, decomposition submodule 520, broadcast submodule 530, and updating submodule 540 as described above. For example, the determining submodule 510, decomposition submodule 520, broadcast submodule 530, and updating submodule 540 may be a circuit board or a combination of multiple circuit boards used to implement the functions described above. In embodiments of this disclosure, the circuit board or combination of multiple circuit boards may include: (1) one or more processors; (2) one or more non-temporary memories connected to the processor; and (3) processor-executable firmware stored in memory.
[0132] It should be noted that in the embodiments of this disclosure, each unit of the performance testing device 500 corresponds to each step of the aforementioned performance testing method. For the specific functions of the performance testing device 500, please refer to the relevant description of the performance testing method, which will not be repeated here. Figure 7 The components and structure of the performance testing apparatus 500 shown are exemplary and not limiting. The performance testing apparatus 500 may include other components and structures as needed. The performance testing apparatus 500 may include more or fewer circuits or units, and the connection relationships between the various circuits or units are not limited and can be determined according to actual needs. The specific configuration of each circuit or unit is not limited; it may be constructed from analog devices, digital chips, or other suitable methods according to circuit principles.
[0133] For example, in the performance testing apparatus provided in at least one example of the above embodiments of this disclosure, the broadcast submodule is further configured to: cause the broadcast subthreads of N processes to broadcast the row swap information of this operation to the broadcast subthreads of other processes among M processes, so that the other processes can start executing the row swap operation; after broadcasting the row swap information of this operation, cause the broadcast subthreads of N processes to broadcast the decomposition matrix of this operation to the broadcast subthreads of other processes.
[0134] For example, in the performance testing apparatus provided in at least one example of the above embodiments of this disclosure, the update submodule is further configured to: in response to obtaining the row swap information of the current operation, cause the main thread of each process to start performing a row swap operation on the target matrix based on the row swap information of the current operation, so as to obtain the swap matrix of the current operation; after obtaining the decomposition matrix of the current operation, cause the main thread of each process to perform a matrix update operation on the target matrix based on the decomposition matrix and the swap matrix of the current operation.
[0135] For example, in the performance testing apparatus provided in at least one example of the embodiments of this disclosure, the multi-row, multi-column block is divided into M groups, the same number as the M processes. Each group forms a block matrix, and each process corresponds to a block matrix. The block matrix corresponding to each process includes a first part and a second part. The decomposition and update module is further configured to: during the execution of a row exchange operation for the first part based on the row exchange information of the current operation, each process performs a matrix update operation for the second part based on the decomposition matrix obtained in the previous operation; and during the execution of a matrix update operation for the first part based on the decomposition matrix of the current operation, each process performs a row exchange operation for the second part based on the row exchange information of the current operation.
[0136] For example, in the performance testing apparatus provided in at least one example of the above embodiments of this disclosure, the update submodule is further configured to: after obtaining the row swap information of the current operation, cause the main thread of each process to perform a row swap operation for the first part based on the row swap information of the current operation to obtain the first swap matrix of the current operation; after obtaining the decomposition matrix of the current operation, cause the main thread of each process to perform a matrix update operation for the first part based on the decomposition matrix of the current operation and the first swap matrix of the current operation; cause the main thread of each process to perform a row swap operation for the second part based on the row swap information of the current operation to obtain the second swap matrix of the current operation; and cause the main thread of each process to perform a matrix update operation for the second part based on the decomposition matrix of the current operation and the second swap matrix of the current operation.
[0137] At least one embodiment of this disclosure also provides a computer system including M process modules, which are configured to perform multiple decomposition and update operations on a target matrix, wherein the target matrix includes multi-row and multi-column blocks arranged in an array, and the multi-row and multi-column blocks correspond to the M processes. Each process module includes a main thread unit and a broadcast sub-thread unit.
[0138] For example, in each decomposition and update operation, the main thread units of the N process modules corresponding to the column where the target block is located in the M process modules are configured to: perform the decomposition operation for the target block and obtain the decomposition matrix and row exchange information for this operation; the broadcast sub-thread units of the M process modules are configured to: perform inter-process broadcast operations for the decomposition matrix and row exchange information; the main thread unit of each process module is configured to: perform row exchange operations and matrix update operations based on the decomposition matrix and row exchange information; where M is an integer greater than 1 and N is a positive integer less than M.
[0139] For example, the computer system can also refer to the heat protection measures for the performance testing method in the above embodiments, which will not be repeated here.
[0140] At least one embodiment of this disclosure also provides an electronic device including a processor and a memory, the memory storing one or more computer program modules. The one or more computer program modules are configured to be executed by the processor to implement the performance testing method described above. This electronic device enables the parallel execution of broadcast operations with other operations (e.g., row switching operations), weakening the sequential relationship between broadcast operations and other operations in the workflow, reducing the time wasted waiting for broadcast operations to execute, and mitigating the impact of communication on the performance of the high-performance computing system (HPL).
[0141] Figure 8 This is a schematic block diagram of an electronic device provided for some embodiments of this disclosure. For example... Figure 8 As shown, the electronic device 600 includes a processor 610 and a memory 620. The memory 620 stores non-transitory computer-readable instructions (e.g., one or more computer program modules). The processor 610 is used to execute the non-transitory computer-readable instructions, which, when executed by the processor 610, perform one or more steps in the performance testing method described above. The memory 620 and the processor 610 can be interconnected via a bus system and / or other forms of connection mechanisms (not shown). For specific implementations and explanations of the various steps of the performance testing method, please refer to the embodiments of the performance testing method described above; repetitions will not be repeated here.
[0142] It should be noted that Figure 6The components of the electronic device 600 shown are merely exemplary and not limiting. The electronic device 600 may have other components as needed for the actual application.
[0143] For example, the processor 610 and the memory 620 can communicate with each other directly or indirectly.
[0144] For example, processor 610 and memory 620 can communicate via a network. The network can include wireless networks, wired networks, and / or any combination of wireless and wired networks. Processor 610 and memory 620 can also communicate with each other via a system bus, and this disclosure is not limiting in this regard.
[0145] For example, the processor 610 and memory 620 can be located on the server side (or in the cloud).
[0146] For example, processor 610 can control other components in electronic device 600 to perform desired functions. For example, processor 610 can be a central processing unit (CPU), a graphics processing unit (GPU), or other form of processing unit with data processing capabilities and / or program execution capabilities. For example, the central processing unit (CPU) can be an x86 or ARM architecture. Processor 610 can be a general-purpose processor or a special-purpose processor, and can control other components in electronic device 600 to perform desired functions.
[0147] For example, memory 620 may include any combination of one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. One or more computer program modules may be stored on the computer-readable storage medium, and processor 610 may run one or more computer program modules to implement various functions of electronic device 600. Various application programs and various data, as well as various data used and / or generated by the application programs, may also be stored in the computer-readable storage medium.
[0148] It should be noted that, in the embodiments of this disclosure, the specific functions and technical effects of the electronic device 800 can be referred to the description of the performance testing method above, and will not be repeated here.
[0149] Figure 9This is a schematic block diagram of another electronic device provided in some embodiments of the present disclosure. The electronic device 700 is, for example, suitable for implementing the performance testing methods provided in the embodiments of the present disclosure. The electronic device 700 may be a terminal device, etc. It should be noted that... Figure 9 The illustrated electronic device 700 is merely an example and does not impose any limitation on the functionality and scope of use of the embodiments of this disclosure.
[0150] like Figure 9 As shown, the electronic device 700 may include a processing unit (e.g., a central processing unit, a graphics processor, etc.) 710, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 720 or a program loaded from a storage device 780 into a random access memory (RAM) 730. The RAM 730 also stores various programs and data required for the operation of the electronic device 700. The processing unit 710, the ROM 720, and the RAM 730 are interconnected via a bus 740. An input / output (I / O) interface 750 is also connected to the bus 740.
[0151] Typically, the following devices can be connected to the I / O interface 750: input devices 760 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 770 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 780 including, for example, magnetic tape, hard disk, etc.; and communication devices 790. The communication device 790 allows the electronic device 700 to communicate wirelessly or wiredly with other electronic devices to exchange data. Although Figure 9 An electronic device 700 with various devices is shown, but it should be understood that it is not required to implement or have all of the devices shown, and the electronic device 700 may alternatively implement or have more or fewer devices.
[0152] For example, according to embodiments of this disclosure, the performance testing method described above can be implemented as a computer software program. For instance, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program including program code for performing the performance testing method described above. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 790, or installed from a storage device 780, or installed from a ROM 720. When the computer program is executed by the processing device 710, the functions defined in the performance testing method provided by embodiments of this disclosure can be implemented.
[0153] At least one embodiment of this disclosure also provides a computer-readable storage medium storing non-transitory computer-readable instructions that, when executed by a computer, can implement the performance testing method described above. This computer-readable storage medium enables the parallel execution of broadcast operations and other operations (e.g., row switching operations), reducing the sequential relationship between broadcast operations and other operations, minimizing time wasted waiting for broadcast operations to execute, and reducing the impact of communication on the performance of the high-performance computing system (HPL).
[0154] Figure 10 This is a schematic diagram of a storage medium provided for some embodiments of this disclosure. For example... Figure 10 As shown, the storage medium 800 stores non-transitory computer-readable instructions 810. For example, when the non-transitory computer-readable instructions 810 are executed by a computer, one or more steps in the performance testing method described above are performed.
[0155] For example, the storage medium 800 can be used in the aforementioned electronic device 600. For example, the storage medium 800 can be... Figure 8 The memory 620 in the illustrated electronic device 600. For example, a description of the storage medium 800 can be found here. Figure 8 The corresponding description of the memory 620 in the illustrated electronic device 600 will not be repeated here.
[0156] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0157] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0158] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
[0159] The following points should be noted regarding this disclosure:
[0160] (1) The accompanying drawings of the embodiments of this disclosure only involve the structures involved in the embodiments of this disclosure. Other structures can be referred to the general design.
[0161] (2) Where there is no conflict, the embodiments of this disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.
[0162] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. The scope of protection of this disclosure should be determined by the scope of protection of the claims.
Claims
1. A performance testing method for a computer system, wherein, The computer system runs M processes, each process including a main thread and a broadcast sub-thread, and the method includes: The M processes perform multiple decomposition and update operations on the target matrix, wherein the target matrix comprises multi-row, multi-column blocks arranged in an array, and the multi-row, multi-column blocks correspond to the M processes. Each decomposition and update operation includes: Determine the target block from the multi-row, multi-column blocks of the target matrix; The main threads of N processes among the M processes, corresponding to the column where the target block is located, perform the decomposition operation on the target block to obtain the decomposition matrix and row exchange information for this operation. The broadcast sub-threads of the M processes perform inter-process broadcast operations for exchanging information about the decomposition matrix and the rows; The main thread of each process performs row swapping and matrix update operations based on the decomposition matrix and the row swapping information. Where M is an integer greater than 1, and N is a positive integer less than M.
2. The performance testing method according to claim 1, wherein, The broadcast sub-threads of the M processes perform inter-process broadcast operations for the decomposition matrix and the row exchange information, including: The broadcast sub-threads of the N processes broadcast the row swapping information of the current operation to the broadcast sub-threads of the other processes among the M processes, so that the other processes can start executing the row swapping operation; After broadcasting the row exchange information of this operation, the broadcast sub-threads of the N processes broadcast the decomposition matrix of this operation to the broadcast sub-threads of the other processes.
3. The performance testing method according to claim 2, wherein, The main thread of each process performs row swapping and matrix update operations based on the decomposition matrix and the row swapping information, including: In response to obtaining the row swap information for the current operation, the main thread of each process begins to perform a row swap operation on the target matrix based on the current row swap information to obtain the swap matrix for the current operation; After obtaining the decomposition matrix of the current operation, the main thread of each process performs a matrix update operation on the target matrix based on the decomposition matrix and the swap matrix of the current operation.
4. The performance testing method according to any one of claims 1 to 3, wherein, The multi-row, multi-column block is divided into M groups, the same number as the M processes. Each group forms a block matrix, and each process corresponds to one block matrix. The block matrix corresponding to each process includes a first part and a second part. The M processes perform multiple decomposition and update operations on the target matrix, including: During the execution of a row swap operation for the first part based on the row swap information of the current operation, each process performs a matrix update operation for the second part based on the decomposition matrix obtained from the previous operation of the current operation. During the matrix update operation for the first part based on the decomposed matrix of the current operation, each process performs a row swap operation for the second part based on the row swap information of the current operation.
5. The performance testing method according to claim 4, wherein, The main thread of each process performs row swapping and matrix update operations based on the decomposition matrix and the row swapping information, including: After obtaining the row swap information for this operation, the main thread of each process performs a row swap operation on the first part based on the row swap information for this operation to obtain the first swap matrix for this operation. After obtaining the decomposition matrix of the current operation, the main thread of each process performs a matrix update operation on the first part based on the decomposition matrix of the current operation and the first exchange matrix of the current operation; The main thread of each process performs a row swap operation on the second part based on the row swap information of the current operation to obtain the second swap matrix of the current operation; The main thread of each process performs a matrix update operation on the second part based on the decomposition matrix of the current operation and the second exchange matrix of the current operation.
6. The performance testing method according to claim 5, wherein, During the execution of a row swap operation for the first part based on the row swap information of the current operation, each process performs a matrix update operation for the second part based on the decomposition matrix obtained from the previous operation, including: During the row swap operation for the first part performed by the main thread of each process based on the row swap information of the current operation, a matrix update operation for the second part is performed based on the decomposition matrix obtained in the previous operation and the second swap matrix obtained in the previous operation. During the matrix update operation for the first part based on the decomposed matrix of the current operation, each process performs a row swap operation for the second part based on the row swap information of the current operation, including: During the matrix update operation for the first part based on the decomposition matrix and the first swap matrix of the current operation, the main thread of each process performs a row swap operation for the second part based on the row swap information of the current operation.
7. The performance testing method according to claim 6, wherein, The process of performing multiple decomposition and update operations on the target matrix by the M processes further includes: During the matrix update operation for the second part based on the decomposition matrix and the second swap matrix of the current operation, each process, after obtaining the row swap information of the next operation, performs a row swap operation for the first part based on the row swap information of the next operation.
8. The performance testing method according to claim 4, wherein, For each process, the first part is at least one column in the block matrix, and the second part is the remaining columns in the block matrix other than the at least one column.
9. The performance testing method according to any one of claims 1 to 3, wherein, The M process arrays are arranged as follows: The broadcast sub-threads of the M processes perform inter-process broadcast operations for the decomposition matrix and the row exchange information, including: Multiple processes located in the same process line broadcast the line exchange information and the decomposition matrix in a point-to-point manner within the process line.
10. The performance testing method according to any one of claims 1 to 3, wherein, The multi-row, multi-column block includes P diagonal blocks located on the diagonal of the target matrix; The M processes perform multiple decomposition and update operations on the target matrix, including: The P diagonal blocks are respectively used as P target blocks, and the M processes sequentially perform the decomposition and update operations P times based on the P target blocks; Where P is an integer greater than 1.
11. The performance testing method according to any one of claims 1 to 3, wherein, The target matrix is stored in row-based storage format.
12. A performance testing apparatus for a computer system, wherein, The computer system runs M processes, each process including a main thread and a broadcast sub-thread; the device includes: A decomposition and update module is configured to cause the M processes to perform multiple decomposition and update operations on a target matrix, wherein the target matrix comprises multi-row, multi-column blocks arranged in an array, and the multi-row, multi-column blocks correspond to the M processes. The decomposition and update module includes: The determination submodule is configured to determine the target block from the multi-row, multi-column blocks of the target matrix; The decomposition submodule is configured to enable the main threads of N processes among the M processes that correspond to the column where the target block is located to perform the decomposition operation on the target block, and obtain the decomposition matrix and row exchange information for this operation; The broadcast submodule is configured to enable the broadcast sub-threads of the M processes to perform inter-process broadcast operations for exchanging information about the decomposition matrix and the rows; The update submodule is configured to enable the main thread of each process to perform row swapping and matrix update operations based on the decomposition matrix and the row swapping information. Where M is an integer greater than 1, and N is a positive integer less than M.
13. The performance testing apparatus according to claim 12, wherein, The broadcast submodule is configured as follows: The broadcast sub-threads of the N processes broadcast the row swapping information of the current operation to the broadcast sub-threads of the other processes among the M processes, so that the other processes can start executing the row swapping operation; After broadcasting the row exchange information of this operation, the broadcast sub-threads of the N processes broadcast the decomposition matrix of this operation to the broadcast sub-threads of the other processes.
14. The performance testing apparatus according to claim 13, wherein, The update submodule is configured as follows: In response to obtaining the row swap information for the current operation, the main thread of each process begins to perform a row swap operation on the target matrix based on the current row swap information to obtain the swap matrix for the current operation; After obtaining the decomposition matrix of the current operation, the main thread of each process performs a matrix update operation on the target matrix based on the decomposition matrix and the swap matrix of the current operation.
15. A computer system comprising: M process modules are configured to perform multiple decomposition and update operations on a target matrix, wherein the target matrix comprises multi-row, multi-column blocks arranged in an array, and the multi-row, multi-column blocks correspond to the M processes. Each process module includes a main thread unit and a broadcast sub-thread unit. In each of the decomposition and update operations, the target block is determined from the multi-row, multi-column blocks of the target matrix; The main thread units of the N process modules corresponding to the column where the target block is located in the M process modules are configured to: perform the decomposition operation on the target block and obtain the decomposition matrix and row exchange information for this operation; The broadcast sub-thread units of the M process modules are configured to perform inter-process broadcast operations for the decomposition matrix and the row exchange information; The main thread unit of each process module is configured to perform row swapping and matrix update operations based on the decomposition matrix and the row swapping information. Where M is an integer greater than 1, and N is a positive integer less than M.
16. An electronic device comprising: processor; Memory, which stores one or more computer program modules; The one or more computer program modules are configured to be executed by the processor to implement the performance testing method according to any one of claims 1-11.
17. A computer-readable storage medium storing non-transitory computer-readable instructions that, when executed by a computer, can implement the performance testing method according to any one of claims 1-11.
Citation Information
Patent Citations
GPU-based batch parallel LU decomposition method of small and medium-sized dense matrixes for simulation system
CN114911619A
Method and equipment for optimizing high-performance Linpack benchmark test program based on heterogeneous platform
CN115827251A