Iterative solver optimization method and system based on Shenwei architecture
By using precision selection strategies, matrix segmentation algorithms and hybrid precision algorithms on Shenwei architecture, it is optimized for the iterative solver of open source mathematical software libraries such as AztecOO, which solves the problems of low computational efficiency, low accuracy and high memory costs caused by double-precision algorithms, and achieves more efficient numerical calculations.
Patent Information
- Application Number
- CN202510119252.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-16
AI Technical Summary
In the prior art, open source mathematical software libraries such as AztecOO use double-precision algorithms in Shenwei's new generation supercomputers, resulting in low computing efficiency, low computing accuracy and high memory costs.
The iterative solver is optimized by using precision selection strategies, matrix segmentation algorithms and hybrid precision algorithms on the Shenwei architecture. The specific steps include obtaining the solution task of sparse linear equation systems, dividing the tasks and allocating memory, calling the calculation from the kernel, and converting the double-precision matrix into a combination of single-precision and double-precision matrix through the accuracy selection strategy and matrix segmentation algorithm, and using a mixed precision algorithm for solving.
While maintaining calculation accuracy, it improves computing efficiency, saves memory resources, and significantly accelerates the computing process.
Smart Images

Figure CN120012426A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of electronic information technology, and in particular, relates to an iterative solver optimization method and system based on a Shenwei architecture. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] In many scientific and engineering applications such as computational fluid dynamics and computational electromagnetics, the solution process of sparse linear equations is the main performance bottleneck. The efficient solution of this process is the key and core to improve the overall performance of the application. The solution of sparse linear equations usually relies on open source mathematical software libraries. Many mathematical software libraries have been developed, and their emergence has brought great convenience to scientific computing. Among them, AztecOO is an important iterative solver for sparse linear equations in the mathematical software library Trilinos, which includes three classic iterative solution algorithms: GMRES, PCG, and BiCGSTAB. With its efficient numerical algorithm, excellent scalability and portability, it is particularly outstanding in scientific and engineering applications such as fluid mechanics and circuit simulation.
[0004] However, the algorithms for solving linear equations in AztecOO and many other open source mathematical software libraries often use double precision because it has superior computational accuracy and numerical stability compared to single precision. Although double-precision floating-point algorithms are favored in scientific computing for their superior computational accuracy and numerical stability, the results of double-precision calculations often exceed the accuracy requirements of many engineering applications. And with the explosive growth of data volume, the size of matrices is getting larger and larger. If stored in double-precision floating-point format, it will cost a huge memory cost just to store the matrix itself. So single precision (and half precision) began to be used in applications.
[0005] Generally speaking, on various commercial processors, the speed of single-precision floating-point operations is higher than that of double-precision floating-point operations. There are three main reasons: First, single-precision floating-point numbers use 32 bits (4 bytes), while double-precision floating-point numbers use 64 bits (8 bytes). The smaller the data bit width processed by the processor in each operation, the less storage bandwidth and computing cycles are required. Second, single precision can move at a higher rate in the memory hierarchy because the amount of data transferred is greatly reduced. Third, single-precision data takes up less memory than double-precision data, which means that more single-precision values can be saved in the cache than double-precision data. Shenwei's new generation supercomputer Shen is a self-developed high-performance computing platform with excellent computing power, suitable for handling large-scale and complex scientific computing problems. The platform is equipped with the SW26010Pro heterogeneous multi-core processor, each processor integrates 6 core groups, containing a total of 390 processing elements. Each core group has one master core and 64 slave cores, all of which use the independently developed SW64 instruction set and share 16GB of main memory. The master core and slave cores support IEEE754 standard half-precision, single-precision, and double-precision floating-point data types.
[0006] However, the inventors found that for a class of applications that use AztecOO software packages as the underlying call to improve numerical calculations, the lower limit of their performance depends on the optimization effect of such software packages under the current architecture. However, the AztecOO in the current new generation of supercomputers in Sunway that is suitable for the SW26010Pro architecture still uses double-precision algorithms, which have high storage and computing costs, and the AztecOO solver has low computing efficiency and low computing accuracy. Summary of the invention
[0007] In order to overcome the deficiencies of the above-mentioned prior art, the present invention provides an iterative solver optimization method and system based on the Shenwei architecture. On the basis of the Shenwei architecture, the iterative solver is optimized by adopting a precision selection strategy, a matrix partitioning algorithm, and a mixed precision algorithm. While ensuring sufficient calculation accuracy, the algorithm calculation speed is improved, memory resources are saved, and the entire solver solution process is accelerated, providing researchers in numerical calculations with more efficient development efficiency.
[0008] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:
[0009] A first aspect of the present invention provides an iterative solver optimization method based on the Shenwei architecture, comprising:
[0010] Obtaining a solution task for a sparse linear equation system, and sending the solution task to a main core;
[0011] The master core divides the solution tasks and allocates memory for them; the master core calls the slave core startup function and performs the calculation;
[0012] The slave core feeds back the calculation results to the master core to obtain the solution of the sparse linear equation system;
[0013] The master core divides the solution tasks and allocates memory for them; the master core calls the slave core startup function and performs the calculation. The specific process is as follows:
[0014] Get the non-zero element values of the initial double-precision sparse matrix and the input vector, where the input vector is randomly generated;
[0015] Based on the non-zero element values of the input sparse matrix, a precision selection strategy is adopted;
[0016] Based on the precision selection strategy, the initial double-precision sparse matrix is converted into a format through a matrix partitioning algorithm;
[0017] A mixed precision algorithm is used to solve the matrix after format conversion.
[0018] As an implementation method, based on the non-zero element values of the input sparse matrix, a precision selection strategy is adopted, specifically, a precision interval range is set, the non-zero elements in the matrix are divided into different intervals, and the corresponding precision is selected according to the data size.
[0019] As an implementation method, the corresponding precision is selected according to the data size, and the formula is:
[0020]
[0021] Among them, value represents the value of the sparse moment, sp represents single precision, that is, the value is within the precision range, and dp represents double precision, that is, the value is not within the precision range.
[0022] As an implementation method, the format of the initial double-precision sparse matrix is converted into a single-precision matrix and a double-precision moment. The specific process is as follows:
[0023] Based on the initial double-precision sparse matrix and the precision interval range, traverse each non-zero element in the sparse matrix;
[0024] Determine whether the non-zero element is within the precision interval. If so, load the element into a single-precision matrix; if not, load it into a double-precision matrix.
[0025] Convert single-precision and double-precision matrices to sparse matrix compressed format.
[0026] As an implementation method, a mixed precision algorithm is used to solve the matrix after format conversion. The specific process is:
[0027] Expand the formal parameters of the original dense function prototype into single-precision matrix array and double-precision matrix array forms;
[0028] Copy a double-precision vector to a single-precision vector;
[0029] Based on the copied single-precision and single-precision matrices, perform two sparse matrix-vector multiplications on the format-converted matrices to obtain the final solution.
[0030] As an implementation method, the single-precision matrix array form includes: a single-precision matrix non-zero value array, a non-zero element column coordinate array, a row pointer array, and a single-precision array.
[0031] As an implementation method, two sparse matrix-vector multiplications are performed on the format-converted matrix to obtain the final solution result. The specific process is as follows:
[0032] Perform the first single-precision sparse matrix-vector multiplication on all non-zero values of the single-precision matrix and the single-precision array to obtain a single-precision vector;
[0033] Perform a second double-precision sparse matrix-vector multiplication of all non-zero values of the double-precision matrix with the double-precision array to obtain a double-precision vector;
[0034] Accumulate the single-precision vector and the double-precision vector to obtain the final solution result.
[0035] A second aspect of the present invention provides an iterative solver optimization system based on the Shenwei architecture, comprising:
[0036] An acquisition task module is used to acquire a solution task of a sparse linear equation system and send the solution task to the main core;
[0037] The computing module is used for the main core to divide the solution tasks and allocate memory for them; the main core calls the slave core startup function and performs the calculation; the main core divides the solution tasks and allocates memory for them; the main core calls the slave core startup function and performs the calculation. The specific process is as follows:
[0038] Get the non-zero element values of the initial double-precision sparse matrix and the input vector, where the input vector is randomly generated;
[0039] Based on the non-zero element values of the input sparse matrix, a precision selection strategy is adopted;
[0040] Based on the precision selection strategy, the initial double-precision sparse matrix is converted into a format through a matrix partitioning algorithm;
[0041] Use mixed precision algorithm to solve the matrix after format conversion;
[0042] The result feedback module is used to feed back the calculation results from the slave core to the master core to obtain the solution of the sparse linear equation system.
[0043] A third aspect of the present invention provides a computer device, comprising a memory, a processor, and a program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in the method described in the first aspect of the present invention are implemented.
[0044] The fourth aspect of the present invention aims to provide a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps in the method described in the first aspect of the present invention.
[0045] One or more of the above technical solutions have the following beneficial effects:
[0046] This embodiment is based on the implementation of a mixed-precision version of the AztecOO solver on the new generation supercomputer of Shenwei equipped with the SW26010Pro processor. Without sacrificing the necessary precision, single precision is introduced into the original double-precision algorithm to implement a high-performance iterative solution algorithm in a mixed-precision format. This method and system can not only improve computing efficiency, save memory resources, and accelerate the computing process while maintaining computing accuracy, but also provide a reference for adopting mixed-precision optimization strategies on domestic supercomputing platforms.
[0047] In this embodiment, a precision selection strategy is proposed based on the size of each specific non-zero element value in the input sparse matrix: an interval range is set, the non-zero elements in the matrix are divided into different intervals, and the appropriate precision is selected according to the value size. This strategy can reduce storage and calculation overhead by using single precision without significantly losing calculation accuracy, while ensuring high precision requirements for larger values.
[0048] In this embodiment, a matrix splitting algorithm splitMatrix is designed to convert an initial single double-precision matrix into two matrices: a combination of a single-precision matrix and a double-precision matrix, thereby saving a considerable amount of storage space for the sparse matrix.
[0049] This embodiment designs a mixed precision MPepetra_dcrsmv algorithm for the original double precision epetra_dcrsmv of the AztecOO solver, which not only saves SPMV memory requirements, but also increases parallelism for the program and improves the SPMV kernel computing efficiency.
[0050] Based on the Shenwei architecture, a heterogeneous multi-core parallel design is performed for the mixed-precision MPepetra_dcrsmv. The master-slave acceleration model is used to accelerate the single-precision matrix As calculation module and the double-precision matrix Ad calculation module. The powerful slave core computing resources of Shenwei are used to accelerate the calculation process.
[0051] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0053] Figure 1 This is a flow chart of an iterative solver optimization method based on the Shenwei architecture in the first embodiment of the present invention;
[0054] Figure 2 Schematic diagram of the matrix splitting algorithm splitMatrix of the first embodiment;
[0055] Figure 3 Schematic diagram of data partitioning of sparse matrix non-zero values and input vector x in the first embodiment;
[0056] Figure 4 This is the acceleration effect test of the first embodiment under the SW26010Pro single-core group, taking the running time of the original double-precision version epetra_dcrsmv as the test benchmark. DETAILED DESCRIPTION
[0057] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.
[0058] It should be noted that the terms used herein are for describing specific embodiments only and are not intended to be limiting of exemplary embodiments according to the present invention.
[0059] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.
[0060] Terminology explanation:
[0061] AztecOO: AztecOO is an iterative solver in the Trilinos mathematical software library for solving sparse linear systems. It iteratively solves sparse linear equations through the Krylov subspace method with preconditioning. The solver includes three iterative solution algorithms: GMRES, PCG, and BiCGSTAB.
[0062] CSR: Compressed sparse row, compresses the sparse matrix in the form of row compression, retaining only the non-zero elements of the sparse matrix.
[0063] epetra_dcrsmv: A fortran77 subroutine built into AztrcOO that implements the sparse matrix-vector multiplication kernel (SPMV).
[0064] SPMV: Sparse Matrix Vector Multiplication (abbreviation);
[0065] MPepetra_dcrsmv: A mixed precision version of epetra_dcrsmv based on the original AztecOO solver. MP is the abbreviation for mixed precision.
[0066] Embodiment 1
[0067] This embodiment discloses an iterative solver optimization method based on the Shenwei architecture.
[0068] In order to more clearly illustrate this embodiment, an iterative solver optimization implementation process based on the Shenwei architecture can be specifically described as follows:
[0069] An iterative solver optimization method based on the Shenwei architecture, comprising:
[0070] S1. Obtain a solution task for a sparse linear equation system, and send the solution task to the main core;
[0071] S2, the main core divides the solution task and allocates memory for it; the main core calls the slave core startup function and performs calculations; the main core divides the solution task and allocates memory for it; the main core calls the slave core startup function and performs calculations. The specific process is as follows:
[0072] Get the non-zero element values of the initial double-precision sparse matrix and the input vector, where the input vector is randomly generated;
[0073] Based on the non-zero element values of the input sparse matrix, a precision selection strategy is adopted;
[0074] Based on the precision selection strategy, the initial double-precision sparse matrix is converted into a format through a matrix partitioning algorithm;
[0075] Use mixed precision algorithm to solve the matrix after format conversion;
[0076] The slave core feeds the calculation results back to the master core to obtain the solution of the sparse linear equations.
[0077] like Figure 1 As shown, in step S1, a solution task of a sparse linear equation group is obtained, and the solution task is sent to the main core.
[0078] In this embodiment, a sparse coefficient matrix derived from the discretization process of the partial differential equation (PDE) in electromagnetic wave propagation is obtained from the matrix market, and is input into the SW26010Pro single-core group in the new generation supercomputer of Sunway, and is solved using the GMRES algorithm in AztecOO.
[0079] like Figure 1 As shown, in step S2, the main core divides the solution task and allocates memory for it; the main core calls the slave core startup function and performs calculations; wherein the main core divides the solution task and allocates memory for it; the main core calls the slave core startup function and performs calculations, the specific process is:
[0080] Get the non-zero element values of the initial double-precision sparse matrix and the input vector, where the input vector is randomly generated;
[0081] Based on the non-zero element values of the input sparse matrix, a precision selection strategy is adopted;
[0082] Based on the precision selection strategy, the initial double-precision sparse matrix is converted into a format through a matrix partitioning algorithm;
[0083] A mixed precision algorithm is used to solve the matrix after format conversion.
[0084] In this embodiment, S2-1, obtain non-zero element values of an initial double-precision sparse matrix and an input vector, wherein the input vector is generated randomly.
[0085] The master-slave core code is written by calling the athread acceleration thread library function interface customized for the Shenwei platform, and the master-slave acceleration model is adopted to accelerate the above-mentioned single-precision matrix As calculation module and double-precision matrix Ad calculation module. According to the special master-slave core hardware architecture of the Shenwei heterogeneous many-core processor, the two calculation modules use the slave core as the acceleration core to improve the computing efficiency. The master core is responsible for reading the non-zero element data of the initial double-precision sparse matrix A and the double-precision input vector x, and allocating memory for them.
[0086] In this embodiment, S2-2, a precision selection strategy is adopted based on the non-zero element values of the input sparse matrix.
[0087] In the existing technology, simply taking a piece of code and reducing its precision by rewriting all double precision to single precision can lead to inaccuracies in floating point calculations, which can in turn negatively affect the convergence of the solver. Therefore, any mixed precision method must identify data and calculations that can be reduced to single, half, or other floating point representations without compromising the accuracy of the underlying calculations.
[0088] The IEEE 754 standard defines the representation and operation rules of floating-point numbers, where single-precision and double-precision floating-point numbers occupy 32 bits and 64 bits respectively. Both are composed of a sign bit, an exponent bit, and a mantissa bit. Single-precision uses an 8-bit exponent and a 23-bit mantissa, while double-precision uses an 11-bit exponent and a 52-bit mantissa. Among them, the IEEE 754 floating-point operation standard guarantees the relative error of the floating-point value, and the limiting formula is:
[0089]
[0090] Where η represents the round-off error unit or machine precision, β is the base of the numbering system, and m represents the precision or number of digits.
[0091] According to formula (1), for single-precision applications, if there are 23 effective binary digits, approximately 7 effective decimal digits are obtained. For double-precision applications, if there are 52 effective binary digits, approximately 16 effective decimal digits are obtained.
[0092] The final result of the iterative solution of sparse linear equations is affected by both relative errors and absolute errors that depend on matrix and vector values. The iterative solution process involves the accumulation of a large number of basic operations such as vector inner products and additions. The IEEE 754 standard ensures that the relative error of floating-point numbers is bounded, but the absolute error is affected by the numerical range. For a floating-point number, it consists of a sign bit s, an exponent bit e, and a mantissa m. The exponent bit determines the range of the floating-point number, and the mantissa stores the significant digits of the floating-point number. The exponent bit can be used to obtain the floating-point number value in the interval [2 e ,2 e+1 ). The mantissa determines the precision of the value within the interval, and a smaller exponent e corresponds to higher precision because the minimum increment of the value change is smaller in a smaller range of values.
[0093] Specifically, the precision p of floating-point numbers in the interval [2e, 2e+1) is expressed as:
[0094]
[0095] Because m is fixed, formula (2) shows that the smaller the exponent e, the higher the precision p, because a smaller numerical range can provide finer representation accuracy. Smaller values usually correspond to smaller representation errors, while larger values may introduce larger errors, especially when performing accumulation and inner product operations, where the errors of larger values may be significantly amplified.
[0096] Based on this, the present invention proposes a precision selection strategy based on the values of non-zero elements in the matrix. By setting an interval range, the non-zero elements in the matrix are divided into different intervals, and the appropriate precision is selected according to the value size. Specifically, an interval episode is set, and the formula is:
[0097]
[0098] Among them, value represents the value of the sparse moment, sp represents single precision, that is, the value is within the precision range, and dp represents double precision, that is, the value is not within the precision range.
[0099] After the above steps, when a certain value of the matrix is within this interval, the precision p is calculated using single precision sp, because the representation error of such values is relatively small and the precision requirement is low. For larger elements, double precision dp is used to reduce the accumulation of representation errors and improve calculation accuracy. This strategy can reduce storage and calculation overhead by using single precision without significantly losing calculation accuracy, while ensuring high precision requirements for larger values.
[0100] In this embodiment, S2-3, based on the precision selection strategy, the format of the initial double-precision sparse matrix is converted by a matrix partitioning algorithm.
[0101] Use the AztecOO solver to solve sparse linear equations. The solver provides a data format conversion function. The solver converts the input matrix COO format to the CSR compressed storage format. This conversion is performed before entering the iterative algorithm and occurs once during runtime. Since the iterative solution algorithm (such as the conjugate gradient method) uses the same input matrix from beginning to end, the one-time time cost of this conversion is evenly distributed over multiple iterations.
[0102] Based on the fine-grained precision selection strategy proposed above, a matrix partitioning algorithm splitMatrix is designed for the original input matrix to convert the initial single double-precision matrix A into a single-precision matrix As and a double-precision matrix Ad.
[0103] The pseudo code of the matrix splitting algorithm splitMatrix is shown in Algorithm 1:
[0104]
[0105] As shown in Algorithm 1, by inputting the original double-precision sparse matrix A and the set precision selection interval episode, traverse each non-zero element in the sparse matrix, compare the value with the precision selection interval episode, and load the element into the single-precision matrix As if it is within the episode interval, otherwise load it into Ad. After traversing all non-zero elements, convert them into CSR matrix compression format.
[0106] Through the matrix splitting algorithm splitMatrix, the conversion of a single double-precision matrix A into two matrices: a single-precision matrix As and a double-precision matrix Ad can be completed, and some storage space can be saved.
[0107] Specifically, (1) based on the initial double-precision sparse matrix and the precision interval range, traverse each non-zero element in the sparse matrix.
[0108] (2) Determine whether the non-zero element is within the precision interval. If so, load the element into a single-precision matrix; if not, load it into a double-precision matrix.
[0109] Call and execute the matrix splitting algorithm splitMatrix to divide and manage the data of the single-precision matrix As calculation module and the double-precision matrix Ad calculation module.
[0110] Among them, data partitioning mainly includes the partitioning of matrix non-zero values and input vector x. Specifically, taking the single-precision matrix As as an example, the number of enabled slave cores in a single core group is defined as N_core, and the sparse matrix As is m rows and n columns. Since the matrix storage format is row-compressed, its non-zero elements are arranged continuously by row in memory, so we divide the matrix into N_core block matrices (block) by row, and each block includes m / Ncore rows and n columns of the matrix As. Each block contains a part of the array value, col, and indx and allocates the part to the LDM of the same slave core. Since the storage space of each slave core is limited, the block is further divided into multiple sub-block matrices (subblock). Since the memory size of each LDM is 256KB, assuming that the space for storing vector x is Memory_x, the size of the subblock is (256-Memory_x)*1024 / 4 bytes. Each slave core loop obtains a data segment of the size of a subblock to the slave core LDM. For the input vector x, since each sub-block matrix of the slave core needs to be calculated with x, and the access to vector x is very irregular, the partition strategy of vector x is crucial. The slave core LDM supports the partitioning of private segments and shared segments, and by partitioning x into shared segments, the efficiency of accessing vector x during the slave core calculation process is improved.
[0111] (3) Convert single-precision matrices and double-precision matrices into sparse matrix compression formats.
[0112] In this embodiment, the single-precision matrix and the double-precision matrix are converted into a sparse matrix compression format, as shown in Table 1. The number of non-zero elements in the matrix is defined as Nnz, the number of elements in the episode interval is Nnz_s, and the number of elements outside the interval is (Nnz-Nnz_s). Because the matrix is stored in CSR format, in addition to storing the matrix non-zero value value array, auxiliary arrays are also required: the non-zero element column coordinate array col, and the row pointer array indx.
[0113] Through the matrix splitting algorithm splitMatrix, you can save
[0114] (12Nnz+20N+4)-(12Nnz-4Nnz_s+28N+8)=4Nnz_s-8N-4 bytes of storage space.
[0115]
[0116]
[0117] After the above steps, the initial single double-precision matrix is converted into a combination of single-precision and double-precision matrices, saving a considerable amount of storage space for the sparse matrix.
[0118] In this embodiment, S2-4, a mixed precision algorithm is used to solve the matrix after format conversion.
[0119] AztecOO is an iterative solver for sparse linear equations. During the iteration process, a large number of sparse matrix-vector multiplications (SPMV) are the main performance bottleneck. The epetra_dcrsmv function inside the AztecOO solver implements the SPMV kernel, in which the non-zero values of the sparse matrix are stored in the dobule double-precision CSR format, as shown in Algorithm 2.
[0120]
[0121] Algorithm 2 describes the original double-precision version epetra_dcrsmv, which is used for CSR format sparse matrix-vector multiplication. The outer loop traverses each row of the sparse matrix, and the inner loop traverses all non-zero elements of the current row, multiplies the current non-zero element with the element at the corresponding position in the input vector x, and then adds it to the current row position of the result vector y.
[0122] In order to reduce the memory requirements of kernel execution and improve execution efficiency, the original double-precision version epetra_dcrsmv is designed into a mixed-precision MPepetra_dcrsmv algorithm. The mixed-precision algorithm is used to solve the matrix after format conversion. The specific process is as follows:
[0123] (1) Expand the formal parameters of the original dense function prototype into single-precision matrix array and double-precision matrix array forms.
[0124] (2) Copy the double-precision vector into a single-precision vector.
[0125] In order to facilitate calculation with the input vector x, the double-precision vector is copied to a single-precision vector x_s.
[0126] (3) Based on a copy of a single-precision and single-precision matrix, two sparse matrix-vector multiplications are performed on the format-converted matrix to obtain the final solution.
[0127] Among them, the matrix after format conversion is multiplied by sparse matrix vector twice to obtain the final solution, which specifically includes:
[0128] Perform the first single-precision sparse matrix-vector multiplication of all non-zero values of the single-precision matrix with the single-precision array to obtain a single-precision vector.
[0129] Perform a second double-precision sparse matrix-vector multiplication of all nonzero values of the double-precision matrix with the double-precision array to obtain a double-precision vector.
[0130] Accumulate the single-precision vector and the double-precision vector to obtain the final solution result.
[0131] Specifically, as shown in Algorithm 3.
[0132]
[0133]
[0134] Algorithm 3 describes the mixed-precision version of MPepetra_dcrsmv, which expands the formal parameters of the original epetra_dcrsmv function prototype. The parameters passed are expanded from the three arrays of the single matrix A to the array form of the two matrices after the partition, and a single-precision copy of the double-precision vector x is made to facilitate the calculation of the single-precision matrix As and x_s. After the mixed-precision design, SPMV is performed twice on the partitioned matrix. The first single-precision SPMV will perform the calculation of all non-zero values of the single-precision matrix As and x_s to obtain the partial result of the final value of the result vector y. The second double-precision SPMV will perform the calculation of all non-zero values of the double-precision matrix Ad and x_d to obtain the result vector y and accumulate the partial result y obtained by the first single-precision SPMV to obtain the final result y.
[0135] After the above steps, some non-zero elements of the matrix are reduced from double precision to single precision, which can greatly save SPMV memory requirements; the matrix is divided into two parts. Although the number of matrix rows in the two matrices is m, the number of non-zero elements in each row of a single matrix will be reduced, so the number of inner loops of the algorithm is reduced, thereby reducing the number of vector inner product calculations, greatly reducing the calculation time of a single-row sparse matrix-vector multiplication; after matrix partitioning, the calculation processes of single-precision matrices and double-precision matrices are completely independent, thereby enhancing the parallelism of the program.
[0136] like Figure 1 As shown, in step S3, the slave core feeds back the calculation result to the master core to obtain the solution of the sparse linear equation system.
[0137] Each slave core cycle obtains a subblock-sized data segment to the slave core LDM. For the input vector x, since each slave core's subblock matrix needs to be calculated with x, and the access to vector x is very irregular, the partitioning strategy for vector x is crucial. The slave core LDM supports the partitioning of private segments and shared segments. By partitioning x into shared segments, the efficiency of accessing vector x during the slave core calculation process is improved.
[0138] The slave core is responsible for obtaining different data blocks from the main memory and asynchronously executing single-precision sparse matrix-vector multiplication SPMV and double-precision sparse matrix-vector multiplication SPMV to accelerate the tasks of the two computing modules.
[0139] The vector result obtained by the first single-precision sparse matrix-vector multiplication and the vector result obtained by the second double-precision sparse matrix-vector multiplication are accumulated to obtain the final solution result, which is fed back to the main core.
[0140] In this embodiment, Figure 4As shown, the optimized iterative solver based on the Shenwei architecture solves any sparse linear equation. When the MPepetra_dcrsmv algorithm is used to optimize the SPMV kernel, a performance improvement of about 1.68 times is obtained compared to the benchmark. It can be seen that the mixed precision design in this embodiment reduces some non-zero elements of the matrix from double precision to single precision, which can greatly save the SPMV memory requirements. At the same time, the matrix is divided into two parts, and the number of non-zero elements in each row of a single matrix will be reduced. Therefore, the number of inner loops of the sparse matrix-vector multiplication SPMV is reduced, thereby reducing the number of vector inner product calculations and reducing the calculation time of a single row of sparse matrix-vector multiplication. After the MPepetra_dcrsmv algorithm is kernelized, the powerful kernel resources of the SW26010Pro processor are fully utilized, the calculation time is further reduced, and a performance improvement of more than 3.89 times is obtained, and the calculation efficiency is significantly improved.
[0141] In terms of the accuracy of the calculation results, for the original double-precision version of the algorithm, when GMRES solves the sparse linear equations problem with the matrix pde2961 as the input matrix, the number of iterations is 286 and the residual is 3.01796e-07. The mixed-precision MPepetra_dcrsmv algorithm is optimized and the same number of iterations is set, and the residual is 3.01797e-07. It can be seen that after the mixed-precision design, it can ensure that the calculation time is accelerated, memory resources are saved, and the calculation efficiency is improved while being close to the accuracy of the original calculation results.
[0142] Embodiment 2
[0143] The purpose of this embodiment is to provide an iterative solver optimization system based on the Shenwei architecture, including:
[0144] An acquisition task module is used to acquire a solution task of a sparse linear equation system and send the solution task to the main core;
[0145] The computing module is used for the main core to divide the solution tasks and allocate memory for them; the main core calls the slave core startup function and performs the calculation; the main core divides the solution tasks and allocates memory for them; the main core calls the slave core startup function and performs the calculation. The specific process is as follows:
[0146] Get the non-zero element values of the initial double-precision sparse matrix and the input vector, where the input vector is randomly generated;
[0147] Based on the non-zero element values of the input sparse matrix, a precision selection strategy is adopted;
[0148] Based on the precision selection strategy, the initial double-precision sparse matrix is converted into a format through a matrix partitioning algorithm;
[0149] Use mixed precision algorithm to solve the matrix after format conversion;
[0150] The result feedback module is used to feed back the calculation results from the slave core to the master core to obtain the solution of the sparse linear equation system.
[0151] Based on providing an iterative solver optimization system based on the Shenwei architecture, the method steps in Example 1 are implemented.
[0152] Embodiment 3
[0153] The purpose of this embodiment is to provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the program.
[0154] Embodiment 4
[0155] The purpose of this embodiment is to provide a computer-readable storage medium.
[0156] A computer-readable storage medium stores a computer program, which executes the steps of the above method when executed by a processor.
[0157] The steps involved in the apparatus of the above embodiment correspond to the method embodiment 1, and the specific implementation method can refer to the relevant description part of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.
[0158] Those skilled in the art should understand that the modules or steps of the present invention described above can be implemented by a general-purpose computer device, or alternatively, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.
[0159] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.
Claims
1. An iterative solver optimization method based on the Shenwei architecture, characterized in that: include: Obtaining a solution task for a sparse linear equation system, and sending the solution task to a main core; The main core divides the solution tasks and allocates memory to them; The master core calls the slave core startup function and performs calculations; The slave core feeds back the calculation results to the master core to obtain the solution of the sparse linear equation system; The master core divides the solution tasks and allocates memory for them; the master core calls the slave core startup function and performs the calculation. The specific process is as follows: Get the non-zero element values of the initial double-precision sparse matrix and the input vector, where the input vector is randomly generated; Based on the non-zero element values of the input sparse matrix, a precision selection strategy is adopted; Based on the precision selection strategy, the initial double-precision sparse matrix is converted into a format through a matrix partitioning algorithm; A mixed precision algorithm is used to solve the matrix after format conversion.
2. The iterative solver optimization method based on the Shenwei architecture as claimed in claim 1, characterized in that: Based on the non-zero element values of the input sparse matrix, a precision selection strategy is adopted. Specifically, the precision interval range is set, the non-zero elements in the matrix are divided into different intervals, and the corresponding precision is selected according to the data size.
3. The iterative solver optimization method based on the Shenwei architecture as claimed in claim 1, characterized in that: Select the corresponding precision according to the data size. The formula is: Among them, value represents the value of the sparse moment, sp represents single precision, that is, the value is within the precision range, and dp represents double precision, that is, the value is not within the precision range.
4. The iterative solver optimization method based on the Shenwei architecture as claimed in claim 1, characterized in that: The format of the initial double-precision sparse matrix is converted into a single-precision matrix and a double-precision moment. The specific process is as follows: Based on the initial double-precision sparse matrix and the precision interval range, traverse each non-zero element in the sparse matrix; Determine whether the non-zero element is within the precision interval. If so, load the element into a single-precision matrix; if not, load it into a double-precision matrix. Convert single-precision and double-precision matrices to sparse matrix compressed format.
5. The iterative solver optimization method based on the Shenwei architecture as claimed in claim 1, characterized in that: The mixed precision algorithm is used to solve the matrix after format conversion. The specific process is as follows: Expand the formal parameters of the original dense function prototype into single-precision matrix array and double-precision matrix array forms; Copy a double-precision vector to a single-precision vector; Based on the copied single-precision and single-precision matrices, perform two sparse matrix-vector multiplications on the format-converted matrices to obtain the final solution.
6. The iterative solver optimization method based on the Shenwei architecture as claimed in claim 1, characterized in that: The single-precision matrix array forms include: single-precision matrix non-zero value array, non-zero element column coordinate array, row pointer array and single-precision array.
7. The iterative solver optimization method based on the Shenwei architecture according to claim 1, characterized in that: Perform two sparse matrix-vector multiplications on the format-converted matrix to obtain the final solution. The specific process is as follows: Perform the first single-precision sparse matrix-vector multiplication on all non-zero values of the single-precision matrix and the single-precision array to obtain a single-precision vector; Perform a second double-precision sparse matrix-vector multiplication of all non-zero values of the double-precision matrix with the double-precision array to obtain a double-precision vector; Accumulate the single-precision vector and the double-precision vector to obtain the final solution result.
8. An iterative solver optimization system based on the Shenwei architecture, characterized in that: include: An acquisition task module is used to acquire a solution task of a sparse linear equation system and send the solution task to the main core; The computing module is used to divide the solution tasks among the main cores and allocate memory to them; The master core calls the slave core startup function and performs calculations. The master core divides the solution tasks and allocates memory for them. The master core calls the slave core startup function and performs calculations. The specific process is as follows: Get the non-zero element values of the initial double-precision sparse matrix and the input vector, where the input vector is randomly generated; Based on the non-zero element values of the input sparse matrix, a precision selection strategy is adopted; Based on the precision selection strategy, the initial double-precision sparse matrix is converted into a format through a matrix partitioning algorithm; Use mixed precision algorithm to solve the matrix after format conversion; The result feedback module is used to feed back the calculation results from the slave core to the master core to obtain the solution of the sparse linear equation system.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are performed.
Citation Information
Cited By
Sublimation AscendCL-based nonlinear optimization processing method, system and device, and storage medium
CN122044891A
Processing method, system and device of nonlinear optimization based on ascend cl and storage medium
CN122044891B