Optimization method and device for solving triangular linear equations
By dividing the result vectors of triangular linear equations into task blocks and circulating to the calculation core for iterative solution, the problem of low solution efficiency in traditional methods is solved, and a more efficient calculation process is achieved.
Patent Information
- Application Number
- CN202411633939.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-15
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-11-15
AI Technical Summary
When solving triangular linear equation systems, the handling waiting time and calculation time of the calculation core are too long, resulting in low solution efficiency.
By dividing the result vectors of the target triangle linear equation system into multiple task blocks, and cyclically allocating these blocks to the computing core of the processing platform, the calculation core is used to iteratively solve the triangle block inverse matrix and input vector block corresponding to each task block until the result vector block of each task block is obtained.
The solution efficiency of triangular linear equation systems is improved, the waiting time and calculation time of the calculation core are reduced, and the inefficiency problems caused by forward iteration methods in traditional methods are avoided.
Smart Images

Figure CN119128343B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of high performance computing technology, and in particular to a method and device for optimizing the solution of a triangular linear equation system. Background Art
[0002] Linear equations are a common problem in mathematics and engineering. They can be expressed in matrix form as A*x=b, where is a known matrix, b is a known vector, and x is an unknown vector. Solving a linear equation is to solve the unknown vector x. When A is a triangular matrix (upper triangular matrix or lower triangular matrix), due to the specific structure of matrix A, a more efficient algorithm can be used to simplify the calculation process to solve such a system of equations. Solving triangular linear equations (trsv) is specially designed for this situation.
[0003] The processing platforms currently used to solve triangular linear equations still have some defects when implementing and optimizing some common basic linear algebra operators such as trsv. Generally speaking, solving triangular linear equations is a recursive process, and each step of calculation depends on the result of the previous step. However, due to hardware design limitations, recursive solving cannot be achieved on traditional processing platforms. At the same time, when solving a result vector block, a forward iteration method is used to substitute the computing units one by one for the solution, and the solution efficiency is very low.
[0004] In the above-mentioned related technologies, the traditional method is used to solve the triangular linear equations. The waiting time for the computing core and the computing time are too long. In addition, the equations are solved one by one in a forward iterative manner on the processing platform, resulting in low solution efficiency. No effective solution has been proposed yet. Summary of the invention
[0005] An embodiment of the present invention provides a method and device for optimizing the solution of a triangular linear equation system, so as to at least solve the technical problems in the related art of using traditional methods to solve a triangular linear equation system, that is, the waiting time for the computing core is too long, the computing time is too long, and the equation system is solved one by one in a forward iterative manner on the processing platform, resulting in low solution efficiency.
[0006] According to one aspect of an embodiment of the present invention, a method for optimizing the solution of a triangular linear equation group is provided, comprising: determining that a triangular linear equation group to be solved is a target triangular linear equation group; cyclically allocating a plurality of task blocks obtained by dividing a result vector of the target triangular linear equation group to a computing core of a processing platform, wherein the result vector is an unknown vector to be solved in the target triangular linear equation group, and the processing platform is used to solve the target triangular linear equation group; allocating the inverse matrix of each triangular block matrix on the diagonal of an input matrix in the target triangular linear equation group and the input vector blocks of the input vector to the computing core where the corresponding task blocks are located, wherein the triangular block matrix of the input matrix and the input vector blocks of the input vector are divided in the same way as the result vector is divided. The input matrix is the coefficient matrix of the target triangular linear equations group, the input vector is the known vector of the target triangular linear equations group, and the triangular block matrix, the input vector block and the task block correspond to each other through the divided labels; the computing core is used to iteratively solve multiple task blocks according to the triangular block inverse matrix corresponding to each task block and the input vector block until the result vector block of each task block is obtained, and then the target input vector is determined to be the solution vector of the target triangular linear equations group, wherein the triangular block inverse matrix refers to the inverse matrix of the triangular block matrix, and the target input vector is a vector obtained by covering the input vector block corresponding to the task block in the input vector using the result vector block obtained by solving each task block in the solution process.
[0007] Optionally, the multiple task blocks obtained by dividing the result vector of the target triangular linear equations group are cyclically distributed to the computing cores of the processing platform, including: dividing the result vector of the target triangular linear equations group to obtain the multiple task blocks; obtaining the multiple computing cores in the processing platform; when a first total number of the multiple task blocks does not exceed a second total number of the multiple computing cores, allocating the multiple task blocks to the computing cores in sequence; when the first total number of the multiple task blocks exceeds the second total number of the multiple computing cores, allocating the multiple task blocks to the computing cores in sequence and cyclically.
[0008] Optionally, the inverse matrices of each triangular block matrix on the diagonal of the input matrix in the target triangular linear equation group and the input vector blocks of the input vector are allocated to the computing core where the corresponding task block is located, including: dividing the matrix on the diagonal of the input matrix in the same division method as the division of the result vector to obtain a plurality of the triangular block matrices; using the computing core to calculate the inverse matrix of each triangular block matrix to obtain a plurality of the triangular block inverse matrices; dividing the input vector in the same division method as the division of the result vector to obtain a plurality of the input vector blocks; and allocating the plurality of triangular block inverse matrices and the plurality of input vector blocks to the computing core where the corresponding task block is located according to the correspondence between the triangular block matrices, the input vector blocks and the task blocks.
[0009] Optionally, the optimization method for solving the triangular linear equations also includes: after obtaining the multiple triangular block inverse matrices and the multiple input vector blocks, storing the multiple triangular block inverse matrices and the multiple input vector blocks in the global memory of the processing platform, wherein the global memory is used to store data whose data volume exceeds a data volume threshold.
[0010] Optionally, before using the computing core to iteratively solve multiple task blocks according to the triangular block inverse matrix and the input vector block corresponding to each task block, until the result vector block of each task block is obtained, and before determining that the target input vector is the solution vector of the target triangular linear equations group, the method for optimizing the solution of the triangular linear equations group also includes: transferring the multiple triangular block inverse matrices and the multiple input vector blocks stored in the global memory to a unified buffer of the processing platform, wherein the unified buffer is used to temporarily store the data to be calculated within a predetermined period of time before the calculation.
[0011] Optionally, the computing core is used to iteratively solve multiple task blocks according to the triangular block inverse matrix corresponding to each task block and the input vector block, until the result vector block of each task block is obtained, and the target input vector is determined to be the solution vector of the target triangular linear equation group, including: a solving step, using the vector calculation unit of the computing core to perform matrix-vector multiplication to solve according to the triangular block inverse matrix corresponding to the task block and the input vector block, to obtain the result vector block of the task block; an updating step, using the result vector block to cover the input vector block corresponding to the task block in the input vector; a judging step, judging whether the result vector blocks of all the task blocks have been obtained, and obtaining a judgment result; repeatedly executing the solving step, the updating step and the judging step until the judgment result indicates that the result vector blocks of all the task blocks have been obtained, determining that the input vector currently output is the target input vector; and determining that the target input vector is the solution vector of the target triangular linear equation group.
[0012] Optionally, after using the vector calculation unit of the computing core to perform matrix-vector multiplication according to the triangular block inverse matrix corresponding to the task block and the input vector block to obtain the result vector block of the task block, the method for optimizing the solution of the triangular linear equations also includes: storing the result vector block in a unified buffer of the processing platform, and before using the result vector block to overwrite the input vector block in the input vector corresponding to the task block, the method for optimizing the solution of the triangular linear equations also includes: transferring the result vector block stored in the unified buffer to the global memory of the processing platform.
[0013] According to another aspect of an embodiment of the present invention, a device for solving and optimizing a triangular linear equation group is also provided, comprising: a first determination unit, for determining that a triangular linear equation group to be solved is a target triangular linear equation group; a first allocation unit, for cyclically allocating a plurality of task blocks obtained by dividing a result vector of the target triangular linear equation group to a computing core of a processing platform, wherein the result vector is an unknown vector to be solved in the target triangular linear equation group, and the processing platform is used to solve the target triangular linear equation group; a second allocation unit, for allocating the inverse matrix of each triangular block matrix on the diagonal of an input matrix in the target triangular linear equation group and the input vector blocks of the input vector to the computing core where the corresponding task blocks are located, wherein the triangular block matrix of the input matrix and the input vector blocks of the input vector are all allocated in accordance with the division of the result vector. The input matrix is obtained by dividing in the same way, the input vector is the known vector of the target triangular linear equations, and the triangular block matrix, the input vector block and the task block correspond to each other through the number after division; the second determination unit is used to use the computing core to iteratively solve multiple task blocks according to the triangular block inverse matrix corresponding to each task block and the input vector block until the result vector block of each task block is obtained, and then determine that the target input vector is the solution vector of the target triangular linear equations, wherein the triangular block inverse matrix refers to the inverse matrix of the triangular block matrix, and the target input vector is a vector obtained by covering the input vector block corresponding to the task block in the input vector using the result vector block obtained by solving each task block in the solution process.
[0014] Optionally, the first allocation unit includes: a first acquisition module, used to divide the result vector of the target triangular linear equations to obtain a plurality of task blocks; a second acquisition module, used to acquire a plurality of computing cores in the processing platform; a first allocation module, used to allocate the plurality of task blocks to the computing cores in sequence when a first total number of the plurality of task blocks does not exceed a second total number of the plurality of computing cores; and a second allocation module, used to allocate the plurality of task blocks to the computing cores in sequence in a cyclic manner when the first total number of the plurality of task blocks exceeds the second total number of the plurality of computing cores.
[0015] Optionally, the second allocation unit includes: a third acquisition module, used to divide the matrix on the diagonal of the input matrix according to the same division method as the division of the result vector, to obtain a plurality of the triangular block matrices; a fourth acquisition module, used to use the computing core to calculate the inverse matrix of each triangular block matrix, to obtain a plurality of the triangular block inverse matrices; a fifth acquisition module, used to divide the input vector according to the same division method as the division of the result vector, to obtain a plurality of the input vector blocks; a third allocation module, used to allocate a plurality of the triangular block inverse matrices and a plurality of the input vector blocks to the computing core where the corresponding task blocks are located according to the correspondence between the triangular block matrices, the input vector blocks and the task blocks.
[0016] Optionally, the device for solving and optimizing a triangular linear equation system further includes: a first storage module, for storing the multiple triangular block inverse matrices and the multiple input vector blocks in the global memory of the processing platform after acquiring the multiple triangular block inverse matrices and the multiple input vector blocks, wherein the global memory is used to store data whose data volume exceeds a data volume threshold.
[0017] Optionally, the device for solving and optimizing the triangular linear equations further includes: a transfer unit, for iteratively solving the plurality of task blocks using the computing core according to the triangular block inverse matrix and the input vector block corresponding to each of the task blocks, until a result vector block of each of the task blocks is obtained, and before determining that the target input vector is the solution vector of the target triangular linear equations, transferring the plurality of triangular block inverse matrices and the plurality of input vector blocks stored in the global memory to a unified buffer of the processing platform, wherein the unified buffer is used to temporarily store the data to be calculated within a predetermined period of time before the calculation is performed.
[0018] Optionally, the second determination unit includes: a solution module, which is used to execute the solution step, and use the vector calculation unit of the computing core to perform matrix-vector multiplication according to the triangular block inverse matrix corresponding to the task block and the input vector block to solve, so as to obtain the result vector block of the task block; an update module, which is used to execute the update step, and use the result vector block to cover the input vector block corresponding to the task block in the input vector; a judgment module, which is used to execute the judgment step, and judge whether the result vector blocks of all the task blocks have been obtained to obtain a judgment result; a first determination module, which is used to repeatedly execute the solution step, the update step and the judgment step until the judgment result indicates that the result vector blocks of all the task blocks have been obtained, and determine that the input vector currently output is the target input vector; a second determination module, which is used to determine that the target input vector is the solution vector of the target triangular linear equation group.
[0019] Optionally, the device for solving and optimizing the triangular linear equations also includes: a second storage module, which is used to perform matrix-vector multiplication according to the inverse matrix of the triangular block corresponding to the task block and the input vector block to solve the problem using the vector calculation unit of the computing core, and then store the result vector block in the unified buffer of the processing platform after obtaining the result vector block of the task block; and a transfer module, which is used to transfer the result vector block stored in the unified buffer to the global memory of the processing platform before using the result vector block to overwrite the input vector block in the input vector corresponding to the task block.
[0020] According to another aspect of an embodiment of the present invention, a system for optimizing the solution of a triangular linear equation group is further provided, wherein the system for optimizing the solution of a triangular linear equation group uses any one of the above-mentioned methods for optimizing the solution of a triangular linear equation group.
[0021] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is further provided, wherein the computer-readable storage medium includes a stored program, wherein the program executes any one of the above-mentioned methods for optimizing the solution of a triangular linear equation system.
[0022] According to another aspect of an embodiment of the present invention, a processor is further provided, wherein the processor is used to run a program, wherein when the program is run, any one of the above-mentioned methods for optimizing the solution of a triangular linear equation system is executed.
[0023] According to another aspect of an embodiment of the present invention, a computer program product is further provided, comprising computer instructions, wherein when the computer instructions are executed by a processor, any one of the above-mentioned methods for optimizing the solution of a triangular linear equation system is executed.
[0024] In an embodiment of the present invention, a triangular linear equation group to be solved is determined to be a target triangular linear equation group; a plurality of task blocks obtained by dividing a result vector of the target triangular linear equation group are cyclically distributed to a computing core of a processing platform, wherein the result vector is an unknown vector to be solved in the target triangular linear equation group, and the processing platform is used to solve the target triangular linear equation group; the inverse matrix of each triangular block matrix on the diagonal of the input matrix in the target triangular linear equation group and the input vector blocks of the input vector are distributed to the computing core where the corresponding task blocks are located, wherein the triangular block matrix of the input matrix and the input vector blocks of the input vector are both divided in the same way as the result vector is divided, The input matrix is the coefficient matrix of the target triangular linear equations, the input vector is the known vector of the target triangular linear equations, and the triangular block matrix, input vector block and task block correspond to each other through the divided labels; the computing core is used to iteratively solve multiple task blocks according to the triangular block inverse matrix and input vector block corresponding to each task block, until the result vector block of each task block is obtained, and then the target input vector is determined to be the solution vector of the target triangular linear equations, wherein the triangular block inverse matrix refers to the inverse matrix of the triangular block matrix, and the target input vector is a vector obtained by covering the input vector block corresponding to the task block in the input vector with the result vector block obtained by solving each task block in the solution process. Through the above technical scheme, the purpose of solving the triangular linear equations is achieved by dividing the result vector into tasks and cyclically allocating each task block to the computing core of the processing platform, so that the result vector block of each task block is iteratively solved by the computing core, and the technical effect of using the vector computing unit to perform matrix-vector multiplication to solve a result vector block of the triangular linear equations at one time and solving the equations serially iteratively by the computing core is achieved, thereby improving the solution efficiency, and further solving the technical problems of the computing core's transportation waiting time and long calculation time in the process of using traditional methods to solve the triangular linear equations in the related technology, and the low solution efficiency caused by the forward iteration method one by one on the processing platform. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0026] Figure 1 It is a hardware structure block diagram of a mobile terminal of a method for solving and optimizing a triangular linear equation system according to an embodiment of the present invention;
[0027] Figure 2 is a flow chart of a method for solving and optimizing a triangular linear equation system according to an embodiment of the present invention;
[0028] Figure 3 is a flow chart of an optional method for solving and optimizing a triangular linear equation system according to an embodiment of the present invention;
[0029] Figure 4 is a schematic diagram of block division according to an embodiment of the present invention;
[0030] Figure 5 is a schematic diagram of vector block division of solution results according to an embodiment of the present invention;
[0031] Figure 6 is a schematic diagram of solving a target triangular linear equation system according to an embodiment of the present invention;
[0032] Figure 7 is a schematic diagram of a device for solving and optimizing a triangular linear equation system according to an embodiment of the present invention. DETAILED DESCRIPTION
[0033] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0034] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0035] As introduced in the background technology, in the related art, the traditional method is used to solve the triangular linear equations. The waiting time and calculation time of the computing core are too long, and the solutions are performed one by one in a forward iteration manner on the processing platform, resulting in low solution efficiency. In view of the above defects, an optimization method and device for solving triangular linear equations are provided in an embodiment of the present invention.
[0036] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0037] The method embodiments provided in the embodiments of the present invention can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 1 is a hardware structure block diagram of a mobile terminal for a method for solving and optimizing a triangular linear equation system according to an embodiment of the present invention. Figure 1 As shown, the mobile terminal may include one or more ( Figure 1 Only one is shown in the figure) a processor 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data, wherein the mobile terminal may also include a transmission device 106 and an input / output device 108 for communication functions. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the mobile terminal. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations shown.
[0038] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the method for solving and optimizing the triangular linear equations in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, the above method is implemented. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories can be connected to the mobile terminal via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof. The transmission device 106 is used to receive or send data via a network. The above-mentioned specific network example may include a wireless network provided by a communication provider of the mobile terminal. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0039] According to an embodiment of the present invention, a method embodiment of a method for optimizing the solution of a triangular linear equation system is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0040] Figure 2 is a flow chart of a method for solving and optimizing a triangular linear equation system according to an embodiment of the present invention. Figure 2 As shown, the method comprises the following steps:
[0041] Step S202, determining that the triangular linear equation group to be solved is a target triangular linear equation group.
[0042] In this embodiment, it can be determined that the triangular linear equation group to be solved is the target triangular linear equation group, and then the target triangular linear equation group is solved.
[0043] For example, suppose the triangular linear equations to be solved are: A^(1|T)*x=b, where x and b are single-precision vectors, A is a single-precision triangular matrix, x is the unknown vector to be solved (i.e., the result vector), b is the input vector, and A is the input matrix. After solving the target triangular linear equations, the solution vector obtained will overwrite the input vector, that is, the solution vector obtained (i.e., the result vector obtained by the solution) will be stored in the space of the original input vector, so the input vector that is finally covered (which can be understood as the target input vector) is the solution vector of the target triangular linear equations, that is, the result vector obtained by the solution.
[0044] Step S204, divide the result vector of the target triangular linear equation group into multiple task blocks and distribute them to the computing core of the processing platform in a loop, wherein the result vector is an unknown vector to be solved in the target triangular linear equation group, and the processing platform is used to solve the target triangular linear equation group.
[0045] In this embodiment, task division can be performed first to assign the divided task blocks to the computing cores of the processing platform. Here, the result vector must first be divided into multiple task blocks, and the task blocks are assigned to the corresponding computing cores for calculation.
[0046] According to the above embodiment of the present invention, in the above step S204, the multiple task blocks obtained by dividing the result vector of the target triangular linear equations group are cyclically allocated to the computing cores of the processing platform, including: dividing the result vector of the target triangular linear equations group to obtain multiple task blocks; obtaining multiple computing cores in the processing platform; when a first total number of the multiple task blocks does not exceed a second total number of the multiple computing cores, allocating the multiple task blocks to the computing cores in sequence; when the first total number of the multiple task blocks exceeds the second total number of the multiple computing cores, allocating the multiple task blocks to the computing cores in sequence in a cyclic manner.
[0047] Combine the following Figure 3 and Figure 4 The above embodiments of the present invention are described in detail. Figure 3 is a flowchart of an optional method for solving and optimizing a triangular linear equation system according to an embodiment of the present invention, Figure 4 It is a schematic diagram of block division according to an embodiment of the present invention.
[0048] like Figure 3 As shown, first of all, task division is performed, and then the divided computing tasks are assigned to the computing cores of the processing platform; when performing task division, it actually includes the division of result vectors, input matrices and input vectors. The result vector will be divided into multiple task blocks, and the input matrix and input vector will also be divided into a corresponding number of blocks to solve the result vectors of the corresponding task blocks.
[0049] like Figure 4 As shown, the result vector will be divided into multiple task blocks for solving, and the size of a single task block needs to be determined according to the size of the unified buffer space. Assuming that a single block of the result vector is n×1 in size, the data type is single precision, and the memory space size of the unified buffer is K bytes, then each time a result vector block is solved, the matrix-vector multiplication performed needs to store a vector of size n×n and a vector of size n in the unified buffer, then the occupied unified buffer size is (n×n+n)×8 bytes, and the restriction condition for the value of n can be obtained as (n×n+n)×8≤K; in order to meet the design requirements of the hardware architecture of the processing platform, n needs to be rounded down to an integer that is a multiple of 32 bytes (it should satisfy the condition of (n×8)%32=0) to improve the hardware's calculation and data handling efficiency.
[0050] When allocating the divided computing tasks to the computing cores of the processing platform, assuming that the number of computing cores is R and the computing cores are numbered i∈{0,1,2,…,R-1}, the task blocks obtained by dividing the result vector will be cyclically allocated to each computing core, that is, task block 0 will be allocated to computing core 0, task 1 will be allocated to computing core 1, and so on. When there are still task blocks that have not been allocated after being allocated to the last computing core, the computing tasks will be allocated again from computing core 0. The task allocation strategy of cyclic allocation can improve the cache hit rate of the processing platform and improve the efficiency of data transfer. For the task blocks for solving the result vector, since the solution iteration of the next step of solving the triangular linear equations will depend on the data of the previous step, these blocks will be solved serially by computing core 0 in sequence. However, when executing the matrix-vector multiplication of the solution and updating the matrix-vector multiplication of the data to be solved in the next step, all computing cores will still be deployed for parallel calculation to improve the computational efficiency of the matrix-vector multiplication.
[0051] Step S206, distribute the inverse matrices of each triangular block matrix on the diagonal of the input matrix in the target triangular linear equation group and the input vector blocks of the input vector to the computing core where the corresponding task blocks are located, wherein the triangular block matrix of the input matrix and the input vector blocks of the input vector are divided in the same way as the result vector is divided, the input matrix is the coefficient matrix of the target triangular linear equation group, the input vector is the known vector of the target triangular linear equation group, and the triangular block matrix, the input vector blocks and the task blocks correspond to each other through the labels after division.
[0052] In this embodiment, the input matrix and the input vector also need to be divided according to the division method of the result vector. The input vector b and the input matrix A can be divided into tasks according to the determined block size n. Assuming that the size of the result vector is N, the input vector b will be divided into N / n vector blocks (i.e., input vector blocks), and the input matrix A diagonal will also be divided into N / n triangular matrix blocks (i.e., triangular block matrices). The triangular block matrices, input vector blocks and task blocks can be corresponded through the block numbers.
[0053] According to the above embodiment of the present invention, in the above step S206, the inverse matrices of each triangular block matrix on the diagonal of the input matrix in the target triangular linear equation group and the input vector blocks of the input vector are allocated to the computing core where the corresponding task blocks are located, including: dividing the matrix on the diagonal of the input matrix in the same division method as the result vector to obtain multiple triangular block matrices; using the computing core to calculate the inverse matrix of each triangular block matrix to obtain multiple triangular block inverse matrices; dividing the input vector in the same division method as the result vector to obtain multiple input vector blocks; and allocating multiple triangular block inverse matrices and multiple input vector blocks to the computing core where the corresponding task blocks are located according to the correspondence between the triangular block matrices, the input vector blocks and the task blocks.
[0054] As above Figure 4 As shown, each triangular block matrix on the diagonal of the input matrix and each input vector block of the input vector will also be allocated to the computing core of the processing platform in the same way as the task blocks of the result vector are allocated; in addition, when solving the inverse matrix of each triangular block matrix on the diagonal of the input matrix (i.e., the triangular block inverse matrix), this cyclic allocation method is also used to allocate them to the computing core for calculation.
[0055] In a specific embodiment of the present invention, the optimization method for solving the triangular linear equations also includes: after obtaining multiple triangular block inverse matrices and multiple input vector blocks, storing the multiple triangular block inverse matrices and multiple input vector blocks to the global memory of the processing platform, wherein the global memory is used to store data whose data volume exceeds a data volume threshold.
[0056] Specifically, as above Figure 3 As shown, the inverse matrices of each triangular block matrix on the diagonal of the input matrix and each input vector block of the input vector are first stored in the global memory of the processing platform.
[0057] Step S208, using the computing core to iteratively solve multiple task blocks according to the triangular block inverse matrix corresponding to each task block and the input vector block, until the result vector block of each task block is obtained, and then the target input vector is determined to be the solution vector of the target triangular linear equations, wherein the triangular block inverse matrix refers to the inverse matrix of the triangular block matrix, and the target input vector is a vector obtained by overwriting the input vector block corresponding to the task block in the input vector using the result vector block obtained by solving each task block during the solution process.
[0058] In this embodiment, when solving the target triangular linear equations, the inverse matrix of the triangular block matrix on the diagonal is read and the matrix-vector multiplication is performed on the corresponding input vector block to solve the current result vector block, and then it is determined whether there are any unsolved vector blocks. If so, the input vector block is updated using the solved result vector block to solve the result vector block of the next task block. The solution is iterated in this way until the result vector blocks of all task blocks are obtained, and the input vector currently covered by the result vector block (i.e., the target input vector) is determined to be the solution vector of the target triangular linear equations.
[0059] According to the above embodiment of the present invention, before the above step S208, that is, before using the computing core to iteratively solve multiple task blocks according to the triangular block inverse matrix and input vector block corresponding to each task block, until the result vector block of each task block is obtained, and before determining the target input vector as the solution vector of the target triangular linear equation group, the method for optimizing the solution of the triangular linear equation group also includes: transferring the multiple triangular block inverse matrices and the multiple input vector blocks stored in the global memory to a unified buffer of the processing platform, wherein the unified buffer is used to temporarily store the data to be calculated within a predetermined period of time before the calculation.
[0060] Specifically, before solving the target triangular linear equations, it is necessary to first transfer the inverse matrices of each triangular block and each input vector stored in the global memory to the unified buffer for storage. This has the following benefits: 1) Reduce data transfer overhead: On the processing platform, the data transfer overhead from the global memory to the computing core is large. By temporarily storing commonly used data in the unified buffer, the number of data transfers from the global memory to the computing core can be reduced, thereby reducing data transfer overhead; 2) Improve memory access efficiency: The unified buffer has a higher access speed. Temporarily storing data in the unified buffer can improve data access efficiency and speed up calculations; 3) Utilize hardware features: In the computing core design of the processing platform, the unified buffer is specially designed for matrix calculations and has a higher bandwidth. By temporarily storing data in the unified buffer, In a unified buffer, the hardware characteristics can be fully utilized to give full play to the computing performance of the computing core; 4) Realize parallel computing: In the solution process, multiple computing cores are required to participate in the calculation at the same time. Temporarily storing data in a unified buffer can easily realize multiple computing cores to access data in parallel, thereby improving computing efficiency; 5) Optimize computing unit scheduling: Temporarily storing data in a unified buffer can easily schedule data between different computing units, optimize the computing tasks of computing units, and improve resource utilization; 6) Reduce computing delay: By temporarily storing commonly used data in a unified buffer, the time waiting for data loading during the calculation process can be reduced, thereby reducing computing delay; 7) Improve cache hit rate: By temporarily storing data in a unified buffer, the cache hit rate of the computing core can be improved, and the data handling overhead when the cache misses can be reduced.
[0061] According to the above embodiment of the present invention, in the above step S208, a computing core is used to iteratively solve multiple task blocks according to the triangular block inverse matrix and the input vector block corresponding to each task block, until the result vector block of each task block is obtained, and the target input vector is determined to be the solution vector of the target triangular linear equation group, including: a solving step, using the vector calculation unit of the computing core to perform matrix-vector multiplication according to the triangular block inverse matrix and the input vector block corresponding to the task block to solve, and obtain the result vector block of the task block; an updating step, using the result vector block to cover the input vector block corresponding to the task block in the input vector; a judging step, judging whether the result vector blocks of all task blocks have been obtained, and obtaining a judgment result; repeatedly executing the solving step, the updating step and the judging step, until the judgment result indicates that the result vector blocks of all task blocks have been obtained, determining that the input vector currently output is the target input vector; determining that the target input vector is the solution vector of the target triangular linear equation group.
[0062] Combine the following Figure 5 and Figure 6 The above embodiments of the present invention are described in detail. Figure 5is a schematic diagram of the solution result vector block according to an embodiment of the present invention, Figure 6 is a schematic diagram of solving a target triangular linear equation system according to an embodiment of the present invention.
[0063] For solving the triangular linear equations of a single task block, it can be expressed as x=A^(-1)×b. In this step, after the unknown vector x is solved, it will be stored in b, that is, matrix-vector multiplication is performed in situ.
[0064] First, the input vector is loaded from the global memory into the unified buffer by performing address offset according to the task block number of the result vector to be solved. If the task block number of the result vector is 0, the loaded input vector is the original input vector; if the task block number of the result vector is not 0, the loaded input vector is the updated input vector; secondly, similarly, the inverse matrix of the triangular matrix on the diagonal of the corresponding input matrix is loaded from the global memory into the unified buffer by performing address offset according to the task block number of the result vector; thirdly, matrix-vector multiplication is performed using the vector calculation unit for the loaded input vector and the inverse matrix of the triangular matrix on the diagonal of the input matrix to obtain the result of the result vector task block, which is then stored in the global memory of the corresponding input vector; then, it is determined whether the task block of the result vector currently being solved is the last task block. If so, it means that all task blocks have been solved so far, and the calculation process will end; if not, it means that there are still task blocks that have not been solved, and the input vector to be solved next time is updated.
[0065] In order to solve the next result vector task block, it is necessary to update the unsolved input vector according to the vector result of the current result vector task block, that is, to use the result vector block obtained by solving the task block to overwrite the corresponding input vector block in the input vector.
[0066] like Figure 4As shown, first, obtain the column number of the last element on the diagonal of the triangular block matrix on the diagonal of the input matrix (which is also equal to the row number of the last element of the result vector block obtained), such as the column number of the last element on the triangular line of the A0 matrix or the A1 matrix; secondly, determine the size of the input vector to be updated according to the obtained column number. The size of the input vector updated each time may be inconsistent, and the input vector block of the size of one or more result vector blocks may be updated. Assuming that the obtained column number is col, extract the least significant bit of the row number col in binary, then the size of the input vector to be updated next time is 2^the least significant bit obtained × 1, and the matrix block size in the input matrix to be used is 2^the least significant bit obtained × 2^the least significant bit obtained, where 2^the least significant bit obtained = col&((~col)+1); if the result vector block b1 or b2 corresponding to the matrix A2 or the matrix A4 has been solved, then update the input vector block (b2, b3) or b3, and the matrix block of the input matrix to be used is the matrix A3 or the matrix A5; again , move the block matrix and input vector block needed for the update from the global memory to the unified buffer space, use the vector calculation unit to perform matrix-vector multiplication, and then perform vector subtraction, subtract the result of matrix-vector multiplication from the input vector block to obtain the updated vector, and then store the updated vector in the global memory of the input vector block. This operation can ensure that the next time the task block is solved, the inverse matrix of the triangular matrix on the diagonal of the input matrix can be directly read and the corresponding updated input vector block can be used to perform matrix-vector multiplication to obtain the result vector block; after solving the result vector block corresponding to A2, the input vector block (b2, b3) with the size of the A3 matrix dimension will be updated, which ensures that the next solution can directly read the inverse matrix corresponding to A4 and the corresponding updated input vector block b2 to perform matrix-vector multiplication to solve a result vector block, and store the result block in the input vector block b2. After updating the corresponding input vector block, continue to solve the result vector block of the next task block, and iterate until all the result vector blocks are solved.
[0067] In general, the overall process of solving the lower triangular linear equations is as follows Figure 6 As shown, the inverse matrix of the triangular matrix on the diagonal of the input matrix is first solved, and then the inverse matrix of the triangular matrix corresponding to the input matrix is multiplied with the corresponding input vector to solve a result vector block. After that, the input vector block to be used for the next solution is updated, and all the result vector blocks are solved iteratively in sequence.
[0068] According to the above-mentioned embodiment of the present invention, after using the vector calculation unit of the computing core to perform matrix-vector multiplication according to the triangular block inverse matrix corresponding to the task block and the input vector block to solve the problem, and obtaining the result vector block of the task block, the method for optimizing the solution of the triangular linear equations also includes: storing the result vector block to a unified buffer of the processing platform, and before using the result vector block to overwrite the input vector block in the input vector corresponding to the task block, the method for optimizing the solution of the triangular linear equations also includes: transferring the result vector block stored in the unified buffer to the global memory of the processing platform.
[0069] Specifically, the result vector blocks obtained by solving the task blocks using the vector calculation unit are first stored in the unified buffer and then transferred to the global memory. Thereafter, the result vector blocks stored in the global memory are used to overwrite the corresponding input vector blocks in the input vector.
[0070] From the above, it can be seen that through the technical solution provided by the above embodiment of the present invention, it can be determined that the triangular linear equation group to be solved is the target triangular linear equation group; the multiple task blocks obtained by dividing the result vector of the target triangular linear equation group are cyclically distributed to the computing core of the processing platform, wherein the result vector is the unknown vector to be solved in the target triangular linear equation group, and the processing platform is used to solve the target triangular linear equation group; the inverse matrix of each triangular block matrix on the diagonal of the input matrix in the target triangular linear equation group and the input vector blocks of the input vector are distributed to the computing core where the corresponding task blocks are located, wherein the triangular block matrix of the input matrix and the input vector blocks of the input vector are divided according to the same division method as the division of the result vector, the input matrix is the coefficient matrix of the target triangular linear equation group, the input vector is the known vector of the target triangular linear equation group, and the triangular block matrix, the input vector blocks and the task blocks are connected by the divided labels. Numbers are corresponded; a computing core is used to iteratively solve multiple task blocks according to the triangular block inverse matrix and input vector blocks corresponding to each task block, until the result vector block of each task block is obtained, and the target input vector is determined to be the solution vector of the target triangular linear equations, wherein the triangular block inverse matrix refers to the inverse matrix of the triangular block matrix, and the target input vector is a vector obtained by covering the input vector block corresponding to the task block in the input vector using the result vector block solved by each task block in the solution process, thereby achieving the purpose of solving the triangular linear equations by dividing the result vector into tasks and cyclically allocating each task block to the computing core of the processing platform, so as to iteratively solve the result vector block of each task block by using the computing core, and realizing the technical effect of using the vector computing unit to execute matrix-vector multiplication to solve a result vector block of the triangular linear equations at one time, and solving it by the computing core in serial iteration, thereby improving the solution efficiency.
[0071] Therefore, the technical solution provided by the above-mentioned embodiment of the present invention solves the technical problems in the related technology of using traditional methods to solve triangular linear equations, such as the long waiting time for the computing core and the long computing time, and the low solving efficiency caused by solving the equations one by one in a forward iterative manner on the processing platform.
[0072] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0073] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0074] According to an embodiment of the present invention, a triangular linear equations solution optimization device for implementing the triangular linear equations solution optimization method is also provided. Figure 7 is a schematic diagram of a device for solving and optimizing a triangular linear equation system according to an embodiment of the present invention, such as Figure 7 As shown, the device includes: a first determination unit 71, a first allocation unit 73, a second allocation unit 75, and a second determination unit 77. The device for solving and optimizing the triangular linear equations is described in detail below.
[0075] The first determining unit 71 is used to determine that the triangular linear equation group to be solved is a target triangular linear equation group.
[0076] The first allocation unit 73 is used to allocate multiple task blocks obtained by dividing the result vector of the target triangular linear equation group to the computing core of the processing platform in a cyclic manner, wherein the result vector is an unknown vector to be solved in the target triangular linear equation group, and the processing platform is used to solve the target triangular linear equation group.
[0077] The second allocation unit 75 is used to allocate the inverse matrices of each triangular block matrix on the diagonal of the input matrix in the target triangular linear equation group and the input vector blocks of the input vector to the computing core where the corresponding task blocks are located, wherein the triangular block matrix of the input matrix and the input vector blocks of the input vector are divided in the same way as the result vector is divided, the input matrix is the coefficient matrix of the target triangular linear equation group, the input vector is the known vector of the target triangular linear equation group, and the triangular block matrix, the input vector blocks and the task blocks correspond to each other through the divided labels.
[0078] The second determination unit 77 is used to use the computing core to iteratively solve multiple task blocks according to the triangular block inverse matrix and input vector blocks corresponding to each task block, until the result vector block of each task block is obtained, and then determine the target input vector as the solution vector of the target triangular linear equation group, wherein the triangular block inverse matrix refers to the inverse matrix of the triangular block matrix, and the target input vector is a vector obtained by overwriting the input vector block corresponding to the task block in the input vector using the result vector block obtained by solving each task block during the solution process.
[0079] It should be noted here that the above-mentioned first determination unit 71, first allocation unit 73, second allocation unit 75, and second determination unit 77 correspond to steps S202 to S208 in the above-mentioned embodiments, and the four units have the same instances and application scenarios implemented by the corresponding steps, but are not limited to the contents disclosed in the above-mentioned embodiments.
[0080] As can be seen from the above, in the scheme recorded in the above-mentioned embodiment of the present invention, the first determination unit can be used to determine that the triangular linear equation group to be solved is the target triangular linear equation group; then the first allocation unit is used to divide the result vector of the target triangular linear equation group into multiple task blocks, which are cyclically allocated to the computing core of the processing platform, wherein the result vector is the unknown vector to be solved in the target triangular linear equation group, and the processing platform is used to solve the target triangular linear equation group; then the second allocation unit is used to allocate the inverse matrix of each triangular block matrix on the diagonal of the input matrix in the target triangular linear equation group and the input vector blocks of the input vector to the computing core where the corresponding task blocks are located, wherein the triangular block matrix of the input matrix and the input vector blocks of the input vector are both obtained by dividing in the same way as the result vector, the input matrix is the coefficient matrix of the target triangular linear equation group, the input vector is the known vector of the target triangular linear equation group, and the triangular block matrix, the input vector blocks and the task blocks are allocated to the computing core where the corresponding task blocks are located. The corresponding numbers are used; finally, the second determination unit is used to use the computing core to iteratively solve multiple task blocks according to the triangular block inverse matrix corresponding to each task block and the input vector block, until the result vector block of each task block is obtained, and the target input vector is determined to be the solution vector of the target triangular linear equation group, wherein the triangular block inverse matrix refers to the inverse matrix of the triangular block matrix, and the target input vector is a vector obtained by covering the input vector block corresponding to the task block in the input vector using the result vector block solved by each task block in the solution process, thereby achieving the purpose of solving the triangular linear equation group by dividing the result vector into tasks and cyclically allocating each task block to the computing core of the processing platform, so as to iteratively solve the result vector block of each task block using the computing core, and realizing the technical effect of using the vector computing unit to perform matrix-vector multiplication to solve a result vector block of the triangular linear equation group at one time, and solving it serially iteratively by the computing core, thereby improving the solution efficiency.
[0081] Therefore, the technical solution provided by the above-mentioned embodiment of the present invention solves the technical problems in the related technology of using traditional methods to solve triangular linear equations, such as the long waiting time for the computing core and the long computing time, and the low solving efficiency caused by solving the equations one by one in a forward iterative manner on the processing platform.
[0082] Optionally, the first allocation unit includes: a first acquisition module, used to divide the result vector of the target triangular linear equations to obtain multiple task blocks; a second acquisition module, used to acquire multiple computing cores in the processing platform; the first allocation module, used to allocate the multiple task blocks to the computing cores in sequence when the first total number of the multiple task blocks does not exceed the second total number of the multiple computing cores; the second allocation module, used to allocate the multiple task blocks to the computing cores in sequence in a cyclic manner when the first total number of the multiple task blocks exceeds the second total number of the multiple computing cores.
[0083] Optionally, the second allocation unit includes: a third acquisition module, used to divide the matrix on the diagonal of the input matrix in the same division method as the result vector, to obtain multiple triangular block matrices; a fourth acquisition module, used to use the computing core to calculate the inverse matrix of each triangular block matrix, to obtain multiple triangular block inverse matrices; a fifth acquisition module, used to divide the input vector in the same division method as the result vector, to obtain multiple input vector blocks; a third allocation module, used to allocate multiple triangular block inverse matrices and multiple input vector blocks to the computing core where the corresponding task blocks are located according to the correspondence between the triangular block matrices, the input vector blocks and the task blocks.
[0084] Optionally, the device for solving and optimizing the triangular linear equations also includes: a first storage module, used to store the multiple triangular block inverse matrices and the multiple input vector blocks to the global memory of the processing platform after acquiring the multiple triangular block inverse matrices and the multiple input vector blocks, wherein the global memory is used to store data whose data volume exceeds a data volume threshold.
[0085] Optionally, the device for solving and optimizing the triangular linear equations further includes: a transfer unit, which is used to iteratively solve multiple task blocks using a computing core according to the triangular block inverse matrix and input vector block corresponding to each task block, until a result vector block of each task block is obtained, and before determining that the target input vector is the solution vector of the target triangular linear equations, transfers the multiple triangular block inverse matrices and multiple input vector blocks stored in the global memory to a unified buffer of the processing platform, wherein the unified buffer is used to temporarily store the data to be calculated within a predetermined period of time before the calculation is performed.
[0086] Optionally, the second determination unit includes: a solution module, which is used to execute the solution step, and use the vector calculation unit of the computing core to perform matrix-vector multiplication according to the triangular block inverse matrix corresponding to the task block and the input vector block to solve, so as to obtain the result vector block of the task block; an update module, which is used to execute the update step, and use the result vector block to overwrite the input vector block corresponding to the task block in the input vector; a judgment module, which is used to execute the judgment step, and judge whether the result vector blocks of all task blocks have been obtained to obtain the judgment result; a first determination module, which is used to repeatedly execute the solution step, the update step and the judgment step until the judgment result indicates that the result vector blocks of all task blocks have been obtained, and determine that the input vector currently output is the target input vector; a second determination module, which is used to determine that the target input vector is the solution vector of the target triangular linear equation group.
[0087] Optionally, the device for solving and optimizing the triangular linear equations also includes: a second storage module, which is used to use the vector calculation unit of the computing core to perform matrix-vector multiplication according to the triangular block inverse matrix corresponding to the task block and the input vector block to solve the problem, and then store the result vector block in a unified buffer of the processing platform after obtaining the result vector block of the task block; and a transfer module, which is used to transfer the result vector block stored in the unified buffer to the global memory of the processing platform before using the result vector block to overwrite the input vector block in the input vector corresponding to the task block.
[0088] According to another aspect of an embodiment of the present invention, a system for optimizing the solution of a triangular linear equation group is further provided. The system for optimizing the solution of a triangular linear equation group uses any of the above-mentioned methods for optimizing the solution of a triangular linear equation group.
[0089] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is further provided, the computer-readable storage medium comprising a stored program, wherein the program executes any one of the above-mentioned methods for optimizing the solution of a triangular linear equation system.
[0090] Optionally, in this embodiment, the computer-readable storage medium may be located in any one of the computer terminals in a computer terminal group in a computer network, or in any one of the communication devices in a communication device group.
[0091] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: determining that the triangular linear equation group to be solved is a target triangular linear equation group; cyclically distributing the multiple task blocks obtained by dividing the result vector of the target triangular linear equation group to the computing core of the processing platform, wherein the result vector is an unknown vector to be solved in the target triangular linear equation group, and the processing platform is used to solve the target triangular linear equation group; distributing the inverse matrix of each triangular block matrix on the diagonal of the input matrix in the target triangular linear equation group and the input vector blocks of the input vector to the computing core where the corresponding task blocks are located, wherein the triangular block matrix of the input matrix and the input vector blocks of the input vector are divided according to the same method as the result vector. The input matrix is divided in the same way as the target triangular linear equations, the input vector is the known vector of the target triangular linear equations, and the triangular block matrix, the input vector block and the task block correspond to each other through the number after division; the computing core is used to iteratively solve multiple task blocks according to the triangular block inverse matrix and the input vector block corresponding to each task block, until the result vector block of each task block is obtained, and then the target input vector is determined to be the solution vector of the target triangular linear equations, wherein the triangular block inverse matrix refers to the inverse matrix of the triangular block matrix, and the target input vector is a vector obtained by covering the input vector block corresponding to the task block in the input vector with the result vector block obtained by solving each task block in the solution process.
[0092] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: dividing the result vector of the target triangular linear equations to obtain multiple task blocks; acquiring multiple computing cores in the processing platform; allocating the multiple task blocks to the computing cores in sequence when a first total number of the multiple task blocks does not exceed a second total number of the multiple computing cores; and cyclically allocating the multiple task blocks to the computing cores in sequence when a first total number of the multiple task blocks exceeds a second total number of the multiple computing cores.
[0093] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: dividing the matrix on the diagonal of the input matrix in the same division manner as the result vector is divided to obtain multiple triangular block matrices; using the computing core to calculate the inverse matrix of each triangular block matrix to obtain multiple triangular block inverse matrices; dividing the input vector in the same division manner as the result vector is divided to obtain multiple input vector blocks; and allocating multiple triangular block inverse matrices and multiple input vector blocks to the computing core where the corresponding task blocks are located according to the correspondence between the triangular block matrices, the input vector blocks and the task blocks.
[0094] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: after obtaining multiple triangular block inverse matrices and multiple input vector blocks, storing the multiple triangular block inverse matrices and multiple input vector blocks to the global memory of the processing platform, wherein the global memory is used to store data whose data volume exceeds a data volume threshold.
[0095] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: transferring multiple triangular block inverse matrices and multiple input vector blocks stored in the global memory to a unified buffer of the processing platform, wherein the unified buffer is used to temporarily store the data to be calculated within a predetermined period of time before the calculation.
[0096] Optionally, in this embodiment, the computer-readable storage medium is configured to store program codes for executing the following steps: a solving step, using the vector calculation unit of the computing core to perform matrix-vector multiplication according to the triangular block inverse matrix corresponding to the task block and the input vector block to solve, and obtain the result vector block of the task block; an updating step, using the result vector block to overwrite the input vector block corresponding to the task block in the input vector; a judging step, judging whether the result vector blocks of all task blocks have been obtained, and obtaining a judgment result; repeating the solving step, the updating step and the judging step until the judgment result indicates that the result vector blocks of all task blocks have been obtained, determining that the input vector currently output is the target input vector; and determining that the target input vector is the solution vector of the target triangular linear equation group.
[0097] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: after using the vector calculation unit of the computing core to perform matrix-vector multiplication according to the triangular block inverse matrix corresponding to the task block and the input vector block to solve the task block, the result vector block is stored in a unified buffer of the processing platform, and before the result vector block is used to overwrite the input vector block in the input vector corresponding to the task block, the result vector block stored in the unified buffer is transferred to the global memory of the processing platform.
[0098] According to another aspect of an embodiment of the present invention, a processor is further provided, and the processor is used to run a program, wherein when the program is run, any one of the above-mentioned methods for solving and optimizing a triangular linear equation system is executed.
[0099] According to another aspect of an embodiment of the present invention, a computer program product is also provided, comprising computer instructions, which, when executed by a processor, execute any one of the above-mentioned methods for optimizing the solution of a triangular linear equation system.
[0100] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0101] In the above embodiments of the present invention, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0102] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units can be a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0103] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0104] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0105] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk, etc. Various media that can store program codes.
[0106] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A method for optimizing the solution of triangular linear equations, characterized in that: include: Determine the triangular linear equation system to be solved as the target triangular linear equation system; Circularly allocating a plurality of task blocks obtained by dividing the result vector of the target triangular linear equation group to the computing core of the processing platform, wherein the result vector is an unknown vector to be solved in the target triangular linear equation group, and the processing platform is used to solve the target triangular linear equation group; Allocate the inverse matrix of each triangular block matrix on the diagonal of the input matrix in the target triangular linear equation group and the input vector block of the input vector to the computing core where the corresponding task block is located, wherein the triangular block matrix of the input matrix and the input vector block of the input vector are obtained by dividing in the same way as the result vector is divided, the input matrix is the coefficient matrix of the target triangular linear equation group, the input vector is the known vector of the target triangular linear equation group, and the triangular block matrix, the input vector block and the task block correspond to each other through the labels after the division; Utilize the computing core to iteratively solve the multiple task blocks according to the triangular block inverse matrix corresponding to each task block and the input vector block, until the result vector block of each task block is obtained, and then determine the target input vector as the solution vector of the target triangular linear equation group, wherein the triangular block inverse matrix refers to the inverse matrix of the triangular block matrix, and the target input vector is a vector obtained by covering the input vector block corresponding to the task block in the input vector with the result vector block obtained by solving each task block during the solution process; The step of allocating the inverse matrix of each triangular block matrix on the diagonal of the input matrix in the target triangular linear equation group and the input vector blocks of the input vector to the computing core where the corresponding task blocks are located comprises: dividing the input vector in the same division manner as that of dividing the result vector to obtain a plurality of input vector blocks; The method further includes: after acquiring the multiple triangular block inverse matrices and the multiple input vector blocks, storing the multiple triangular block inverse matrices and the multiple input vector blocks in the global memory of the processing platform, wherein the global memory is used to store data whose data volume exceeds a data volume threshold.
2. The method for solving and optimizing a triangular linear equation system according to claim 1, characterized in that: The result vector of the target triangular linear equation system is divided into multiple task blocks and cyclically distributed to the computing core of the processing platform, including: Dividing the result vector of the target triangular linear equation group to obtain a plurality of task blocks; Acquire a plurality of the computing cores in the processing platform; In a case where a first total number of the plurality of task blocks does not exceed a second total number of the plurality of computing cores, allocating the plurality of task blocks to the computing cores in sequence; When the first total number of the plurality of task blocks exceeds the second total number of the plurality of computing cores, the plurality of task blocks are allocated to the computing cores in sequence and in a cyclic manner.
3. The method for solving and optimizing a triangular linear equation system according to claim 1, characterized in that: Allocating the inverse matrix of each triangular block matrix on the diagonal line of the input matrix in the target triangular linear equation group and the input vector block to the computing core where the corresponding task block is located, including: Dividing the matrix on the diagonal of the input matrix in the same division manner as the division of the result vector to obtain a plurality of triangular block matrices; Utilizing the computing core to calculate the inverse matrix of each triangular block matrix, to obtain a plurality of triangular block inverse matrices; According to the corresponding relationship between the triangular block matrix, the input vector block and the task block, a plurality of the triangular block inverse matrices and a plurality of the input vector blocks are allocated to the computing core where the corresponding task block is located.
4. The method for solving and optimizing a triangular linear equation system according to claim 1, characterized in that: After using the computing core to iteratively solve the plurality of task blocks according to the triangular block inverse matrix corresponding to each of the task blocks and the input vector blocks until a result vector block of each of the task blocks is obtained, before determining that the target input vector is the solution vector of the target triangular linear equation group, the method further includes: The plurality of triangular block inverse matrices and the plurality of input vector blocks stored in the global memory are transferred to a unified buffer of the processing platform, wherein the unified buffer is used to temporarily store the data to be calculated within a predetermined period of time before the calculation is performed.
5. The method for solving and optimizing a triangular linear equation system according to claim 1, characterized in that: The computing core is used to iteratively solve the plurality of task blocks according to the triangular block inverse matrix corresponding to each task block and the input vector block until a result vector block of each task block is obtained, and then a target input vector is determined as a solution vector of the target triangular linear equation group, including: A solving step, using the vector calculation unit of the computing core to perform matrix-vector multiplication according to the triangular block inverse matrix corresponding to the task block and the input vector block to obtain the result vector block of the task block; An updating step, using the result vector block to cover the input vector block corresponding to the task block in the input vector; A judgment step, judging whether the result vector blocks of all the task blocks have been obtained, and obtaining a judgment result; Repeating the solving step, the updating step and the judging step until the judging result indicates that the result vector blocks of all the task blocks have been obtained, and determining that the input vector currently outputted is the target input vector; The target input vector is determined to be the solution vector of the target triangular linear equation system.
6. The method for solving and optimizing a triangular linear equation system according to claim 5, characterized in that: After using the vector calculation unit of the computing core to perform matrix-vector multiplication according to the inverse matrix of the triangular block corresponding to the task block and the input vector block to obtain the result vector block of the task block, the method further includes: storing the result vector block in a unified buffer of the processing platform, Before overwriting the input vector block corresponding to the task block in the input vector with the result vector block, the method further includes: transferring the result vector block stored in the unified buffer to the global memory of the processing platform.
7. A device for solving and optimizing triangular linear equations, characterized in that: include: A first determining unit, used for determining the triangular linear equation group to be solved as a target triangular linear equation group; A first allocation unit is used to allocate multiple task blocks obtained by dividing the result vector of the target triangular linear equation group to the computing core of the processing platform in a cyclic manner, wherein the result vector is an unknown vector to be solved in the target triangular linear equation group, and the processing platform is used to solve the target triangular linear equation group; A second allocation unit is used to allocate the inverse matrix of each triangular block matrix on the diagonal of the input matrix in the target triangular linear equation group and the input vector block of the input vector to the computing core where the corresponding task block is located, wherein the triangular block matrix of the input matrix and the input vector block of the input vector are obtained by dividing in the same way as the result vector is divided, the input matrix is the coefficient matrix of the target triangular linear equation group, the input vector is the known vector of the target triangular linear equation group, and the triangular block matrix, the input vector block and the task block correspond to each other through the labels after division; A second determination unit is used to use the computing core to iteratively solve the multiple task blocks according to the triangular block inverse matrix corresponding to each of the task blocks and the input vector block, until the result vector block of each of the task blocks is obtained, and then determine that the target input vector is the solution vector of the target triangular linear equation group, wherein the triangular block inverse matrix refers to the inverse matrix of the triangular block matrix, and the target input vector is a vector obtained by covering the input vector block corresponding to the task block in the input vector using the result vector block obtained by solving each of the task blocks during the solution process; The second allocation unit includes: a fifth acquisition module, configured to divide the input vector in the same division manner as the division of the result vector, to obtain a plurality of input vector blocks; The device also includes: a first storage module, which is used to store the multiple triangular block inverse matrices and the multiple input vector blocks in the global memory of the processing platform after acquiring the multiple triangular block inverse matrices and the multiple input vector blocks, wherein the global memory is used to store data whose data volume exceeds a data volume threshold.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the program executes the method for solving the triangular linear equation system as described in any one of claims 1 to 6.
9. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by the processor, the method for solving and optimizing the triangular linear equations described in any one of claims 1 to 6 is performed.
Citation Information
Patent Citations
Algebra system solution method and system based on KNL platform
CN106897163A
Implementation method for matrix decomposition and lower triangular matrix inversion
CN114996649A