Matrix processing method, matrix processing device and computing equipment
By decomposing the matrix into a form containing diagonal blocks and non-diagonal blocks, and using the inverted relationship of sub-matrix for elimination, the problem of low solution efficiency of sparse linear equation systems is solved, and the efficiency improvement of the matrix decomposition process is achieved.
Patent Information
- Application Number
- CN202311640644.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-30
- Publication Date
- 2025-05-30
AI Technical Summary
The existing technology has low solution efficiency when dealing with large-scale sparse linear equation systems, resulting in bottlenecks in the scale and speed of problem processing, affecting core competitiveness.
By decomposing the matrix into a matrix containing two diagonal blocks and two non-diagonal blocks, and using the sub-matrixes in non-diagonal blocks to perform elimination operations and converting them into a more simplified matrix form, reducing the computational complexity.
It effectively reduces the redundant information in the matrix decomposition process, reduces the computational complexity, and improves the efficiency of matrix decomposition.
Smart Images

Figure CN120067508A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a method for processing a matrix, a matrix processing device, and a computing device. Background Art
[0002] Complex physical, chemical, biological, economic and other problems in nature can be described by ordinary differential equations and partial differential equations. By numerically discretizing differential equations such as finite difference and finite element, a linear equation system of Ax = b is obtained, where A is often a large-scale sparse matrix, with the scale ranging from hundreds of thousands to billions. Generally speaking, the solution of sparse linear equation systems (i.e., linear equation systems where A is a sparse matrix) is the basis and core of industrial fields such as computer aided engineering (CAE), electronic design automation (EDA), or computational fluid dynamics (CFD). To a certain extent, the ability to solve sparse linear equation systems determines the scale and speed of problems that industrial software can handle, and thus determines the core competitiveness of industrial software.
[0003] In view of this, how to improve the ability of industrial software to solve sparse linear equation systems is an urgent problem to be solved currently. Summary of the Invention
[0004] This application provides a method for processing a matrix, a matrix processing device, and a computing device, which are used to improve the efficiency of matrix decomposition.
[0005] In a first aspect, this application provides a method for processing a matrix. Obtain a first matrix, where the first matrix includes two diagonal blocks and two non-diagonal blocks. After the first matrix is block-partitioned, each non-diagonal block can be divided into two sub-matrices (a first sub-matrix and a second sub-matrix). Due to the characteristics of the matrix in the elliptic property equation, the two non-diagonal blocks in the first matrix are in an inverted relationship with each other, so the first sub-matrices of the two non-diagonal blocks are in an inverted relationship with each other, and the second sub-matrices of the two non-diagonal blocks are in an inverted relationship with each other.
[0006] In the non-diagonal block, there is redundant information. For the two sub-matrices (a first sub-matrix and a second sub-matrix) of the same non-diagonal block, one sub-matrix in the non-diagonal block can be represented by the other sub-matrix in the non-diagonal block and the elimination coefficient. Therefore, in the embodiments of this application, the two first sub-matrices of the two non-diagonal blocks in the first matrix need to be eliminated, and then the two first sub-matrices in the first matrix will be zeroed. Specifically, through the second sub-matrix of a non-diagonal block in the first matrix, the first sub-matrix in the same non-diagonal block can be eliminated.
[0007] In each non - diagonal block, after the first sub - matrix is eliminated by the second sub - matrix, the first matrix will be transformed into a second matrix and a first elimination matrix. That is, the product of the second matrix and the first elimination matrix can obtain the first matrix. In other words, the second matrix and the first elimination matrix are the inverse matrices of the first matrix.
[0008] Next, perform LU decomposition on the second matrix to obtain a third matrix and a first decomposition matrix. Then, the third matrix and the first decomposition matrix are the inverse matrices of the second matrix. In the embodiments of the present application, first, eliminate two first sub - matrices from the first matrix to obtain the second matrix. Compared with the first matrix, the second matrix has reduced redundant information, thus also reducing the computational complexity of performing LU decomposition based on the second matrix and reducing the computational complexity in the matrix decomposition process. Further, since the second matrix does not contain the information of the first sub - matrix, for the LU decomposition of the second matrix, some sub - matrices in the diagonal blocks of the second matrix can also be eliminated, thereby further reducing the redundant information in the third matrix.
[0009] Based on the first aspect, in an alternative embodiment, after the first matrix is block - partitioned, 9 sub - matrices arranged in three rows and three columns are obtained. The two non - diagonal blocks are the first non - diagonal block and the second non - diagonal block. The first sub - matrix in the first non - diagonal block is the sub - matrix in the third row and the first column, the second sub - matrix in the first non - diagonal block is the sub - matrix in the third row and the second column, the first sub - matrix in the second non - diagonal block is the sub - matrix in the first row and the third column, and the second sub - matrix in the second non - diagonal block is the sub - matrix in the second row and the third column. In practical applications, the first matrix includes several zero elements and multiple non - zero elements. After the first matrix is block - partitioned, 9 sub - matrices arranged in three rows and three columns are obtained. The embodiments of the present application do not limit the block - partition layout of the first matrix and the positions of each element (zero element or non - zero element). Essentially, through matrix block - partitioning, the first matrix is divided into 9 regions (sub - matrices) arranged in three rows and three columns.
[0010] Based on the first aspect, in an alternative embodiment, during the process of transforming the first matrix into the second matrix and the first elimination matrix, first, calculate the elimination coefficient of the first sub - matrix relative to the second sub - matrix in the same non - diagonal block. Then, the first sub - matrix can be represented as the product of the second sub - matrix and the elimination sparsity. Since the two non - diagonal blocks in the first matrix are in an inverted relationship, the first sub - matrices between the two non - diagonal blocks are in an inverted relationship, and the second sub - matrices between the two non - diagonal blocks are in an inverted relationship, the elimination coefficients used to eliminate the first sub - matrix in the two non - diagonal blocks are the same. Therefore, the first elimination matrix can be generated according to the elimination coefficient. By processing the first matrix with the first elimination matrix, the second matrix can be obtained. At this time, the two first sub - matrices originally in the first matrix have been eliminated (set to zero) in the second matrix.
[0011] Based on the first aspect, in an alternative implementation, a precision threshold for eliminating elements of the first sub-matrix can be preset in advance. According to the preset precision threshold, an approximate elimination coefficient of the first sub-matrix relative to the second sub-matrix can be obtained. Then, the approximate first sub-matrix can be obtained by multiplying the second sub-matrix by the approximate elimination sparse matrix. Among them, the larger the precision threshold, the lower the required computational amount, but the lower the decomposition precision of the first sub-matrix; the smaller the precision threshold, the higher the required computational amount, but the higher the decomposition precision of the first sub-matrix. In practical applications, the size of the precision threshold can be determined according to the scenario requirements or computing power configuration, so as to make a trade-off between the computational amount and the decomposition precision and improve the efficiency of matrix decomposition.
[0012] Based on the first aspect, in an alternative implementation, continue to perform LU decomposition on the third matrix, and then the two second sub-matrices in the original first matrix can be eliminated to obtain the fourth matrix and the second decomposition matrix. Thus, the first matrix is converted into the fourth matrix, the first elimination matrix, the first decomposition matrix, and the second decomposition matrix, further reducing the redundant information in the fourth matrix.
[0013] Based on the first aspect, in an alternative implementation, after decomposing the first matrix into the fourth matrix, the first elimination matrix, the first decomposition matrix, and the second decomposition matrix, the Schur complement matrix in the fourth matrix can be used for the decomposition task of the parent node of the first matrix. Specifically, obtain the Schur complement matrix in the fourth matrix, and transfer the Schur complement matrix in the fourth matrix to the decomposition task of the parent node of the first matrix. Then, according to the Schur complement matrix in the fourth matrix, generate the fifth matrix. The decomposition task of the parent node of the first matrix is the decomposition task for the fifth matrix. Since the redundant information of the fourth matrix has been greatly reduced after the first matrix is decomposed into the fourth matrix, transferring the Schur complement matrix in the fourth matrix to the decomposition task of the parent node (the fifth matrix) of the first matrix reduces the degrees of freedom of the parent node of the first matrix, thereby reducing the computational complexity of the decomposition task of the fifth matrix.
[0014] Based on the first aspect, in an alternative embodiment, after decomposing the first matrix into a third matrix, a first elimination matrix, and a first decomposition matrix, the Schur complement matrix in the fourth matrix can be used for the decomposition task of the parent node of the first matrix. Specifically, obtain the Schur complement matrix in the third matrix, and transfer the Schur complement matrix in the third matrix to the decomposition task of the parent node of the first matrix, so as to generate a sixth matrix according to the Schur complement matrix in the three matrices. Then, the decomposition task of the parent node of the first matrix is the decomposition task for the sixth matrix. Since the redundant information of the third matrix has been greatly reduced after the first matrix is decomposed into the third matrix, transferring the Schur complement matrix in the third matrix to the decomposition task of the parent node (the sixth matrix) of the first matrix reduces the degree of freedom of the parent node of the first matrix, thereby reducing the computational complexity of the decomposition task of the sixth matrix.
[0015] Based on the first aspect, in an alternative embodiment, similar to eliminating the first sub-matrix in the first matrix, in the embodiments of the present application, the first sub-matrix in the sixth matrix can be continuously eliminated to further reduce the computational complexity of the decomposition task of the sixth matrix. The specific decomposition process of the sixth matrix is similar to the decomposition process of eliminating the first sub-matrix in the first matrix. The sixth matrix can be divided into two diagonal blocks and two non-diagonal blocks. After the sixth matrix is block-partitioned, each non-diagonal block of the sixth matrix can be divided into two sub-matrices (a first sub-matrix and a second sub-matrix). Due to the characteristics of the matrix in the elliptic property equation, the two non-diagonal blocks in the sixth matrix are in an inverted relationship with each other, so the first sub-matrices of the two non-diagonal blocks are in an inverted relationship with each other, and the second sub-matrices of the two non-diagonal blocks are in an inverted relationship with each other. Then, based on the second sub-matrix in the sixth matrix, the first sub-matrix in the sixth matrix is eliminated to obtain a seventh matrix and a second elimination matrix, and the LU decomposition is performed on the seventh matrix to obtain an eighth matrix and a third decomposition matrix. Thus, the sixth matrix is decomposed into an eighth matrix, a second elimination matrix, and a third decomposition matrix.
[0016] Based on the first aspect, in an alternative embodiment, the sixth matrix can be decomposed using a traditional existing numerical decomposition process, that is, the LU decomposition is performed on the sixth matrix to obtain a ninth matrix.
[0017] Based on the first aspect, in an alternative embodiment, in order to improve the efficiency of numerically decomposing a matrix, for the decomposition scenario of a sparse matrix, the sparse matrix can be matrix rearranged and symbolically decomposed to obtain an elimination tree corresponding to the sparse matrix. The elimination tree includes multiple levels of tree nodes, and each tree node corresponds to a dense matrix in the sparse matrix. The matrix processing method in the embodiments of the present application can obtain the dense matrix corresponding to a certain tree node as the first matrix. Alternatively, the matrix processing method in the embodiments of the present application can also directly decompose an existing dense matrix block.
[0018] Based on the first aspect, in an optional implementation, before generating the elimination tree corresponding to the sparse matrix, the discrete grid information corresponding to the sparse matrix can be obtained first, and the discrete grid information indicates the positions of the non-zero elements in the sparse matrix. Then, in the embodiments of the present application, it is not necessary to generate the adjacency graph corresponding to the sparse matrix, thereby reducing the computational overhead of generating the adjacency graph. For the matrix rearrangement and symbolic factorization processes of the sparse matrix, they are executed based on the discrete grid information of the sparse matrix, improving the processing efficiency and quality of matrix rearrangement for the sparse matrix, reducing the computational amount in the numerical factorization stage, and improving the solution speed.
[0019] Based on the first aspect, in an optional implementation, the decomposition result of the sparse matrix obtained by the matrix processing method of the present application can be further used as the preconditioner of the sparse matrix in the iterative method, and the preconditioner includes the above-mentioned third matrix, the first elimination matrix, and the first factorization matrix.
[0020] In a second aspect, the present application provides a matrix processing device, including:
[0021] An acquisition unit, configured to acquire a first matrix, where the first matrix includes two diagonal blocks and two non-diagonal blocks, and each non-diagonal block includes a first sub-matrix and a second sub-matrix, and the first sub-matrices between the two non-diagonal blocks are in an inverted relationship with each other, and the second sub-matrices between the two non-diagonal blocks are in an inverted relationship with each other;
[0022] A processing unit, configured to eliminate the first sub-matrix in the first matrix based on the second sub-matrix in the first matrix to obtain a second matrix and a first elimination matrix;
[0023] The processing unit is further configured to perform LU factorization on the second matrix to obtain a third matrix and a first factorization matrix.
[0024] Based on the second aspect, in an optional implementation, the first matrix includes 9 sub-matrices arranged in three rows and three columns, the two non-diagonal blocks are the first non-diagonal block and the second non-diagonal block, the first sub-matrix in the first non-diagonal block is the sub-matrix in the third row and the first column, the second sub-matrix in the first non-diagonal block is the sub-matrix in the third row and the second column, the first sub-matrix in the second non-diagonal block is the sub-matrix in the first row and the third column, and the second sub-matrix in the second non-diagonal block is the sub-matrix in the second row and the third column.
[0025] Based on the second aspect, in an optional implementation, the processing unit is specifically configured to:
[0026] Determine the elimination coefficient of the first sub-matrix relative to the second sub-matrix in the same non-diagonal block of the first matrix;
[0027] Generate a first elimination matrix according to the elimination coefficient;
[0028] Process the first matrix with a first elimination matrix to obtain a second matrix.
[0029] Based on the second aspect, in an alternative implementation, the processing unit is specifically configured to:
[0030] Obtain a precision threshold for a first sub-matrix in the first matrix;
[0031] Based on the precision threshold, determine an elimination coefficient of the first sub-matrix in the first matrix relative to a second sub-matrix in the first matrix.
[0032] Based on the second aspect, in an alternative implementation, the processing unit is further configured to:
[0033] Perform LU decomposition on a third matrix to obtain a fourth matrix and a second decomposition matrix.
[0034] Based on the second aspect, in an alternative implementation, the processing unit is further configured to:
[0035] Obtain a Schur complement matrix in the fourth matrix;
[0036] Generate a fifth matrix according to the Schur complement matrix in the fourth matrix, where the fifth matrix corresponds to another tree node in the elimination tree of the sparse matrix, and the tree node corresponding to the fourth matrix is the parent node of the tree node of the first matrix.
[0037] Based on the second aspect, in an alternative implementation, the processing unit is further configured to:
[0038] Obtain a Schur complement matrix in the third matrix;
[0039] Generate a sixth matrix according to the Schur complement matrix in the third matrix, where the sixth matrix corresponds to another tree node in the elimination tree of the sparse matrix, and the tree node corresponding to the sixth matrix is the parent node of the tree node of the first matrix.
[0040] Based on the second aspect, in an alternative implementation, the processing unit is further configured to:
[0041] Divide the sixth matrix into two diagonal blocks and two non-diagonal blocks, where each non-diagonal block includes a first sub-matrix and a second sub-matrix, the first sub-matrices between the two non-diagonal blocks are in an inverted relationship with each other, and the second sub-matrices between the two non-diagonal blocks are in an inverted relationship with each other;
[0042] Eliminate the first sub-matrix in the sixth matrix based on the second sub-matrix in the sixth matrix to obtain a seventh matrix and a second elimination matrix;
[0043] Perform LU decomposition on the seventh matrix to obtain an eighth matrix and a third decomposition matrix.
[0044] Based on the second aspect, in an optional implementation, the processing unit is further configured to:
[0045] Perform LU decomposition on the sixth matrix to obtain a ninth matrix.
[0046] Based on the second aspect, in an optional implementation, the processing unit is further configured to:
[0047] Obtain a sparse matrix;
[0048] Generate an elimination tree for the sparse matrix, where the elimination tree includes multiple levels of tree nodes, and each tree node corresponds to a dense matrix in the sparse matrix.
[0049] Based on the second aspect, in an optional implementation, the processing unit is specifically configured to:
[0050] Obtain discrete grid information corresponding to the sparse matrix, where the discrete grid information indicates the positions of non-zero elements in the sparse matrix;
[0051] Generate an elimination tree corresponding to the sparse matrix according to the discrete grid information.
[0052] Based on the second aspect, in an optional implementation, the processing unit is further configured to:
[0053] Determine a preconditioner corresponding to the sparse matrix, where the preconditioner includes a third matrix, a first elimination matrix, and a first decomposition matrix.
[0054] The content such as the information interaction and execution process of the embodiments shown in this aspect is based on the same concept as the embodiments shown in the first aspect. Therefore, for the description of the beneficial effects shown in this aspect, please refer to the above-mentioned first aspect, and details are not repeated here.
[0055] In a third aspect, an embodiment of the present application provides a computing device, including: a processor, the processor is coupled to a memory, and the memory is used to store instructions. When the instructions are executed by the processor, the computing device implements the method in the above-mentioned first aspect or any possible implementation manner of the first aspect.
[0056] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which instructions are stored. When the instructions are executed, the computer executes the method in the above-mentioned first aspect, or any possible implementation manner of the first aspect.
[0057] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes computer program code. When the computer program code runs on a computer, the computer executes the method in the above-mentioned first aspect, or any possible implementation manner of the first aspect.
[0058] Sixth aspect, an embodiment of the present application provides a chip, including: a processor, the processor is coupled to a memory, and the memory is used to store instructions. When the instructions are executed by the processor, the chip implements the method in the above first aspect or any possible implementation manner of the first aspect. Description of the Drawings
[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0060] Figure 1 It is a flowchart of a sparse direct method for sparse linear equations;
[0061] Figure 2 It is another flowchart of a sparse direct method for sparse linear equations;
[0062] Figure 3 It is a flowchart of the numerical decomposition stage in the block low-rank algorithm;
[0063] Figure 4 It is a flowchart of the matrix processing method in the embodiment of the present application;
[0064] Figure 5 It is a flowchart of the numerical decomposition stage in the present application;
[0065] Figure 6 It is a schematic diagram of experimental result comparison of the matrix processing method in the embodiment of the present application;
[0066] Figure 7 It is a schematic diagram of experimental result comparison of the matrix processing method in the embodiment of the present application;
[0067] Figure 8 It is a schematic structural diagram of a matrix processing device provided in the embodiment of the present application;
[0068] Figure 9 It is a schematic logical structure diagram of a computing device provided in the embodiment of the present application. Detailed Embodiments
[0069] Embodiments of the present application provide a method for processing a matrix, a matrix processing device, and a computing device, which are used to improve the efficiency of matrix decomposition.
[0070] The embodiments of the present application will be described below in conjunction with the accompanying drawings in the embodiments of the present application. The terms used in the embodiments of the present application are only used to explain the specific embodiments of the present application, rather than to limit the embodiments of the present application. As is known to those of ordinary skill in the art, with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0071] In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. "At least one (item)" or similar expressions thereof refer to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0072] In the description of the present application, the terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above - mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here, for example, can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non - exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0073] A sparse matrix refers to a matrix in which most of the elements are zero elements. Specifically, if the number of elements with a value of 0 in the matrix is much larger than the number of non - zero elements, then the matrix is called a sparse matrix; on the contrary, if the number of non - zero elements accounts for the majority, then the matrix is called a dense matrix. Among them, the ratio of the total number of non - zero elements to the total number of all elements in the matrix is the density of the matrix.
[0074] Complex physical, chemical, biological, economic, and other problems in nature can be described by ordinary differential equations and partial differential equations. By numerically discretizing differential equations using finite differences, finite elements, etc., a linear system of equations Ax = b is obtained, where A is often a large-scale sparse matrix, with the scale ranging from several hundred thousand to several billion. In the fields of scientific computing and engineering technology, such as image processing, meteorology, turbulent flow simulation, astrophysics, or reservoir simulation, etc., the solution of many practical problems can be reduced to solving a certain sparse linear system of equations (i.e., a linear system of equations with A being a sparse matrix).
[0075] Therefore, the solution of sparse linear systems of equations is one of the key technologies in numerical simulation of many practical problems in natural science and social science. For example, in the structural design scenarios of high-rise buildings, bridges, dams, or flood embankments, it is necessary to simulate the deformation and stress conditions; for example, in the scenarios of oil and gas resource exploration and analysis, numerical weather forecasting, or the dynamic analysis of aircraft, it is necessary to use the fluid mechanics equations for simulation; for example, in the scenario of stellar atmosphere analysis, it is often necessary to use laws such as radiative hydrodynamics and particle statistical equilibrium for simulation. When analyzing and simulating these problems, mathematical models are usually established using partial differential equations. In the process of discretely solving partial differential equations, the sparse linear system of equations solving algorithm plays a very important role, and often, it is the performance bottleneck in the entire calculation process.
[0076] Thus, it can be seen that the efficient solution of sparse linear systems of equations is one of the very important topics in computational mathematics and engineering applications. The solution of sparse linear systems of equations is the foundation and core of industrial fields such as computer-aided engineering (CAE), electronic design automation (EDA), or computational fluid dynamics (CFD). To a certain extent, the ability to solve sparse linear systems of equations determines the scale and speed of problems that industrial software can handle, and thus determines the core competitiveness of industrial software.
[0077] Specifically, the methods for solving sparse linear equations mainly include two categories: direct method and iterative method. Among them, the direct method refers to obtaining the exact solution of the sparse linear equations through a finite number of operations including matrix factorization and triangular equation system solving without considering computational rounding errors, so it is also called the exact method; while the iterative method refers to given an initial solution vector, constructing a vector sequence through certain calculations (generally obtaining a series of approximate solutions approaching the exact value through successive iterations), and the limit of the vector sequence is the theoretically exact solution of the equation system. The iterative method transforms the solution of the original problem into an iterative form, and updates the value of the unknown x each time until the required accuracy is achieved. For the iterative method, the construction of a high-quality preconditioner is the key to its fast convergence.
[0078] Next, the currently common methods for solving sparse linear equations will be introduced separately.
[0079] 1. Sparse direct method: The sparse direct method is essentially LU triangular matrix factorization. Please refer to Figure 1 , Figure 1 for a schematic flow chart of the sparse direct method for sparse linear equations. As Figure 1 shown, the sparse direct method of the sparse matrix needs to go through the matrix rearrangement stage, symbolic factorization stage, numerical factorization stage, and triangular equation system solving stage. Among them, the matrix rearrangement stage is used to ensure the sparsity of the matrix during the triangular factorization process by permuting the rows and columns of the matrix; the symbolic factorization stage is used to predict the positions of the non-zero elements of the triangular matrix after factorization, so as to allocate memory space in advance to store the non-zero elements. Symbolic factorization can also analyze the dependency relationships between different levels of numerical factorization tasks, so as to optimize the parallel strategy.
[0080] Please refer to Figure 2 , Figure 2 for another schematic flow chart of the sparse direct method for sparse linear equations. As Figure 2 shown, matrix A, as the input of the sparse direct method, first generates the adjacency graph corresponding to matrix A, and performs graph analysis based on the adjacency graph of matrix A to rearrange matrix A. During the matrix rearrangement process, the rows and columns of the sparse matrix are permuted to maintain the sparsity of the sparse matrix and improve the parallelism of the algorithm in subsequent calculations. Then, in the symbolic factorization stage, the sparse structure after matrix factorization is predicted, storage space is allocated in advance, and the elimination tree corresponding to matrix A is generated. The elimination tree records data dependencies (dependency relationships). The elimination tree includes multiple tree nodes, and each tree node corresponds to a certain level of numerical factorization task in the sparse matrix A. In the elimination tree, there is an edge connection between the parent node and the child node, indicating that there is a data dependency relationship between the numerical factorization task corresponding to the parent node and the numerical factorization task corresponding to the child node. In Figure 2In the numerical factorization stage shown, the numerical factorization task is executed starting from the bottom layer. Each node of the elimination tree corresponds to an LU factorization task of a dense small matrix. Specifically, it is necessary to perform the factorization of the diagonal block, the solution of the non-diagonal block, and the update of the Schur matrix. The output of the numerical factorization stage is several inverse matrices after the factorization of matrix A (for example, the upper triangular matrix and the lower triangular matrix obtained by LU factorization). In the triangular system of equations solving stage, the lower triangular matrix equation Ly = b and the upper triangular matrix equation Ux = y are solved to obtain the solution vector x.
[0081] The computational complexity of the sparse direct method is high and the solution time is long. For example, the time complexity of solving a two-dimensional discrete partial differential equation is O(N 1.5 ), and the space complexity is O(N 4 / 3 ); the time complexity of solving a three-dimensional discrete partial differential equation is O(N·logN), and the space complexity is O(N
[0082] 2. Block Low Rank (BLR) algorithm: The BLR algorithm is a direct sparse matrix solving method based on block low rank matrices. The BLR algorithm optimizes the numerical factorization stage of the sparse direct method on the basis of the aforementioned sparse direct method. Please refer to Figure 3 , Figure 3 for the flow chart of the numerical factorization stage in the block low rank algorithm. As Figure 3 shown, the BLR algorithm partitions the dense matrix that appears in the numerical factorization stage of the sparse direct method, and performs low rank compression on some of the small matrix blocks, thereby simplifying the numerical factorization stage to the LU factorization of the diagonal block, the low rank solution of the non-diagonal block, and the low rank update of the Schur matrix. Through low rank compression, the BLR algorithm reduces the computational complexity of each tree node of the elimination tree, so the computational complexity in many scenarios is lower than that of the sparse direct method.
[0083] However, although the computational complexity of the BLR algorithm is lower than that of the sparse direct method, the time complexity of the BLR algorithm for solving three-dimensional discrete differential equations can only reach O(N 1.33 ) in the best case, which limits the problem scale that the BLR algorithm can solve.
[0084] On the other hand, 3. The sparse direct method and the BLR algorithm start completely from the matrix, decoupled from the actual problem, and cannot fully reduce the time cost of the matrix rearrangement stage.
[0085] 3. Multigrid preconditioner: The construction of the multigrid preconditioner is to apply the multigrid method to matrix A to obtain an approximate matrix-vector multiplication operation of the inverse of A. This approximate matrix-vector multiplication operation can be used as a preconditioner for the iterative solver.
[0086] Multigrid preconditioners are only applicable to equations with strong elliptic properties and will fail for equations with weak elliptic properties.
[0087] 4. Incomplete LU factorization preconditioner: The construction of the incomplete LU factorization preconditioner is to limit the filling of non-zero elements during the triangular factorization of the sparse matrix A. At the cost of factorization accuracy, it improves the factorization speed and reduces memory occupancy, thus obtaining an approximate LU triangular factorization. Then, the back substitution operation of the triangular factorization is used as the preconditioner for the iterative solver. The accuracy of the incomplete LU factorization preconditioner is uncontrollable, resulting in the inability to guarantee the fast convergence of the iterative solver for matrices with relatively poor properties.
[0088] In view of this, the embodiments of the present application provide a method for processing matrices, a matrix processing device, and a computing device, which are used to improve the efficiency of matrix factorization. The matrix processing method of the embodiments of the present application is applicable to the solution of equations with elliptic properties (such as Poisson equations or low wavenumber Helmholtz equations), and is specifically used to decompose and simplify the matrix A (including dense matrices or sparse matrices) in the equations with elliptic properties, thereby reducing the solution complexity of the equations with elliptic properties. Exemplarily, in practical applications, the actual problems in the fields of incompressible fluid, meteorological analysis, seismic analysis (frequency domain acoustic wave and elastic wave simulation), or reservoir analysis can be transformed into the above equations with elliptic properties (Ax = b), so as to decompose the matrix A in the equations with elliptic properties through the matrix processing method of the embodiments of the present application and reduce the solution complexity of the equations.
[0089] The matrix processing method provided by the embodiments of the present application can be integrated into industrial software such as computer-aided engineering (CAE), electronic design automation (EDA), or computational fluid dynamics (CFD) as a solver for equations with elliptic properties in the industrial software. The solver is run by a computing module with specific computing capabilities, taking the matrix A in the equations with elliptic properties as the input of the solver, and the computing module executes the solver to implement the matrix processing method in the embodiments of the present application, so as to output the decomposition result (inverse matrix) of the matrix A.
[0090] Exemplarily, the computing module includes, but is not limited to, a central processing unit (CPU), a graphics processing unit (GPU), or an artificial intelligence (AI) chip. Further, the computing module may also adopt a heterogeneous computing architecture. For example, a CPU+GPU architecture may be adopted, a CPU+AI chip architecture may be adopted, or a CPU+GPU+AI chip architecture may be adopted, etc., which is not specifically limited herein.
[0091] Exemplarily, the matrix processing method in the embodiments of the present application can be applied to the scenario of vehicle autonomous driving. Specifically, a computing module (such as a CPU) for executing the matrix processing method in the embodiments of the present application is configured on the vehicle. The vehicle captures the motion information (including position, attitude, motion speed, or motion direction, etc.) of other vehicles on the road through a camera, sensors, etc. In the scenario of vehicle autonomous driving, it is necessary to predict the future motion trajectory of other vehicles based on the motion information of other vehicles. Thus, the computing module on the vehicle constructs a sparse linear equation (Ax = b) with elliptical properties based on the motion information of other vehicles, where x is the future motion trajectory of other vehicles. The computing module on the vehicle runs a solver of CAE, simplifies and decomposes the sparse matrix A through the matrix processing method in the embodiments of the present application, thereby accelerating the solution of the sparse linear equation (Ax = b) to obtain the future motion trajectory (x) of other vehicles, so that the vehicle can execute the decision of autonomous driving based on the future motion trajectory of other vehicles.
[0092] Next, the matrix processing method provided in the embodiments of the present application will be introduced. Please refer to Figure 4 , Figure 4 which is a schematic flowchart of the matrix processing method in the embodiments of the present application. As Figure 4 shown, the matrix processing method in the embodiments of the present application includes, but is not limited to, steps 101 to 103.
[0093] 101. Obtain a first matrix.
[0094] Obtain a first matrix and decompose the first matrix through the matrix processing method in the embodiments of the present application. Generally speaking, in order to improve the efficiency of numerical decomposition of a matrix, for the decomposition scenario of a sparse matrix, the sparse matrix can be rearranged and sign-decomposed to obtain an elimination tree corresponding to the sparse matrix. The elimination tree includes tree nodes at multiple levels, and each tree node corresponds to a dense matrix in the sparse matrix. The matrix processing method in the embodiments of the present application can obtain the dense matrix corresponding to a certain tree node as the first matrix. Alternatively, the matrix processing method in the embodiments of the present application can also directly decompose a pre-existing dense matrix block.
[0095] In a possible implementation, before generating the elimination tree corresponding to the sparse matrix, the discrete grid information corresponding to the sparse matrix can be obtained first, and the discrete grid information indicates the positions of the non-zero elements in the sparse matrix. Then, in the embodiments of the present application, it is not necessary to generate the adjacency graph corresponding to the sparse matrix, thereby reducing the computational overhead of generating the adjacency graph. For the matrix rearrangement and sign-decomposition processes of the sparse matrix, they are executed based on the discrete grid information of the sparse matrix, improving the processing efficiency and quality of matrix rearrangement of the sparse matrix, reducing the computational amount in the numerical decomposition stage, and improving the solution speed.
[0096] The first matrix includes two diagonal blocks and two non-diagonal blocks. After the first matrix is block-partitioned, each non-diagonal block can be divided into two sub-matrices (the first sub-matrix and the second sub-matrix). Due to the characteristics of the matrix in the elliptic property equation, the two non-diagonal blocks in the first matrix are in an inverted relationship with each other, so the first sub-matrices of the two non-diagonal blocks are in an inverted relationship with each other, and the second sub-matrices of the two non-diagonal blocks are in an inverted relationship with each other. Specifically, after the first matrix is block-partitioned, 9 sub-matrices arranged in three rows and three columns are obtained. The two non-diagonal blocks are the first non-diagonal block and the second non-diagonal block. The first sub-matrix in the first non-diagonal block is the sub-matrix in the third row and the first column, the second sub-matrix in the first non-diagonal block is the sub-matrix in the third row and the second column, the first sub-matrix in the second non-diagonal block is the sub-matrix in the first row and the third column, and the second sub-matrix in the second non-diagonal block is the sub-matrix in the second row and the third column.
[0097] As Figure 4 shown, before the first matrix is block-partitioned, the first matrix is where A pp and A qq are the two diagonal blocks of the first matrix, A qp and are the two non-diagonal blocks of the first matrix, and A qp and are in an inverted relationship with each other. After the first matrix is block-partitioned, is obtained, where A qpAfter matrix partitioning, the (off-diagonal block) is obtained. After matrix partitioning, the (off-diagonal block) is obtained. Off-diagonal block A qp In the (equivalent to the first off-diagonal block), (The first sub-matrix) and the off-diagonal block (Equivalent to the second off-diagonal block), in which (The first sub-matrix) are in an inverted relationship with each other. Off-diagonal block A qp In the (equivalent to the first off-diagonal block), (The second sub-matrix) and the off-diagonal block (Equivalent to the second off-diagonal block), in which (The second sub-matrix) are in an inverted relationship with each other. And the diagonal block A pp After matrix partitioning, it is obtained
[0098] In practical applications, the first matrix includes several zero elements and multiple non-zero elements. After matrix partitioning of the first matrix, 9 sub-matrices arranged in three rows and three columns are obtained. The embodiments of the present application do not limit the partitioning layout of the first matrix and the positions of each element (zero element or non-zero element). Essentially, through matrix partitioning, the first matrix is divided into 9 regions (sub-matrices) in three rows and three columns.
[0099] 102. Eliminate the first sub-matrix in the first matrix to obtain a second matrix and a first elimination matrix.
[0100] In the off-diagonal block, there is redundant information. For two sub-matrices (the first sub-matrix and the second sub-matrix) of the same off-diagonal block, one sub-matrix in the off-diagonal block can be represented by the other sub-matrix in the off-diagonal block and the elimination coefficient. Therefore, in the embodiments of the present application, the two first sub-matrices of the two off-diagonal blocks in the first matrix need to be eliminated, and then the two first sub-matrices in the first matrix will be zeroed. Specifically, through the second sub-matrix of an off-diagonal block in the first matrix, the first sub-matrix in the same off-diagonal block can be eliminated.
[0101] In each off-diagonal block, after eliminating the first sub-matrix through the second sub-matrix, the first matrix will be converted into a second matrix and a first elimination matrix. That is, the product of the second matrix and the first elimination matrix can obtain the first matrix. In other words, the second matrix and the first elimination matrix are the inverse matrices of the first matrix.
[0102] Next, the process of converting the first matrix into the second matrix and the first elimination matrix is introduced. First, the elimination coefficient of the first sub-matrix relative to the second sub-matrix in the same non-diagonal block is calculated, and the first sub-matrix can be represented as the product of the second sub-matrix and the elimination sparsity. Since the two non-diagonal blocks in the first matrix are in an inverted relationship, the first sub-matrices of the two non-diagonal blocks are in an inverted relationship, and the second sub-matrices of the two non-diagonal blocks are in an inverted relationship, the elimination coefficients used to eliminate the first sub-matrix in the two non-diagonal blocks are the same. Therefore, the first elimination matrix can be generated according to the elimination coefficient. By processing the first matrix with the first elimination matrix, the second matrix can be obtained. At this time, the two first sub-matrices in the original first matrix have been eliminated (set to zero) in the second matrix.
[0103] In a possible implementation, a precision threshold for eliminating the first sub-matrix can be preset. According to the preset precision threshold, an approximate elimination coefficient of the first sub-matrix relative to the second sub-matrix can be obtained, and the approximate first sub-matrix can be obtained by multiplying the second sub-matrix by the approximate elimination sparsity. Among them, the larger the precision threshold, the lower the required computational amount, but the lower the decomposition precision of the first sub-matrix; the smaller the precision threshold, the higher the required computational amount, but the higher the decomposition precision of the first sub-matrix. In practical applications, the size of the precision threshold can be determined according to the scenario requirements or computing power configuration to facilitate the trade-off between the computational amount and the decomposition precision and improve the efficiency of matrix decomposition.
[0104] As Figure 4 shown, in the same non-diagonal block, the elimination coefficient of the first sub-matrix relative to the second sub-matrix is calculated to obtain T p , and according to the elimination coefficient T p , the first elimination matrix is obtained. The first elimination matrix includes and Among them, is the inversion of Q p . By processing the first matrix with the first elimination matrix, the second matrix is obtained, which can be specifically expressed as the following formula:
[0105]
[0106] 103. Perform LU decomposition on the second matrix to obtain the third matrix and the first decomposition matrix.
[0107] Next, perform LU decomposition on the second matrix, the third matrix and the first decomposition matrix, and the third matrix and the first decomposition matrix are the inverse matrices of the second matrix. As Figure 4 shown, is the second matrix. Perform LU decomposition on (the second matrix), which is expressed as Among them, and is the first decomposition matrix in the LU decomposition process of the second matrix. Thus, the first matrix is converted into a third matrix, a first elimination matrix, and a first decomposition matrix. In the embodiments of the present application, two first sub-matrices of the first matrix are first eliminated to obtain the second matrix. Compared with the first matrix, the second matrix has less redundant information, thereby reducing the computational complexity of LU decomposition based on the second matrix and reducing the computational complexity in the matrix decomposition process. Further, since the second matrix does not contain the information of the first sub-matrix, for the LU decomposition of the second matrix, some sub-matrices of the diagonal blocks in the second matrix can also be eliminated, thereby further reducing the redundant information in the third matrix. As Figure 4 shown, the two first sub-matrices in the original first matrix have been eliminated in step 102 to obtain the second matrix. After performing LU decomposition on the second matrix, a third matrix is obtained That is, the sub-matrix in the first row and second column and the sub-matrix in the second row and first column of the second matrix are also eliminated in the LU decomposition process, thereby further reducing the redundant information in the third matrix.
[0108] In a possible implementation, continuing to perform LU decomposition on the third matrix can eliminate two second sub-matrices in the original first matrix to obtain a fourth matrix and a second decomposition matrix. Thus, the first matrix is converted into a fourth matrix, a first elimination matrix, a first decomposition matrix, and a second decomposition matrix, further reducing the redundant information in the fourth matrix. Exemplarily, the third matrix after LU decomposition, a fourth matrix is obtained, specifically wherein, the second decomposition matrix includes and
[0109] As can be seen from the above, in the elimination tree corresponding to the sparse matrix, each tree node corresponds to a decomposition task of a dense matrix. In the elimination tree, there is a data dependency relationship between the decomposition tasks of the parent node and the child node, that is, the decomposition task of the dense matrix corresponding to the parent node depends on the decomposition result of the dense matrix corresponding to the child node. Then the first matrix corresponds to a certain tree node in the elimination tree, and the decomposition result corresponding to this tree node is further used for the decomposition task of the parent node of this tree node. Please refer to Figure 5 , Figure 5 which is a schematic flowchart of the numerical decomposition stage in the embodiments of the present application. As Figure 5 shown, from the bottom layer of the tree node (i.e., L = 0, corresponding to matrix A 0=A) Start to execute the decomposition task, and iterate layer by layer from bottom to top towards the parent node. Based on the numerical decomposition process of the traditional sparse direct method, an additional process of matrix elimination (eliminating the first sub-matrix) is added in the embodiments of the present application. In practical applications, the process of matrix elimination (eliminating the first sub-matrix) in the embodiments of the present application can be applied to the decomposition tasks corresponding to all tree nodes except the top vertex in the elimination tree, or it can also be applied to the decomposition tasks corresponding to some tree nodes in the elimination tree, and the decomposition tasks of the remaining tree nodes can continue to adopt the traditional existing numerical decomposition process, which is not limited in the embodiments of the present application.
[0110] In a possible implementation, after decomposing the first matrix into a fourth matrix, a first elimination matrix, a first decomposition matrix, and a second decomposition matrix, the Schur complement matrix in the fourth matrix can be used for the decomposition task of the parent node of the first matrix. Specifically, obtain the Schur complement matrix in the fourth matrix, and transfer the Schur complement matrix in the fourth matrix to the decomposition task of the parent node of the first matrix, so as to generate a fifth matrix according to the Schur complement matrix in the fourth matrix. Then, the decomposition task of the parent node of the first matrix is the decomposition task for the fifth matrix. Since the redundant information of the fourth matrix has been greatly reduced after the first matrix is decomposed into the fourth matrix, transferring the Schur complement matrix in the fourth matrix to the decomposition task of the parent node (the fifth matrix) of the first matrix reduces the degree of freedom of the parent node of the first matrix, thereby reducing the computational complexity of the decomposition task of the fifth matrix. Exemplarily, the obtained fourth matrix is Then the Schur complement matrix in the fourth matrix is
[0111] In a possible implementation, after decomposing the first matrix into a third matrix, a first elimination matrix, and a first decomposition matrix, the Schur complement matrix in the fourth matrix can be used for the decomposition task of the parent node of the first matrix. Specifically, obtain the Schur complement matrix in the third matrix, and transfer the Schur complement matrix in the third matrix to the decomposition task of the parent node of the first matrix, so as to generate a sixth matrix according to the Schur complement matrix in the third matrix. Then, the decomposition task of the parent node of the first matrix is the decomposition task for the sixth matrix. Since the redundant information of the third matrix has been greatly reduced after the first matrix is decomposed into the third matrix, transferring the Schur complement matrix in the third matrix to the decomposition task of the parent node (the sixth matrix) of the first matrix reduces the degree of freedom of the parent node of the first matrix, thereby reducing the computational complexity of the decomposition task of the sixth matrix. Exemplarily, the obtained third matrix is Then the Schur complement matrix in the third matrix is
[0112] Further, similar to the elimination of the first sub-matrix in the first matrix, in the embodiments of the present application, the first sub-matrix in the sixth matrix can be continuously eliminated to further reduce the computational complexity of the decomposition task of the sixth matrix. The specific decomposition process of the sixth matrix is similar to the decomposition process of eliminating the first sub-matrix in the first matrix. The sixth matrix can be divided into two diagonal blocks and two non-diagonal blocks. After the matrix block division of the sixth matrix, each non-diagonal block of the sixth matrix can be divided into two sub-matrices (the first sub-matrix and the second sub-matrix). Due to the characteristics of the matrix in the elliptic property equation, the two non-diagonal blocks in the sixth matrix are in an inverted relationship with each other, so the first sub-matrices of the two non-diagonal blocks are in an inverted relationship with each other, and the second sub-matrices of the two non-diagonal blocks are in an inverted relationship with each other. Then, based on the second sub-matrix in the sixth matrix, the first sub-matrix in the sixth matrix is eliminated to obtain a seventh matrix and a second elimination matrix, and the LU decomposition is performed on the seventh matrix to obtain an eighth matrix and a third decomposition matrix. Thus, the sixth matrix is decomposed into an eighth matrix, a second elimination matrix, and a third decomposition matrix.
[0113] In a possible implementation, the traditional existing numerical decomposition process can be used for the sixth matrix, that is, the LU decomposition is performed on the sixth matrix to obtain a ninth matrix.
[0114] Next, the decomposition task of the tree nodes can be continuously iterated in the above manner, starting from the bottommost tree node to the topmost tree node to complete convergence, and the decomposition result of the sparse matrix is obtained. Then, the decomposition result of the sparse matrix can be further used as a preconditioner for the sparse matrix in the iterative method. The preconditioner includes the above-mentioned third matrix, the first elimination matrix, and the first decomposition matrix.
[0115] Please refer to Figure 6 , Figure 6 which is a schematic diagram of a comparison of experimental results of the matrix processing method in the embodiments of the present application. Taking the Figure 6 shown sparse matrix as the input of the solver, compared with the solution result of the sparse linear equation calculated by the traditional iterative method, the number of iterations of the solution result of the sparse linear equation calculated based on the matrix processing method in the embodiments of the present application is significantly reduced, the accuracy is improved, and the calculation time is shortened. It can be seen that the matrix processing method in the embodiments of the present application can greatly improve the performance of the solver.
[0116] Please refer to Figure 7 , Figure 7 which is a schematic diagram of a comparison of experimental results of the matrix processing method in the embodiments of the present application. As Figure 7 shown, in various industrial scenarios, the matrix processing method in the embodiments of the present application can significantly accelerate the solution time compared with the traditional matrix decomposition method.
[0117] Next, in order to better implement the above solution of the embodiments of the present application, the embodiments of the present application also provide related devices for implementing the above solution. Specifically, please refer to Figure 8 , Figure 8 which is a schematic structural diagram of a matrix processing device provided by an embodiment of the present application. As Figure 8 shown, the matrix processing device includes:
[0118] An acquisition unit 201, configured to acquire a first matrix, where the first matrix includes two diagonal blocks and two non-diagonal blocks, and each non-diagonal block includes a first sub-matrix and a second sub-matrix, and the first sub-matrices between the two non-diagonal blocks are in an inverted relationship with each other, and the second sub-matrices between the two non-diagonal blocks are in an inverted relationship with each other;
[0119] A processing unit 202, configured to eliminate the first sub-matrix in the first matrix based on the second sub-matrix in the first matrix to obtain a second matrix and a first elimination matrix;
[0120] The processing unit 202 is further configured to perform LU decomposition on the second matrix to obtain a third matrix and a first decomposition matrix.
[0121] In a possible implementation, the first matrix includes 9 sub-matrices arranged in three rows and three columns, and the two non-diagonal blocks are a first non-diagonal block and a second non-diagonal block. The first sub-matrix in the first non-diagonal block is the sub-matrix in the third row and the first column, and the second sub-matrix in the first non-diagonal block is the sub-matrix in the third row and the second column. The first sub-matrix in the second non-diagonal block is the sub-matrix in the first row and the third column, and the second sub-matrix in the second non-diagonal block is the sub-matrix in the second row and the third column.
[0122] In a possible implementation, the processing unit 202 is specifically configured to:
[0123] Determine the elimination coefficient of the first sub-matrix relative to the second sub-matrix in the same non-diagonal block of the first matrix;
[0124] Generate a first elimination matrix according to the elimination coefficient;
[0125] Process the first matrix through the first elimination matrix to obtain a second matrix.
[0126] In a possible implementation, the processing unit 202 is specifically configured to:
[0127] Obtain a precision threshold for the first sub-matrix in the first matrix;
[0128] Based on the precision threshold, determine the elimination coefficient of the first sub-matrix in the first matrix relative to the second sub-matrix in the first matrix.
[0129] In a possible implementation, the processing unit 202 is further configured to:
[0130] Perform LU decomposition on the third matrix to obtain a fourth matrix and a second decomposition matrix.
[0131] In a possible implementation, the processing unit 202 is further configured to:
[0132] Obtain a Schur complement matrix in the fourth matrix;
[0133] Generate a fifth matrix according to the Schur complement matrix in the fourth matrix. The fifth matrix corresponds to another tree node in the elimination tree of the sparse matrix, and the tree node corresponding to the fourth matrix is the parent node of the tree node of the first matrix.
[0134] In a possible implementation, the processing unit 202 is further configured to:
[0135] Obtain a Schur complement matrix in the third matrix;
[0136] Generate a sixth matrix according to the Schur complement matrix in the third matrix. The sixth matrix corresponds to another tree node in the elimination tree of the sparse matrix, and the tree node corresponding to the sixth matrix is the parent node of the tree node of the first matrix.
[0137] In a possible implementation, the processing unit 202 is further configured to:
[0138] Divide the sixth matrix into two diagonal blocks and two non-diagonal blocks, where each non-diagonal block includes a first sub-matrix and a second sub-matrix, and the first sub-matrices between the two non-diagonal blocks are in an inverted relationship with each other, and the second sub-matrices between the two non-diagonal blocks are in an inverted relationship with each other;
[0139] Eliminate the first sub-matrix in the sixth matrix based on the second sub-matrix in the sixth matrix to obtain a seventh matrix and a second elimination matrix;
[0140] Perform LU decomposition on the seventh matrix to obtain an eighth matrix and a third decomposition matrix.
[0141] In a possible implementation, the processing unit 202 is further configured to:
[0142] Perform LU decomposition on the sixth matrix to obtain a ninth matrix.
[0143] In a possible implementation, the processing unit 202 is further configured to:
[0144] Obtain the sparse matrix;
[0145] Generate an elimination tree for the sparse matrix. The elimination tree includes multiple levels of tree nodes, and each tree node corresponds to a dense matrix in the sparse matrix.
[0146] In a possible implementation, the processing unit 202 is specifically configured to:
[0147] Obtaining discrete grid information corresponding to the sparse matrix, where the discrete grid information indicates the positions of non-zero elements in the sparse matrix;
[0148] Generate the elimination tree corresponding to the sparse matrix based on the discrete grid information.
[0149] In a possible implementation, the processing unit 202 is further configured to:
[0150] A preconditioner corresponding to the sparse matrix is determined, where the preconditioner includes the third matrix, the first elimination matrix, and the first decomposition matrix.
[0151] It should be noted that the information interaction, execution process, etc. between the modules / units in the matrix processing device are the same as those in the present application. Figure 4 The corresponding method embodiments are based on the same concept. For specific contents, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.
[0152] See also Figure 9 , Figure 9 A schematic diagram of a logical structure of a computing device 30 provided in an embodiment of the present application. The computing device 30 may be deployed with Figure 8 The matrix processing device described in the corresponding embodiment is used to implement Figure 4 Corresponding embodiments: The computer device 30 comprises: a memory 301 , a processor 302 , a communication interface 303 and a bus 304 . The memory 301 , the processor 302 , and the communication interface 303 are connected to each other through the bus 304 .
[0153] The memory 301 may be a read only memory (ROM), a static storage device, a dynamic storage device or a random access memory (RAM). The memory 301 may store a program. When the program stored in the memory 301 is executed by the processor 302, the processor 302 and the communication interface 303 are used to execute steps 101-103 of the above data processing method embodiment.
[0154] The processor 302 may be a central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), a graphics processing unit (GPU), a digital signal processor (DSP), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or any combination thereof, and is used to execute relevant programs to implement one or more of the steps 101-103 in the method embodiment for matrix processing in this application. The steps of the data processing method disclosed in combination with the embodiments of this application may be executed by a compiler and an executor, where the compiler and the executor may be executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read only memory, a programmable read only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 301, and the processor 302 reads the information in the memory 301 and combines its hardware to execute one or more of the steps 101-103 in the method embodiment for matrix processing in this application.
[0155] The communication interface 303 uses a transceiver device such as, but not limited to, a transceiver to implement communication between the computer device 30 and other devices or communication networks.
[0156] The bus 304 can implement a path for transmitting information between various components of the computer device 30 (for example, the memory 301, the processor 302, and the communication interface 303). The bus 304 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 9 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0157] It should be noted that the information interaction, execution process, etc. between the modules / units in the computing device are based on the same concept as the corresponding method embodiment in this application. For specific content, reference may be made to the description in the method embodiment shown above in this application, which will not be elaborated here. Figure 4 Figure 4
[0158] The embodiments of the present application also provide a computer program product containing instructions. The computer program product can be software or a program product containing instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computer device, it causes at least one computer device to execute the method described in the foregoing Figure 4 embodiment as shown.
[0159] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive), etc. The computer-readable storage medium includes instructions that direct the computing device to execute the method described in the foregoing Figure 4 embodiment as shown.
[0160] The communication device provided by the embodiments of the present application can specifically be a chip, and the chip includes: a processing unit and a communication unit. The processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, a pin, or a circuit, etc. The processing unit can execute the computer-executable instructions stored in the storage unit to cause the chip to execute the above Figure 4 embodiment as shown. Optionally, the storage unit is a storage unit inside the chip, such as a register, a cache, etc. The storage unit can also be a storage unit outside the chip within the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0161] It should be further noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Additionally, in the drawings of the device embodiments provided by the embodiments of the present application, the connection relationships between the modules indicate that they have communication connections, which can specifically be implemented as one or more communication buses or signal lines.
[0162] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments of the present application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions accomplished by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be various, such as analog circuits, digital circuits or dedicated circuits, etc. However, for the embodiments of the present application, software programs are more often the better implementation means. Based on such an understanding, the technical solutions of the embodiments of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disc of a computer, etc., and includes several instructions for causing a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0163] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0164] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, training device or data center to another website, computer, training device or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium can be a magnetic medium (for example, floppy disk, hard disk, magnetic tape), an optical medium (for example, DVD), or a semiconductor medium (for example, solid state disk (SSD)), etc.
Claims
1. A method for processing a matrix, characterized in that, it includes: obtaining a first matrix, the first matrix including two diagonal blocks and two non-diagonal blocks, wherein each of the non-diagonal blocks includes a first sub-matrix and a second sub-matrix, and the first sub-matrices between the two non-diagonal blocks are in an inverted relationship with each other, and the second sub-matrices between the two non-diagonal blocks are in an inverted relationship with each other; eliminating the first sub-matrix in the first matrix based on the second sub-matrix in the first matrix to obtain a second matrix and a first elimination matrix; performing LU decomposition on the second matrix to obtain a third matrix and a first decomposition matrix.
2. The method according to claim 1, characterized in that, the first matrix includes 9 sub-matrices arranged in three rows and three columns, the two non-diagonal blocks are a first non-diagonal block and a second non-diagonal block, the first sub-matrix in the first non-diagonal block is the sub-matrix in the third row and the first column, the second sub-matrix in the first non-diagonal block is the sub-matrix in the third row and the second column, the first sub-matrix in the second non-diagonal block is the sub-matrix in the first row and the third column, and the second sub-matrix in the second non-diagonal block is the sub-matrix in the second row and the third column.
3. The method according to claim 1 or 2, characterized in that, eliminating the first sub-matrix in the first matrix based on the second sub-matrix in the first matrix to obtain a second matrix and a first elimination matrix, includes: determining the elimination coefficient of the first sub-matrix relative to the second sub-matrix in the same non-diagonal block of the first matrix; generating a first elimination matrix according to the elimination coefficient; processing the first matrix through the first elimination matrix to obtain a second matrix.
4. The method according to claim 3, characterized in that, determining the elimination coefficient of the first sub-matrix in the first matrix relative to the second sub-matrix in the first matrix, includes: obtaining a precision threshold for the first sub-matrix in the first matrix; determining the elimination coefficient of the first sub-matrix in the first matrix relative to the second sub-matrix in the first matrix based on the precision threshold.
5. The method according to any one of claims 1 to 4, characterized in that, the method further includes: performing LU decomposition on the third matrix to obtain a fourth matrix and a second decomposition matrix.
6. The method according to claim 5, characterized in that, the first matrix corresponds to a tree node in the elimination tree of a sparse matrix, and the method further includes: obtaining a Schur complement matrix in the fourth matrix; generating a fifth matrix according to the Schur complement matrix in the fourth matrix, the fifth matrix corresponding to another tree node in the elimination tree of the sparse matrix, and the tree node corresponding to the fourth matrix is the parent node of the tree node of the first matrix.
7. The method according to any one of claims 1 to 4, characterized in that, the first matrix corresponds to a tree node in the elimination tree of a sparse matrix, and the method further includes: obtaining a Schur complement matrix in the third matrix; Generate a sixth matrix according to the Schur complement matrix in the third matrix, where the sixth matrix corresponds to another tree node in the elimination tree of the sparse matrix, and the tree node corresponding to the sixth matrix is the parent node of the tree node of the first matrix.
8. The method according to claim 7, wherein, the method further includes: Partition the sixth matrix into two diagonal blocks and two non-diagonal blocks, where each non-diagonal block includes a first sub-matrix and a second sub-matrix, and the first sub-matrices of the two non-diagonal blocks are inverse to each other, and the second sub-matrices of the two non-diagonal blocks are inverse to each other; Eliminate the first sub-matrix in the sixth matrix based on the second sub-matrix in the sixth matrix to obtain a seventh matrix and a second elimination matrix; Perform LU decomposition on the seventh matrix to obtain an eighth matrix and a third decomposition matrix.
9. The method according to claim 7, wherein, the method further includes: Perform LU decomposition on the sixth matrix to obtain a ninth matrix.
10. The method according to any one of claims 6 to 9, wherein, the method further includes: Obtain a sparse matrix; Generate an elimination tree for the sparse matrix, where the elimination tree includes multiple levels of tree nodes, and each tree node corresponds to a dense matrix in the sparse matrix.
11. The method according to claim 10, wherein, generating an elimination tree for the sparse matrix includes: Obtain the discrete grid information corresponding to the sparse matrix, where the discrete grid information indicates the positions of the non-zero elements in the sparse matrix; Generate the elimination tree corresponding to the sparse matrix according to the discrete grid information.
12. The method according to claim 10 or 11, wherein, the method further includes: Determine the preconditioner corresponding to the sparse matrix, where the preconditioner includes the third matrix, the first elimination matrix, and the first decomposition matrix.
13. A matrix processing device, wherein, comprising: An acquisition unit for acquiring a first matrix, where the first matrix includes two diagonal blocks and two non-diagonal blocks, where each non-diagonal block includes a first sub-matrix and a second sub-matrix, and the first sub-matrices of the two non-diagonal blocks are inverse to each other, and the second sub-matrices of the two non-diagonal blocks are inverse to each other; A processing unit for eliminating the first sub-matrix in the first matrix based on the second sub-matrix in the first matrix to obtain a second matrix and a first elimination matrix; The processing unit is further configured to perform LU decomposition on the second matrix to obtain a third matrix and a first decomposition matrix.
14. The device according to claim 13, wherein, The first matrix includes 9 sub - matrices arranged in three rows and three columns. The two non - diagonal blocks are the first non - diagonal block and the second non - diagonal block. The first sub - matrix in the first non - diagonal block is the sub - matrix in the third row and the first column. The second sub - matrix in the first non - diagonal block is the sub - matrix in the third row and the second column. The first sub - matrix in the second non - diagonal block is the sub - matrix in the first row and the third column. The second sub - matrix in the second non - diagonal block is the sub - matrix in the second row and the third column.
15. The apparatus according to claim 13 or 14, wherein, the processing unit is specifically configured to: determine the elimination coefficient of the first sub - matrix relative to the second sub - matrix in the same non - diagonal block of the first matrix; generate a first elimination matrix according to the elimination coefficient; process the first matrix through the first elimination matrix to obtain a second matrix.
16. The apparatus according to claim 15, wherein, the processing unit is specifically configured to: acquire a precision threshold for the first sub - matrix in the first matrix; determine the elimination coefficient of the first sub - matrix in the first matrix relative to the second sub - matrix in the first matrix based on the precision threshold.
17. The apparatus according to any one of claims 13 to 16, wherein, the processing unit is further configured to: perform LU decomposition on the third matrix to obtain a fourth matrix and a second decomposition matrix.
18. The apparatus according to claim 17, wherein, the processing unit is further configured to: acquire a Schur complement matrix in the fourth matrix; generate a fifth matrix according to the Schur complement matrix in the fourth matrix. The fifth matrix corresponds to another tree node in the elimination tree of the sparse matrix, and the tree node corresponding to the fourth matrix is the parent node of the tree node of the first matrix.
19. The apparatus according to any one of claims 13 to 16, wherein, the processing unit is further configured to: acquire a Schur complement matrix in the third matrix; generate a sixth matrix according to the Schur complement matrix in the third matrix. The sixth matrix corresponds to another tree node in the elimination tree of the sparse matrix, and the tree node corresponding to the sixth matrix is the parent node of the tree node of the first matrix.
20. The apparatus according to claim 19, wherein, the processing unit is further configured to: divide the sixth matrix into two diagonal blocks and two non - diagonal blocks. Among them, each non - diagonal block includes a first sub - matrix and a second sub - matrix. The first sub - matrices of the two non - diagonal blocks are in an inverted relationship with each other, and the second sub - matrices of the two non - diagonal blocks are in an inverted relationship with each other; eliminate the first sub - matrix in the sixth matrix based on the second sub - matrix in the sixth matrix to obtain a seventh matrix and a second elimination matrix; perform LU decomposition on the seventh matrix to obtain an eighth matrix and a third decomposition matrix.
21. The apparatus according to claim 19, wherein, the processing unit is further configured to: perform LU decomposition on the sixth matrix to obtain a ninth matrix.
22. The apparatus according to any one of claims 18 to 21, It is characterized in that the processing unit is further configured to: obtain a sparse matrix; generate an elimination tree for the sparse matrix, the elimination tree including multiple levels of tree nodes, and each tree node corresponding to a dense matrix in the sparse matrix.
23. The apparatus according to claim 22, it is characterized in that the processing unit is specifically configured to: obtain discrete grid information corresponding to the sparse matrix, the discrete grid information indicating the positions of non-zero elements in the sparse matrix; generate an elimination tree corresponding to the sparse matrix according to the discrete grid information.
24. The apparatus according to claim 22 or 23, it is characterized in that the processing unit is further configured to: determine a preconditioner corresponding to the sparse matrix, the preconditioner including the third matrix, the first elimination matrix, and the first decomposition matrix.
25. A computing device, it is characterized in that it includes a processor, the processor is coupled with a memory, the memory is used for storing instructions; the processor is used for executing the instructions in the memory, so that the apparatus executes the method according to any one of claims 1 to 12.
26. A computer-readable storage medium, it is characterized in that the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 12 is implemented.
27. A computer program product, it is characterized in that the computer program product stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, the method according to any one of claims 1 to 12 is implemented.
Citation Information
Cited By
Power supply network structure weakness detection method based on multi-diagonal-block matrix decomposition
CN120873346A
Aircraft simulation method, device and equipment based on parallel block sparse matrix incomplete LU decomposition and medium
CN121051876A
Parallel solving method of chip dynamic power consumption model based on Cholesky decomposition
CN121279220A