Sparse matrix efficient storage and decomposition device and method for super nodes
By generating the method of eliminating trees and merging nodes to form supernodes, the inefficiency problem of sparse matrix storage format and decomposition algorithm in the prior art is solved, and efficient dense matrix computing and calculation performance improvement is achieved.
Patent Information
- Application Number
- CN202510487851.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The prior art has problems such as inefficient memory efficiency, insufficient computing performance and lack of support for hypernode structures in terms of sparse matrix storage formats and decomposition algorithms, and is particularly difficult to adapt to memory waste caused by irregular sparse structures and dynamic storage adjustments.
The elimination tree is generated through structural analysis, and the continuous nodes with the same L structure are merged to form supernodes. Only effective dense blocks and adjacency lists are stored. It supports direct call of dense BLAS/LAPACK functions such as POTRF/SYTRF, TRSM, and SYRK to achieve efficient dense matrix operations.
It improves computing performance, reduces memory footprint, realizes the efficiency of symmetric sparse matrix decomposition, low memory footprint and algorithm universality, and solves the technical bottleneck of traditional sparse matrix storage format.
Smart Images

Figure CN120011700A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data storage and processing, and in particular to a sparse matrix efficient storage and decomposition device and method for super nodes. Background Art
[0002] At present, in the field of direct solution of sparse matrices, the storage format and decomposition algorithm for symmetric sparse matrices in the existing technology have significant technical limitations. Although COO, CSR, CSC, ELL, BSR and other formats are widely used in specific scenarios, they generally have problems such as low memory efficiency, insufficient computing performance and lack of support for super node structures. For example, the COO format redundantly stores row / column indexes, resulting in excessive memory usage; although CSR and CSC support sparse BLAS libraries, their computational efficiency is significantly lower than that of dense BLAS libraries, and they cannot directly call efficient functions to implement key operations; ELL and BSR formats waste memory due to fixed block size or padding values, and are particularly difficult to adapt to irregular sparse structures; and hybrid formats (such as HYB) attempt to combine the advantages of different formats, but the implementation is complex and the computational overhead increases. In addition, the existing technology has not been optimized for the dense block structure of super nodes, and it is impossible to achieve efficient management of data dependencies by pre-allocating storage space or explicitly recording adjacency lists, resulting in frequent switching between sparse and dense algorithms during the decomposition process, which increases the computational overhead. Although existing patents such as CN119337040A and CN119311221A propose improvement solutions for sparse matrix storage, their core is still focused on general sparse patterns. They fail to design structured storage and decomposition algorithms for super nodes, cannot directly call dense BLAS libraries to improve computing efficiency, and do not solve problems such as memory waste or unclear data transmission paths caused by dynamic storage adjustments. Summary of the invention
[0003] In view of this, the present invention proposes a sparse matrix efficient storage and decomposition device and method for super nodes, which realizes the efficiency and compactness of sparse matrix decomposition of super nodes. The present invention provides the following technical solutions: A sparse matrix efficient storage and decomposition device for super nodes, the device comprising: A matrix preprocessing module, for receiving an original matrix and performing matrix rearrangement thereon to generate an elimination tree; A supernode construction module, used to construct supernodes based on the elimination tree and determine the L matrix of each supernode; Storage format definition module, used to define the symmetric sparse matrix storage format of super nodes; A data loading module, used for loading non-zero elements in the original matrix into the L matrix of the corresponding super node to store the L matrix in a dense format; The matrix decomposition module is used to traverse the dense L matrix and perform matrix decomposition on it.
[0004] Optionally, the matrix preprocessing module specifically includes: A rearrangement unit, used for rearranging the original matrix by using an approximate minimum degree method or a nested partitioning method, so as to move a dense part of the original matrix close to a diagonal line; A node determination unit, used for determining, for each node, the child node that is first eliminated in the decomposition process as a parent node; The elimination tree generation unit is used to traverse the parent node relationship of all nodes and generate an elimination tree starting with a root node, wherein the child nodes of each node constitute its subtree.
[0005] Optionally, the super node construction module specifically includes: A structural analysis unit, for analyzing the row and column distribution of non-zero elements in its subtree based on the parent-child relationship of the elimination tree, so as to determine the non-zero element structure of the node in the L matrix; A supernode merging unit is used to merge adjacent nodes into a supernode if the L structures of these nodes are exactly the same; A parameter definition unit, used to define a starting node number parameter and a length parameter of a supernode; An adjacency list generating unit, configured to generate an adjacency list according to the indexes of all non-zero rows below the supernode in the elimination tree and record the length of the adjacency list; The storage allocation unit is used to allocate storage space to each super node according to a length parameter of the super node and a length of the adjacency list to form an L matrix.
[0006] Optionally, the data loading module specifically includes: A supernode attribution determination unit, used for traversing each column of the original matrix to determine the supernode to which the non-zero element belongs; A non-zero element positioning unit, used to determine the element positioning of the non-zero element in the corresponding supernode to obtain the local coordinates of the non-zero element; A data writing unit is used to write the non-zero element value into the corresponding position of the L matrix of the super node according to the local coordinates.
[0007] Optionally, the matrix decomposition module specifically includes: A traversal control unit, used to process each super node from the top left to the bottom right starting from the root node according to the parent-child relationship of the elimination tree; A decomposition function calling unit, used for calling a decomposition function POTRF or SYTRF to decompose the diagonal block of the current supernode into a lower triangular matrix L1; A matrix updating unit, configured to call a triangular matrix solving function DTRSM to update non-diagonal blocks based on the lower triangular matrix L1 to obtain an L2 matrix; A contribution matrix calculation unit, used for calculating a contribution matrix by calling a matrix multiplication function SYRK based on the L2 matrix; The data transmission unit is used to determine the transmission path and transmission target of the contribution matrix, and copy each column of the contribution matrix to the specified row of the L matrix of the transmission target super node according to the index of the adjacency list.
[0008] The present invention further discloses a method for efficient storage and decomposition of sparse matrices for super nodes, which is applied to the above-mentioned sparse matrix efficient storage and decomposition device for super nodes, and the method comprises: The matrix preprocessing module receives the original matrix and performs matrix rearrangement on it to generate an elimination tree; The supernode construction module constructs supernodes based on the elimination tree and obtains the L matrix of each supernode; The storage format definition module defines the symmetric sparse matrix storage format of the super node; The data loading module loads the non-zero elements in the original matrix into the L matrix of the corresponding supernode to store the L matrix in a dense format; The matrix decomposition module traverses the dense format L matrix and performs matrix decomposition on it.
[0009] Optionally, the method also includes: a supernode attribution determination unit traverses each column of the original matrix to determine the supernode to which the non-zero element belongs; a non-zero element positioning unit determines the element positioning of the non-zero element in the corresponding supernode to obtain the local coordinates of the non-zero element; and a data writing unit writes the non-zero element value into the corresponding position of the L matrix of the supernode according to the local coordinates.
[0010] The present invention further discloses a computer-readable storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the functions of the above-mentioned device are realized.
[0011] The present invention further discloses an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the functions of the above-mentioned device when executing the program.
[0012] The present invention further discloses a computer program product, comprising a computer program, wherein the computer program realizes the functions of the above-mentioned device when executed by a processor.
[0013] According to the technical solution of the present invention, an elimination tree is generated through structural analysis and continuous nodes with the same L structure are merged to form a supernode, only valid dense blocks and adjacency lists are stored, and by pre-allocating L matrix space and explicitly recording the adjacency list, it supports direct calling of dense BLAS / LAPACK functions such as POTRF / SYTRF, TRSM, SYRK, etc., and converts the original discrete sparse operations into efficient dense matrix operations, which solves the problem that the CSR / CSC format depends on the inefficient sparse BLAS library, and the computing performance is greatly improved compared with the traditional method. Furthermore, the continuous storage of supernodes and the contribution matrix transfer path driven by the adjacency list ensure that the data dependency is clear and there is no need to dynamically adjust the storage, ultimately achieving the high efficiency, low memory usage and algorithm versatility of symmetric sparse matrix decomposition. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] For the purpose of illustration and not limitation, the present invention is now described in conjunction with the embodiments of the present invention and the accompanying drawings, in which: Figure 1 It is a structural schematic diagram of a sparse matrix efficient storage and decomposition device for super nodes in an embodiment of the present invention; Figure 2 It is a flowchart of a method for efficient storage and decomposition of sparse matrices of super nodes in an embodiment of the present invention; Figure 3 It is a schematic diagram of the structure of an electronic device in an embodiment of the present invention. DETAILED DESCRIPTION
[0015] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the implementation mode of the present application will be clearly and completely described below in conjunction with the drawings in the implementation mode of the present application. Obviously, the described implementation mode is only a part of the implementation mode of the present application, not all the implementation modes. Based on the implementation mode in the present application, all other implementation modes obtained by ordinary technicians in the field without creative work should fall within the scope of protection of the present application.
[0016] It should be noted that, in the absence of conflict, the embodiments of the present application and the features in the embodiments can be combined with each other. The embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0017] refer to Figure 1 This embodiment discloses a sparse matrix efficient storage and decomposition device for super nodes, the device includes a matrix preprocessing module 21, a super node construction module 22, a storage format definition module 23, a data loading module 24 and a matrix decomposition module 25, which are described below respectively.
[0018] The matrix preprocessing module 21 is used to receive the original matrix and perform matrix rearrangement on it to generate an elimination tree. The matrix preprocessing module 21 specifically includes a rearrangement unit, a node determination unit and an elimination tree generation unit.
[0019] Specifically, the input original matrix is a symmetric sparse matrix, which is input in the compressed sparse column CSC format. The matrix is preprocessed, the starting position of the non-zero elements in each column is recorded, and the global position of all non-zero elements, including the row number, is stored. The original matrix is rearranged by the rearrangement unit in the matrix preprocessing module to reduce the padding so that the dense part of the matrix is close to the diagonal. Specifically, the original matrix is rearranged by the approximate minimum degree method or the nested dissection method, among which the approximate minimum degree method (AMD): is suitable for matrices with no obvious local structure but small-scale dense blocks. Nested dissection method (Nested Dissection): is suitable for matrices with strong locality (such as grid structure). Based on the rearranged matrix, the node determination unit is called to determine the first child node eliminated in the decomposition process as the parent node for each node; then the parent node relationship of all nodes is traversed by the elimination tree generation unit to generate an elimination tree starting with the root node, in which the child nodes of each node constitute its subtree. By rearranging the matrix, we can effectively reduce padding, avoid memory waste of ELL / BSR, and eliminate the tree-driven decomposition order to ensure clear data dependencies to support the subsequent construction of supernodes.
[0020] The supernode construction module 22 is used to construct supernodes based on the elimination tree and determine the L matrix of each supernode. The supernode construction module includes a structure analysis unit, a supernode merging unit, a parameter definition unit and an adjacency list generation unit.
[0021] The structural analysis unit is used to traverse all non-zero elements of the subtree for each node in the rearranged matrix based on the parent-child relationship of the elimination tree, and record its global row and column distribution. The row range is the minimum row number and maximum row number of the non-zero elements in the subtree, and the column range is the column number of the node to the column number of the last continuous node in its subtree. And the total number of non-zero elements in the subtree is counted based on the recorded global row and column distribution. The supernode merging unit is used to compare the L structures of adjacent nodes. If the minimum row number and the maximum row number of the subtree are the same, and the total number of non-zero elements of the subtree is the same, then the adjacent nodes are merged into a supernode. The parameter definition unit is used to define the starting node number parameter and length parameter of the supernode, wherein the starting node number parameter start is the global column number of the first node of the supernode, and the length parameter len is the number of continuous nodes contained in the supernode. An adjacency list generation unit generates an adjacency list and records the length of the adjacency list according to the indexes of all non-zero rows below the super node in the elimination tree; and a storage allocation unit allocates storage space to each super node according to the length parameter of the super node and the length of the adjacency list to form an L matrix.
[0022] Furthermore, for the original matrix, when performing effective data storage, if it is a symmetric positive definite matrix, the L matrix only stores the lower triangular part, and the diagonal elements default to 1. If it is a symmetric non-positive definite matrix, the L matrix stores the lower triangular and diagonal parts, and the diagonal elements store the values of the diagonal matrix D.
[0023] The storage format definition module 23 is used to define the symmetric sparse matrix storage format of the super node.
[0024] Specifically, the storage format definition module 23 includes a structure description definition unit, which is used to define the data structure description of the super node in C language, including: int start; / / The starting node number of the supernode; for example, if the supernode covers columns 3 to 5, start=3.
[0025] int len; / / The length of the supernode; that is, the number of nodes contained in the supernode.
[0026] int adjlen; / / The length of the adjacency list of the supernode; that is, the total number of adjacent non-zero rows under the supernode. For example, when the adjacency list contains rows [5,7,9], adjlen=3.
[0027] int adjlist; / / Supernode adjacency list; that is, the integer array of the adjacency list, storing the global index of the non-zero row below.
[0028] float / double L; / / The first address of the L array of the super node.
[0029] This implementation is based on the data structure description of the supernode defined above, and further defines the storage strategy for the diagonal blocks and non-diagonal blocks of the L matrix. Among them, the matrix position of the diagonal block is the first len×len submatrix. For a symmetric positive definite matrix, only the lower triangular part is stored, and the diagonal elements are 1 by default (no explicit storage is required). For non-diagonal blocks, their position is the last adjlen×len submatrix, and their storage rule is to store the original values of the adjacent non-zero rows below the supernode. To ensure that the storage format of the L matrix is compatible with the BLAS / LAPACK function, the L matrix is stored row-first in this implementation.
[0030] The data loading module 24 specifically includes a supernode affiliation determination unit, a non-zero element location unit and a data writing unit.
[0031] The supernode attribution determination unit is used to load the non-zero elements in the original matrix into the L matrix of the corresponding supernode to store the L matrix in a dense format. The supernode attribution determination unit is used to traverse each column of the original matrix and traverse all supernodes at the same time to determine whether the non-zero elements in the current column belong to a supernode. If so, the non-zero elements are matched to the supernode. The element location of the non-zero element in the corresponding supernode is determined by the non-zero element positioning unit to obtain the local coordinates of the non-zero element. If the global row number of the non-zero element is within the column range of the supernode, the non-zero element is stored in the diagonal block area, that is, the above-mentioned len×len area. If the global row number of the non-zero element is in the adjacency list of the supernode, the element is stored in the non-diagonal block area, that is, the post-adjlen×len part of the above-mentioned L matrix. For non-zero elements whose storage location is in the diagonal block, the global row number of the element is subtracted from the starting node number of the supernode to obtain its local row number. Similarly, the starting node number of the supernode is subtracted from its global column number to obtain its local column number. For non-zero elements whose storage location is in a non-diagonal block, the index adj_idx of the global row number of the element in the adjacency list is determined by binary search, and the calculation method of its local row number is the len of the supernode plus adj_idx, and the local column number of the element is also the global column number minus the starting node number of the supernode. In this way, the element location of the non-zero element in the corresponding supernode is determined to obtain the local coordinates of the non-zero element. Further, the non-zero element value is written into the corresponding position of the L matrix of the supernode according to the local coordinates through the data writing unit to store the L matrix in a dense format.
[0032] The matrix decomposition module 25 specifically includes a traversal control unit, a decomposition function calling unit, a matrix updating unit, a contribution matrix calculating unit and a data transmission unit.
[0033] The matrix decomposition module 25 is used to perform traversal on the dense format L matrix and perform matrix decomposition on it. Specifically, the traversal control unit starts from the root node and processes each super node from the upper left to the lower right according to the parent-child relationship of the elimination tree. It should be understood that here it starts from the root node in the upper left corner of the matrix and gradually processes each super node to the lower right to ensure that the parent node of each super node has been decomposed during the decomposition process.
[0034] For a symmetric positive definite matrix, the decomposition function POTRF is called by the decomposition function calling unit to decompose the diagonal block of the current supernode into L1×L1 T For symmetric non-positive definite matrices, SYTRF is called to decompose the diagonal blocks of the current supernode into L1×D×L1 T Based on the above decomposed matrix, the triangular matrix solver function DTRSM is called by the matrix update unit to update the non-diagonal blocks to obtain the L2 matrix. The update formula is: . Further, through the contribution matrix calculation unit, based on the L2 matrix, the matrix multiplication function SYRK is called to calculate the contribution matrix. The calculation formula of the contribution matrix is: . Further, the transfer path and transfer target of the contribution matrix are determined by the data transfer unit, and each column of the contribution matrix is copied to the specified row of the L matrix of the transfer target supernode according to the index of the adjacency list. Specifically, according to the adjacency list adjlist, each column of CB is associated with the global row number in the adjacency list, and each adjacent row adjlist[i] belongs to the diagonal block or non-diagonal block of the L matrix of a parent node or ancestor supernode. By eliminating the parent-child relationship of the tree or the global row number of the adjacency list, the target supernode is located, and the columns of CB are further copied to the corresponding rows of the target supernode.
[0035] In summary, this embodiment proposes an efficient storage and decomposition device for sparse matrices for supernodes. Through the coordinated design of structural analysis, supernode partitioning, compact storage and efficient computing, the technical bottlenecks of traditional sparse matrix storage formats in memory efficiency, computing acceleration and supernode support are solved. It first performs matrix rearrangement on the original matrix through the matrix preprocessing module 21, generates an elimination tree to optimize filling and clarify node dependencies; then, the supernode construction module 22 analyzes the distribution of non-zero elements of the subtree of each node based on the elimination tree, merges continuous nodes with the same L structure into supernodes, and records the starting position, length, adjacency list and storage space parameters of the supernode; then, the storage format definition module 23 defines the C language data structure of the supernode, and stores the L matrix in a dense format, where the diagonal of a symmetric positive definite matrix defaults to 1, and a non-positive definite matrix explicitly stores the diagonal matrix. The value of the matrix D; then the non-zero elements of the original matrix are loaded into the L matrix of the corresponding supernode through the conversion of global to local coordinates through the data loading module 24, only valid dense blocks are stored and the data dependency is located using the adjacency list; finally, the tree-driven top-down traversal is eliminated through the matrix decomposition module 25, and the diagonal blocks of each supernode are decomposed into L1 by calling POTRF / SYTRF, and the non-diagonal blocks are updated to L2 through TRSM, and then the contribution matrix CB is generated through SYRK, and the CB column blocks are accurately transferred to the specified rows of the target supernode according to the adjacency list, completing the iterative process of decomposition. Among them, the SDB format only stores the dense areas of the diagonal blocks and non-diagonal blocks of the supernode, avoiding the waste of filling in the traditional format, and explicitly recording the data dependency path through the adjacency list without redundant indexes. Through the pre-allocated dense L matrix space, efficient BLAS / LAPACK functions such as POTRF, TRSM, and SYRK are directly called to convert the original discrete sparse operations into vectorized dense calculations, the decomposition speed is improved, and the cache locality is optimized through the continuous storage of supernodes, further reducing the memory access overhead. It supports differentiated storage of symmetric positive definite and non-positive definite matrices, and ensures the correctness of the data transmission path by eliminating the tree-driven decomposition order, avoiding dynamic memory adjustment, and achieving stable solution of large-scale matrices. This solution changes the dilemma of separation of storage and computation in traditional sparse matrix decomposition through the deep combination of super-node structured storage and dense BLAS, and achieves the unity of storage compactness and computational efficiency.
[0036] refer to Figure 2 This embodiment further discloses a method for efficient storage and decomposition of a sparse matrix for a supernode, which is applied to the above-mentioned sparse matrix efficient storage and decomposition device for a supernode, and the method comprises: S100: receiving an original matrix and performing matrix rearrangement on it to generate an elimination tree, including: rearranging the original matrix using an approximate minimum degree method or a nested decomposition method to bring the dense part of the original matrix closer to the diagonal; determining, for each node, its first child node to be eliminated in the decomposition process as a parent node; traversing the parent node relationship of all nodes to generate an elimination tree starting with a root node, wherein the child nodes of each node constitute its subtree; S200: constructing a supernode based on the elimination tree and obtaining an L matrix of each supernode, including: analyzing the row and column distribution of non-zero elements in its subtree based on the parent-child relationship of the elimination tree to determine the non-zero element structure of the node in the L matrix; if the L structures of adjacent nodes are exactly the same, merging these nodes into a supernode; defining a starting node number parameter and a length parameter of the supernode; generating an adjacency list and recording the length of the adjacency list according to the indexes of all non-zero rows below the supernode in the elimination tree; allocating storage space to each supernode according to the length parameter of the supernode and the length of the adjacency list to form an L matrix; S300: define the symmetric sparse matrix storage format of the supernode, including: define the data structure description of the supernode in C language, including: int start; / / the starting node number of the supernode; int len; / / the length of the supernode; int adjlen; / / the length of the adjacency list of the supernode; int adjlist; / / supernode adjacency list; float / double L; / / The first address of the L array of the super node; S400: Loading non-zero elements in the original matrix into the L matrix of the corresponding supernode to store the L matrix in a dense format, including: traversing each column of the original matrix to determine the supernode to which the non-zero element belongs; determining the element location of the non-zero element in the corresponding supernode to obtain the local coordinates of the non-zero element; writing the non-zero element value into the corresponding position of the L matrix of the supernode according to the local coordinates; S500: traverse the dense matrix L and perform matrix decomposition on it, including: According to the parent-child relationship of the elimination tree, starting from the root node, each supernode is processed from the upper left to the lower right; the decomposition function POTRF or SYTRF is called to decompose the diagonal blocks of the current supernode into the lower triangular matrix L1; based on the lower triangular matrix L1, the triangular matrix solving function DTRSM is called to update the non-diagonal blocks to obtain the L2 matrix; based on the L2 matrix, the matrix multiplication function SYRK is called to calculate the contribution matrix; the transfer path and transfer target of the contribution matrix are determined, and each column of the contribution matrix is copied to the specified row of the L matrix of the transfer target supernode according to the index of the adjacency list.
[0037] Figure 3 A schematic diagram of the physical structure of an electronic device provided in an embodiment of the present invention, such as Figure 3 As shown, the electronic device 50 includes: a processor 501 (processor), a memory 502 (memory) and a bus 503; The processor 501 and the memory 502 communicate with each other via the bus 503 ; the processor 501 is used to call program instructions in the memory 502 to execute the methods provided by the above-mentioned method implementation methods.
[0038] This embodiment provides a non-transitory computer-readable storage medium, which stores computer instructions. The computer instructions enable a computer to execute the methods provided by the above-mentioned method embodiments.
[0039] A person skilled in the art can understand that all or part of the steps for implementing the above-mentioned method implementation method can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium, which, when executed, executes the steps of the above-mentioned method implementation method; and the aforementioned storage medium includes: ROM, RAM, magnetic disk or optical disk, etc., various storage media that can store program codes.
[0040] The device implementation described above is merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, i.e., they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present implementation scheme. Those of ordinary skill in the art may understand and implement it without creative effort.
[0041] Through the description of the above implementation modes, those skilled in the art can clearly understand that each implementation mode can be implemented by means of software plus a necessary general hardware platform, or of course by hardware. Based on such an understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each implementation mode or some parts of the implementation mode.
[0042] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions may occur depending on design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A sparse matrix efficient storage and decomposition device for super nodes, characterized in that: The device comprises: A matrix preprocessing module, for receiving an original matrix and performing matrix rearrangement thereon to generate an elimination tree; A supernode construction module, used to construct supernodes based on the elimination tree and determine the L matrix of each supernode; Storage format definition module, used to define the symmetric sparse matrix storage format of super nodes; A data loading module, used for loading non-zero elements in the original matrix into the L matrix of the corresponding super node to store the L matrix in a dense format; The matrix decomposition module is used to traverse the dense L matrix and perform matrix decomposition on it.
2. The sparse matrix efficient storage and decomposition device for super nodes according to claim 1, characterized in that: The matrix preprocessing module specifically includes: A rearrangement unit, used for rearranging the original matrix by using an approximate minimum degree method or a nested partitioning method, so as to move a dense part of the original matrix close to a diagonal line; A node determination unit, used for determining, for each node, the child node that is first eliminated in the decomposition process as a parent node; The elimination tree generation unit is used to traverse the parent node relationship of all nodes and generate an elimination tree starting with a root node, wherein the child nodes of each node constitute its subtree.
3. The sparse matrix efficient storage and decomposition device for super nodes according to claim 1, characterized in that: The super node construction module specifically includes: A structural analysis unit, for analyzing the row and column distribution of non-zero elements in its subtree based on the parent-child relationship of the elimination tree, so as to determine the non-zero element structure of the node in the L matrix; A supernode merging unit is used to merge adjacent nodes into a supernode if the L structures of these nodes are exactly the same; A parameter definition unit, used to define a starting node number parameter and a length parameter of a supernode; An adjacency list generating unit, configured to generate an adjacency list according to the indexes of all non-zero rows below the supernode in the elimination tree and record the length of the adjacency list; The storage allocation unit is used to allocate storage space to each super node according to a length parameter of the super node and a length of the adjacency list to form an L matrix.
4. The sparse matrix efficient storage and decomposition device for super nodes according to claim 1, characterized in that: The data loading module specifically includes: A supernode attribution determination unit, used for traversing each column of the original matrix to determine the supernode to which the non-zero element belongs; A non-zero element positioning unit, used to determine the element positioning of the non-zero element in the corresponding supernode to obtain the local coordinates of the non-zero element; A data writing unit is used to write the non-zero element value into the corresponding position of the L matrix of the super node according to the local coordinates.
5. The sparse matrix efficient storage and decomposition device for super nodes according to claim 3, characterized in that: The matrix decomposition module specifically includes: A traversal control unit, used to process each super node from the top left to the bottom right starting from the root node according to the parent-child relationship of the elimination tree; A decomposition function calling unit, used for calling a decomposition function POTRF or SYTRF to decompose the diagonal block of the current supernode into a lower triangular matrix L1; A matrix updating unit, configured to call a triangular matrix solving function DTRSM to update non-diagonal blocks based on the lower triangular matrix L1 to obtain an L2 matrix; A contribution matrix calculation unit, used for calculating a contribution matrix by calling a matrix multiplication function SYRK based on the L2 matrix; The data transmission unit is used to determine the transmission path and transmission target of the contribution matrix, and copy each column of the contribution matrix to the specified row of the L matrix of the transmission target super node according to the index of the adjacency list.
6. A method for efficient storage and decomposition of sparse matrices for supernodes, applied to the device according to any one of claims 1 to 5, characterized in that: The method includes: The matrix preprocessing module receives the original matrix and performs matrix rearrangement on it to generate an elimination tree; The supernode construction module constructs supernodes based on the elimination tree and obtains the L matrix of each supernode; The storage format definition module defines the symmetric sparse matrix storage format of the super node; The data loading module loads the non-zero elements in the original matrix into the L matrix of the corresponding supernode to store the L matrix in a dense format; The matrix decomposition module traverses the dense format L matrix and performs matrix decomposition on it.
7. The method for efficient storage and decomposition of sparse matrices for supernodes according to claim 6, characterized in that: Also includes: The supernode attribution determination unit traverses each column of the original matrix to determine the supernode to which the non-zero element belongs; The non-zero element positioning unit determines the element positioning of the non-zero element in the corresponding supernode to obtain the local coordinates of the non-zero element; The data writing unit writes the non-zero element value into the corresponding position of the L matrix of the super node according to the local coordinates.
8. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the function of the device described in any one of claims 1 to 5 is realized.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the function of the device described in any one of claims 1 to 5 is realized.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the computer program implements the functions of the device according to any one of claims 1 to 5.
Citation Information
Patent Citations
Data storage and reading method and device, equipment and medium
CN119311221A
Computing device, method, equipment, chip and system
CN119337040A
Sparse matrix parallel solving method and device based on upper and lower triangular decomposition
CN114329327A
Matrix solving method, computer equipment, storage medium and program product
CN118152716A
Method and device for processing sparse matrix for heterogeneous distributed platform
CN118410265A