Method and system for realizing Cholesky decomposition
By defining and optimizing the Cholesky decomposition module using the Chisel language, the problem of inefficiency of the Cholesky decomposition algorithm on the hardware is solved, and an efficient, stable and scalable Cholesky decomposition hardware circuit is realized.
Patent Information
- Application Number
- CN202510169902.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-30
AI Technical Summary
When the prior art realizes Cholesky decomposition on hardware, the calculation complexity is high, the memory consumption is large, the data dependence is strong, and the hardware design language is difficult to write and debug, resulting in inefficiency.
The Cholesky decomposition module, matrix multiplication module, matrix subtraction module and main controller module are defined using the Chisel language. The calculation is optimized through the LDL decomposition method and the Strassen algorithm to generate efficient hardware circuits.
It improves the execution speed and efficiency of Cholesky's decomposition algorithm on hardware, has good scalability and portability, and is suitable for large-scale matrix operations and linear equation system solutions.
Smart Images

Figure CN120067511A_ABST
Abstract
Description
Technical Field
[0001] The present invention mainly relates to the technical fields of mathematics and computer technology, and particularly relates to a method and system for implementing Cholesky decomposition. Background Art
[0002] Cholesky decomposition is a method of decomposing a Hermitian positive definite matrix (a Hermitian positive definite matrix) into the product of a lower triangular matrix and its conjugate transpose. It is useful in efficient numerical solutions. For example, in Monte Carlo simulations, it is necessary to perform Cholesky decomposition on the covariance matrix to generate normally distributed random numbers.
[0003] Currently, there are the following problems with the methods for implementing Cholesky decomposition on hardware: (1) The method itself has a high computational complexity and memory consumption. Therefore, when implementing on hardware, it is necessary to consider how to optimize the computational speed and memory usage.
[0004] (2) The method itself has strong data dependencies. Therefore, when implementing on hardware, it is necessary to consider how to utilize parallelism and pipelining techniques to improve efficiency.
[0005] (3) Hardware design languages are usually difficult to write and debug. Therefore, when implementing on hardware, it is necessary to consider how to simplify the design process and verification methods. In the prior art, there have been some methods and systems for implementing the Cholesky decomposition algorithm using different hardware design languages. However, these methods and systems usually have the following deficiencies: (1) Writing the Cholesky decomposition algorithm using traditional hardware design languages (such as VHDL or Verilog) is cumbersome and inefficient. These languages lack high-level abstractions and modular features. (2) Converting the Cholesky decomposition algorithm written in C language to a hardware circuit using high-level synthesis tools (such as Catapult C or Vivado HLS) may result in waste of hardware resources and degradation of performance. These tools lack control over hardware characteristics and optimizations. (3) Writing the Cholesky decomposition algorithm using other hardware construction languages (such as Bluespec or Lava) may encounter language limitations and compatibility issues. These languages lack support for existing hardware design tools and libraries.
[0006] Chisel language is a hardware construction language that can be used to describe and generate digital circuits. It has the characteristics of high-level abstraction and low-level control. Summary of the Invention
[0007] The object of the present invention is to provide a method and system for implementing Cholesky decomposition, which utilizes the characteristics of high-level abstraction and low-level control of Chisel language, as well as its compatibility with existing hardware design tools and libraries, to implement an efficient, stable, scalable, and portable method and system for Cholesky decomposition.
[0008] To achieve the above object, the present invention provides a method for implementing Cholesky decomposition, including the following steps: Step 1: Input a positive definite matrix A, and divide the positive definite matrix A into 4 sub-block matrices, including A00, A01, A10, and A11. Each sub-block can be recursively divided; Step 2: Define a Cholesky decomposition module using Chisel language. The Cholesky decomposition module receives a sub-block matrix as input and outputs a lower triangular matrix L and a diagonal matrix D; Step 3: Define a matrix multiplication module using Chisel language. The matrix multiplication module receives two sub-block matrices as input and outputs the product of the two sub-block matrices; Step 4: Define a matrix subtraction module using Chisel language. The matrix subtraction module receives two sub-block matrices as input and outputs the difference between the two sub-block matrices; Step 5: Define a main controller module using Chisel language to execute the Cholesky decomposition algorithm and output the final lower triangular matrix L.
[0009] Further, it also includes Step 6: Use the generator function of Chisel language to convert the method for Cholesky decomposition written in Chisel language into a description file of a hardware circuit, including Verilog language and VHDL language.
[0010] Further, the positive definite matrix A in Step 1 is a Hermitian positive definite matrix.
[0011] Further, in Step 2, the LDL decomposition method is used inside the Cholesky decomposition module to calculate the lower triangular matrix L and the diagonal matrix D. The steps of the LDL decomposition method include: S1: Assume the size of the positive definite matrix A is n×n, and the size of the sub-block matrix is m×m. Then n = K×m, where K is a positive integer; S2: Initialize the lower triangular matrix L and the diagonal matrix D as zero matrices; S3: Set i from 0 to m - 1, and the operations to be performed are: Calculate D[i][i] = A[i][i] - L[i][0:i] * D[0:i][0:i] * L[i][0:i]T̂; Set \(j\) from \(i + 1\) to \(m - 1\), and the operations to be performed are as follows: Calculate \(L[j][i]=(A[j][i]- L[j][0:i]*D[0:i][0:i]*L[i][0:i]^T) / D[i][i]\); S4: Output the lower triangular matrix \(L\) and the diagonal matrix \(D\).
[0012] Furthermore, the matrix multiplication module in step 3 uses the Strassen algorithm internally to accelerate matrix multiplication. The steps of the Strassen algorithm are as follows: S5: Divide the two input sub-block matrices into 4 smaller sub-block matrices, then ; S6: Calculate 7 intermediate results, \(P1=(X00 + X11)*(Y00 + Y11)\), \(P2=(X10 + X11)*Y00\), \(P3=X00*(Y01 - Y11)\), \(P4=X11*(Y10 - Y11)\), \(P5=(X00 + X01)*Y11\), \(P6=(X10 - X00)*(Y00 + Y01)\), \(P7=(X01 - X11)*(Y10 + Y11)\); S7: According to the 7 intermediate results, calculate the four parts of the output sub-block matrix, which are , where \(Z00 = P1 + P4 - P5 + P7\), \(Z01 = P3 + P5\), \(Z10 = P2 + P4\), \(Z11 = P1 - P2 + P3 + P6\).
[0013] Furthermore, the steps for the main controller module to execute the Cholesky decomposition algorithm in step 5 are as follows: S8: Call the Cholesky decomposition module to decompose the sub-block matrix \(A00\) to obtain the lower triangular matrix \(L00\) and the diagonal matrix \(D00\); S9: Call the matrix multiplication module to perform a multiplication operation on the diagonal matrix \(D00\) and the sub-block matrix \(A10\), and the resulting product matrix is \(X10\); S10: Call the matrix multiplication module to perform a multiplication operation on the lower triangular matrix \(L00\) and the product \(X10\), and the resulting product matrix is \(Y10\); S11: Call the matrix subtraction module to perform a subtraction operation on the sub-block matrix \(A11\) and the product \(Y10\), and the resulting difference matrix is \(Z11\); S12: Call the Cholesky decomposition module to decompose \(Z11\) to obtain the lower triangular matrix \(L11\) and the diagonal matrix \(D11\); S13: Call the matrix multiplication module to perform a multiplication operation on the lower triangular matrix \(L00\) and the sub-block matrix \(A01\), and the resulting product matrix is \(X01\); S14: Invoke the matrix multiplication module to perform a multiplication operation on the diagonal matrix D00 and the product X01, and the resulting product matrix is Y01; S15: Output L00, Y01, X10, L11 as the final lower triangular matrix L.
[0014] Furthermore, different parameters can be used to adjust the performance and resource consumption of the Cholesky decomposition method. The parameters include the size of the sub-block matrix, the number of pipeline stages, and the data bit width.
[0015] To achieve the above object, the present invention also provides a system for implementing Cholesky decomposition, including: An input interface, configured to receive a positive definite matrix A and divide the positive definite matrix A into 4 sub-block matrices, including A00, A01, A10, A11, and each sub-block can be recursively divided; An output interface, configured to output a lower triangular matrix L and combine L00, Y01, X10, L11 into a complete matrix; A memory, configured to store the input and output sub-block matrices; A Cholesky decomposition module, configured to decompose the input sub-block matrix and output the corresponding lower triangular matrix and diagonal matrix; A matrix multiplication module, configured to perform a multiplication operation on the input or output sub-block matrix and output the corresponding product matrix; A matrix subtraction module, configured to perform a subtraction operation on the input or output sub-block matrix and output the corresponding difference matrix; A main controller module, configured to coordinate the data transmission and control signals between the input interface, output interface, memory, and each module.
[0016] Beneficial effects: The present invention provides a method and system for implementing Cholesky decomposition. By utilizing the high-level abstraction and low-level control characteristics of the Chisel language, as well as its compatibility with existing hardware design tools and libraries, it can improve the execution speed and efficiency of the Cholesky decomposition algorithm on hardware, and has good scalability and portability to implement an efficient, stable, scalable, and portable Cholesky decomposition method and system. The present invention can be applied in fields that require large-scale matrix operations and solving linear equations, such as Monte Carlo simulation, etc. Description of the Drawings
[0017] Figure 1 It is a schematic diagram of the Cholesky decomposition module involved in the embodiment of the present invention; Figure 2It is a structural diagram of a linear algebra accelerator system that implements Cholesky decomposition using the Chisel language in an embodiment of the present invention. Detailed implementation manners
[0018] The preferred mechanisms and implementation methods of the present invention will be further described below in conjunction with the accompanying drawings and specific implementation manners.
[0019] As Figures 1 to 2 shown, an embodiment of the present invention discloses a technical solution for a method and system for implementing Cholesky decomposition. Embodiment 1
[0020] A method for implementing Cholesky decomposition includes the following steps: Step 1: Input a Hermitian positive definite matrix A, and divide the positive definite matrix A into 4 sub-block matrices A00, A01, A10, and A11. Each sub-block can be recursively divided. Assume the size of A is n×n, and the size of the sub-block matrix is m×m, then n = k×m, where k is a positive integer. For example, when n = 16 and m = 4, then k = 4, and A can be divided into the following form: , where each sub-block matrix is a 4×4 matrix.
[0021] Step 2: Define a Cholesky decomposition module using the Chisel language. The Cholesky decomposition module receives a sub-block matrix as input and outputs a lower triangular matrix L and a diagonal matrix D.
[0022] The Cholesky decomposition module internally uses the LDL decomposition method to calculate the lower triangular matrix L and the diagonal matrix D. The steps of the LDL decomposition method include: (1) Initialize the lower triangular matrix L and the diagonal matrix D as zero matrices; (2) Set i from 0 to m - 1, and the operations performed are: Calculate D[i][i] = A[i][i] - L[i][0:i] * D[0:i][0:i] * L[i][0:i]T̂; Set j from i + 1 to m - 1, and the operations performed are: Calculate L[j][i] = (A[j][i] - L[j][0:i] * D[0:i][0:i] * L[i][0:i]T̂) / D[i][i]; (3) Output the lower triangular matrix L and the diagonal matrix D.
[0023] Step 3: Define a matrix multiplication module using Chisel language. The matrix multiplication module receives two sub-block matrices as inputs and outputs the product of the two sub-block matrices.
[0024] The Strassen algorithm is used inside the matrix multiplication module to accelerate matrix multiplication. The steps of the Strassen algorithm include: (1) Divide the two input sub-block matrices into 4 smaller sub-block matrices. For example: ; (2) Calculate 7 intermediate results: P1 = (X00 + X11) * (Y00 + Y11), P2 = (X10 + X11) * Y00, P3 = X00 * (Y01 - Y11), P4 = X11 * (Y10 - Y11), P5 = (X00 + X01) * Y11, P6 = (X10 - X00) * (Y00 + Y01), P7 = (X01 - X11) * (Y10 + Y11); (3) Calculate the four parts of the output sub-block matrix based on the 7 intermediate results, which are , where Z00 = P1 + P4 - P5 + P7, Z01 = P3 + P5, Z10 = P2 + P4, Z11 = P1 - P2 + P3 + P6.
[0025] Step 4: Define a matrix subtraction module using Chisel language. The matrix subtraction module receives two sub-block matrices as inputs and outputs the difference between the two sub-block matrices. Simple element-wise subtraction is used inside this module to calculate matrix subtraction.
[0026] Step 5: Define a main controller module using Chisel language to execute the Cholesky decomposition algorithm and output the final lower triangular matrix L.
[0027] The steps for the main controller module to execute the Cholesky decomposition algorithm are: (1) Call the Cholesky decomposition module to decompose the sub-block matrix A00 to obtain the lower triangular matrix L00 and the diagonal matrix D00; (2) Call the matrix multiplication module to perform a multiplication operation on the diagonal matrix D00 and the sub-block matrix A10, and the resulting product matrix is X10; (3) Call the matrix multiplication module to perform a multiplication operation on the lower triangular matrix L00 and the product X10, and the resulting product matrix is Y10; (4) Call the matrix subtraction module to perform a subtraction operation on the sub-block matrix A11 and the product Y10, and the resulting difference matrix is Z11; (5) Call the Cholesky decomposition module to decompose Z11 to obtain the lower triangular matrix L11 and the diagonal matrix D11; (6) Call the matrix multiplication module to perform a multiplication operation on the lower triangular matrix L00 and the sub-block matrix A01, and the resulting product matrix is X01; (7) Call the matrix multiplication module to perform a multiplication operation on the diagonal matrix D00 and the product X01, and the resulting product matrix is Y01; (8) Output L00, Y01, X10, and L11 as the final lower triangular matrix L.
[0028] Use the generator function of the Chisel language to convert the Cholesky decomposition algorithm written in the Chisel language into a description file of the hardware circuit, such as Verilog or VHDL.
[0029] This embodiment uses the following parameters to adjust the performance and resource consumption of the algorithm: the size of the sub-block matrix is m = 4; the number of pipeline stages is p = 2; the data bit width is w = 32; The size of the sub-block matrix refers to the dimension of the sub-matrix divided from a large matrix. By setting the matrix block dimension, the balance between the calculation parallelism and the memory resource occupation can be balanced; The number of pipeline stages refers to the number of steps into which the instruction execution process is divided in the computer architecture. Each step is processed by an independent hardware unit, and the number of these steps is the number of pipeline stages. By dividing the calculation unit into a multi-stage pipeline structure, the balance between the task throughput rate and the logical resource consumption can be controlled; The data bit width refers to the number of binary digits of the data that can be processed or transmitted at one time in a computer system or digital circuit. By configuring the bit precision of the numerical calculation, the balance between the algorithm precision and the consumption of storage and transmission bandwidth can be adjusted.
[0030] This embodiment uses an FPGA as the target platform and uses the following configuration options to generate a Verilog description file: the clock frequency f = 100 MHz; the optimization strategy s = area.
[0031] The structure of the hardware circuit in this embodiment is as Figure 2 shown. The hardware circuit of this embodiment can perform Cholesky decomposition on a Hermitian positive definite matrix and output a lower triangular matrix. The performance and resource consumption of the hardware circuit of this embodiment are shown in Table 1. It can be seen from Table 1 that the hardware circuit of this embodiment has a high calculation speed and low resource consumption, indicating that the method and system of the present invention have good efficiency and scalability.
[0032] Embodiment 2
[0033] A system for implementing Cholesky decomposition includes: An input interface for receiving a Hermitian positive definite matrix A and partitioning the Hermitian positive definite matrix A into four sub-block matrices A00, A01, A10, and A11, where each sub-block can be recursively partitioned; An output interface for outputting a lower triangular matrix L and combining L00, Y01, X10, and L11 into a complete matrix; A memory for storing the input and output sub-block matrices; A Cholesky decomposition module (CholeskyModule) for decomposing the input sub-block matrices and outputting the corresponding lower triangular matrix and diagonal matrix; A matrix multiplication module (MatMulModule) for performing multiplication operations on the input or output sub-block matrices and outputting the corresponding product matrix; A matrix subtraction module (MatSubModule) for performing subtraction operations on the input or output sub-block matrices and outputting the corresponding difference matrix; A main controller module (MainController) for coordinating data transmission and control signals among the input interface, output interface, memory, and each module.
[0034] This embodiment uses the AXI protocol as the input and output interfaces and SRAM as the memory. This embodiment uses an FPGA as the target platform and uses the following configuration options to generate a Verilog description file: clock frequency f = 200 MHz; optimization strategy s = speed.
[0035]
[0036] The present invention provides a method and system for implementing Cholesky decomposition. Cholesky decomposition is a method of decomposing a Hermitian positive definite matrix (a Hermitian positive definite matrix) into the product of a lower triangular matrix and its conjugate transpose, which is useful in efficient numerical solutions. The Chisel language is a hardware construction language that can be used to describe and generate digital circuits. The present invention provides a method of writing a Cholesky decomposition algorithm using the Chisel language and converting it into a hardware circuit. The present invention also provides a linear algebra accelerator system based on the Cholesky decomposition algorithm. The present invention can improve the execution speed and efficiency of the Cholesky decomposition algorithm on hardware and has good scalability and portability. The present invention can be applied to fields that require large-scale matrix operations and solving linear equations, such as Monte Carlo simulations, etc.
[0037] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. However, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for implementing Cholesky decomposition, characterized in that: The following steps are involved: Step 1: Input a positive definite matrix A and divide it into 4 sub-block matrices, including A00, A01, A10, and A11. Each sub-block can be recursively divided. Step 2: Use Chisel language to define a Cholesky decomposition module, which receives a sub-block matrix as input and outputs a lower triangular matrix L and a diagonal matrix D; Step 3: Use Chisel language to define a matrix multiplication module, which receives two sub-block matrices as input and outputs the product of the two sub-block matrices; Step 4: Use Chisel language to define a matrix subtraction module, which receives two sub-block matrices as input and outputs the difference between the two sub-block matrices; Step 5: Define a main controller module using Chisel language to execute the Cholesky decomposition algorithm and output the final lower triangular matrix L.
2. A method for implementing Cholesky decomposition according to claim 1, characterized in that: The method also includes step 6: using the generator function of the Chisel language to convert the Cholesky decomposition method written in the Chisel language into a description file of the hardware circuit, including the Verilog language and the VHDL language.
3. A method for implementing Cholesky decomposition according to claim 1, characterized in that: The positive definite matrix A in step 1 is a Hermitian positive definite matrix.
4. A method for implementing Cholesky decomposition according to claim 1, characterized in that: In step 2, the Cholesky decomposition module uses the LDL decomposition method to calculate the lower triangular matrix L and the diagonal matrix D. The steps of the LDL decomposition method include: S1: Assume that the size of the positive definite matrix A is n×n, and the size of the sub-block matrix is m×m, then n=K×m, where K is a positive integer; S2: Initialize the lower triangular matrix L and the diagonal matrix D to zero matrices; S3: Set i from 0 to m-1, and perform the following operations: Calculate D[i][i]=A[i][i]-L[i][0:i]*D[0:i][0:i]*L[i][0:i]T̂; Set j from i+1 to m-1, and perform the following operations: Calculate L[j][i]=(A[j][i]- L[j][0:i]*D[0:i][0:i]*L[i][0:i]T̂) / D[i][i]; S4: Output the lower triangular matrix L and the diagonal matrix D.
5. A method for implementing Cholesky decomposition according to claim 1, characterized in that: The matrix multiplication module in step 3 uses the Strassen algorithm to accelerate matrix multiplication. The steps of the Strassen algorithm include: S5: Split the two input sub-block matrices into four smaller sub-block matrices, then ; S6: Calculate 7 intermediate results, P1=(X00+X11)∗(Y00+Y11), P2=(X10+X11)∗Y00, P3=X00∗(Y01−Y11), P4=X11∗(Y10−11), P5=(X00+X01)*Y11, P6=(X10-X00)*(Y00+Y01), P7=(X01-X11)*(Y10+Y11); S7: Based on the 7 intermediate results, the four parts of the output sub-block matrix are calculated as , where Z00=P1+P4−P5+P7, Z01=P3+P5, Z10=P2+P4, and Z11=P1−P2+P3+P6.
6. A method for implementing Cholesky decomposition according to claim 1, characterized in that: In step 5, the steps of executing the Cholesky decomposition algorithm by the main controller module are: S8: Call the Cholesky decomposition module to decompose the sub-block matrix A00 to obtain the lower triangular matrix L00 and the diagonal matrix D00; S9: calling the matrix multiplication module to perform multiplication operation on the diagonal matrix D00 and the sub-block matrix A10, and the obtained product matrix is X10; S10: calling the matrix multiplication module to perform multiplication operation on the lower triangular matrix L00 and the product X10, and the obtained product matrix is Y10; S11: calling the matrix subtraction module to perform a subtraction operation on the sub-block matrix A11 and the product Y10, and the obtained difference matrix is Z11; S12: Call the Cholesky decomposition module to decompose Z11 to obtain the lower triangular matrix L11 and the diagonal matrix D11; S13: calling the matrix multiplication module to perform multiplication operation on the lower triangular matrix L00 and the sub-block matrix A01, and the obtained product matrix is X01; S14: calling the matrix multiplication module to perform multiplication operation on the diagonal matrix D00 and the product X01, and the obtained product matrix is Y01; S15: Output L00, Y01, X10, L11 as the final lower triangular matrix L.
7. A method for implementing Cholesky decomposition according to claim 1, characterized in that: The performance and resource consumption of the Cholesky decomposition method can be adjusted using different parameters, including the size of the sub-block matrix, the number of pipeline stages, and the data bit width.
8. A system for implementing Cholesky decomposition, characterized in that: include: An input interface is used to receive a positive definite matrix A and divide the positive definite matrix A into 4 sub-block matrices, including A00, A01, A10, and A11, and each sub-block can be recursively divided; Output interface, used to output a lower triangular matrix L and combine L00, Y01, X10, and L11 into a complete matrix; A memory for storing input and output sub-block matrices; Cholesky decomposition module, used to decompose the input sub-block matrix and output the corresponding lower triangular matrix and diagonal matrix; A matrix multiplication module, used for performing multiplication operation on the input or output sub-block matrices and outputting the corresponding product matrix; A matrix subtraction module, used for performing subtraction operations on input or output sub-block matrices and outputting corresponding difference matrices; The main controller module is used to coordinate the input interface, output interface, memory, data transmission and control signals between various modules.