Minimum delay fine-grained parallel Givens rotation method for super-large scale matrix QR decomposition
By adopting external and internal loop collaborative parallel design and pipeline parallel architecture in super-large-scale matrix QR decomposition, the problems of low hardware resource utilization and large calculation delay in the existing technology are solved, and efficient matrix decomposition calculation is achieved.
Patent Information
- Application Number
- CN202510071035.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-16
AI Technical Summary
The prior art has problems such as low hardware resource utilization and large calculation delay in the ultra-large-scale matrix QR decomposition, especially in the insufficient parallel computing development of inner-layer loop blocks, resulting in a long wait time for the startup of the latter-level PE unit.
The collaborative parallel design of external and internal loops is adopted. The outer loop is responsible for the calculation of a row vector of the Q matrix and R matrix, and the inner loop is responsible for the Givens rotation of the two row vectors. The pipeline parallelization is realized through cascadingly connected PE units, and the element update module design is gradually reduced, making full use of the data correlation characteristics.
It significantly reduces hardware latency, optimizes resource utilization, and improves the computing efficiency of ultra-large-scale matrix QR decomposition.
Smart Images

Figure CN119988807A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of signal and information processing, and in particular to a super-large-scale matrix QR decomposition minimum delay fine-grained parallel Givens rotation method. Background Art
[0002] The rapid advancement of modern science and technology has led to an explosive growth in data size, and the performance standards for intensive computing have continued to rise. Matrix operations occupy a core position in intensive computing, and their computing performance largely determines the overall performance of the system. Large-scale matrix QR decomposition has a wide range of applications in signal processing, image processing, computational fluid dynamics, and computational structural mechanics.
[0003] In the implementation of QR decomposition, there are different situations for matrices of different sizes. For small-scale QR decomposition, a two-dimensional Systolic array is usually used for implementation. Although this two-dimensional array structure can provide high parallelism, it consumes a lot of hardware resources, and the number of processing units in the array generally reaches O(n2). When the matrix scale increases, it is difficult to implement the QR decomposition of the matrix with a two-dimensional array structure due to the limitations of chip area and power consumption. For larger matrices, the two-dimensional structure array is often folded into a one-dimensional structure with the help of folding and mapping methods. The number of linear array processing units implemented by this method is O(n). However, when the matrix scale increases further, the implementation difficulty of this method is still very high. At present, for the QR decomposition of ultra-large-scale matrices, a fine-grained parallel structure is often used. In this fine-grained parallel structure, the PE unit has the characteristics of simple structure and strong scalability, but it also has certain limitations: the existing design mainly focuses on optimizing the pipeline operation of the outer loop block, while the parallel computing of the inner loop block is underdeveloped. This design means that the startup of the next-level PE unit must rely on the results of two iterative calculations of the previous-level PE unit, that is, first wait for the previous-level PE unit to complete a set of Givens rotations, and then wait for the next set of Givens rotations to generate the required valid data. Due to this dependency, the startup waiting time of the next-level PE unit is long, and as the matrix size increases and the number of PE units increases, this waiting time is significantly extended, greatly limiting the computational efficiency of the QR decomposition hardware accelerator. Summary of the invention
[0004] The purpose of the present invention is to provide a minimum delay fine-grained parallel Givens Rotation method, which realizes efficient calculation of ultra-large-scale matrix QR decomposition through the collaborative parallel design of internal and external loop slices, significantly reduces hardware delay and optimizes resource utilization.
[0005] The technical solution of the present invention is: providing a super-large-scale matrix QR decomposition minimum delay fine-grained parallel Givens rotation method, which uses an operation architecture with an outer loop and an inner loop to perform matrix decomposition, and the method includes: the outer loop is responsible for the calculation of a row vector of the Q matrix and the R matrix, and the inner loop is responsible for the Givens rotation of the two row vectors;
[0006] When performing QR decomposition on a very large-scale matrix, the outer loop of the matrix is decomposed into several outer loop blocks, each of which contains multiple outer loop slices; each outer loop slice is calculated by a PE unit and is responsible for updating a row vector of the matrix;
[0007] Each PE unit is connected sequentially in a cascade manner, and the calculation tasks of the outer loop slices are processed in parallel inside the outer loop block; the first-level PE unit reads the initial matrix row vector from the external memory and updates it. After receiving part of the intermediate results generated by the previous-level PE unit, each subsequent level of PE unit immediately starts the calculation, updates the row vector it is responsible for, and transmits the intermediate results to the next level of PE unit in real time until all the calculation tasks of the outer loop block are completed;
[0008] Each PE unit further divides the outer loop slice into multiple inner loop slices. Each PE unit contains an update factor calculation module and several element update modules. Each inner loop slice is responsible for performing Givens rotation operations on two row vectors in the matrix. After the update factor calculation module calculates the Givens rotation factor, it directly passes the data to the element update module. The element update module processes part of the update task of the current row vector step by step, and passes the intermediate data to the next level PE unit in real time. The next level element update module receives the data generated by the previous level element update module step by step and starts the operation in real time to complete all the calculation tasks of the inner loop slice.
[0009] After all PE units complete the above calculations, the updated results generated by the last level PE unit are written into the off-chip memory for storage.
[0010] In any of the above technical solutions, further, the number of element update modules in each PE unit decreases step by step, specifically:
[0011] The first-level PE unit contains n element update modules; the second-level PE unit contains n-1 element update modules; the p-th level PE unit contains (n-p+1) element update modules; and the last-level PE unit contains only 1 element update module.
[0012] In any of the above technical solutions, further, the first-level PE unit reads the initial matrix data from the off-chip memory, and the remaining PE units immediately receive and start calculations after the previous-level PE unit generates intermediate calculation data, and the last-level PE unit writes the calculation results back to the off-chip memory to provide input data for subsequent outer loop blocks.
[0013] In any of the above technical solutions, further, the outer loop block is executed in a serial manner, and after all outer loop slices of an outer loop block are calculated, the calculation of the next outer loop block is started;
[0014] Inside the outer loop block, PE units execute the computational tasks of the outer loop slice in parallel in a cascaded manner.
[0015] In any of the above technical solutions, further, the n PE element update modules in the PE unit responsible for the i-th outer loop slice implement parallel calculation of an inner loop block in the following manner:
[0016] In the first inner loop block, the first-level element update module performs Givens rotation on the i-th vector and the (i+1)-th vector received by the current PE unit; the updated i-th vector is passed to the second-level element update module of the PE unit, and the (i+1)-th vector is passed to the PE unit responsible for updating the (i+1)-th vector;
[0017] In the s-th inner loop block, the first-level element update module performs Givens rotation on the (i+n×(s-1)+1)-th vector received by the current PE unit and the i-th vector stored in the on-chip memory in the (s-1)-th inner loop block;
[0018] The k-th level element update module (k≠1) performs Givens rotation on the (i+k)-th vector received by the current PE unit and the i-th vector updated by the (k-1)-th level element update module of the current PE unit;
[0019] The k-th PE unit passes the updated i-th vector to the (k+1)-th element update module of the current PE unit, and passes the (i+k)-th vector to the PE unit responsible for the (i+1)-th outer loop slice;
[0020] The n-th level element update module saves the i-th vector generated by the update in the on-chip memory for the calculation of the next inner loop block;
[0021] Each inner loop block is executed serially. After the calculation of the previous inner loop block is completed, the calculation of the next inner loop block is started until the calculation of all inner loop blocks is completed.
[0022] In any of the above technical solutions, further, the intermediate vector calculated by the non-last-stage PE unit is transmitted to the next-stage PE unit through the data path; the intermediate vector calculated by the last-stage PE unit is stored in the off-chip memory to provide input for the calculation of the subsequent outer loop block;
[0023] The vector updated by the givens rotation performed by the non-last-level element update module in the inner loop is passed to the next-level element update module, and the i-th vector updated by the last-level element update module is stored in the on-chip memory.
[0024] The beneficial effects of the present invention are:
[0025] The present invention optimizes the design of the key steps in the QR decomposition of ultra-large-scale matrices and proposes a minimum-delay fine-grained parallel Givens Rotation method. This method solves the problems of low hardware resource utilization and high computational latency in the QR decomposition of ultra-large-scale matrices through innovative architecture design and parallel computing design.
[0026] The present invention combines the pipeline design of the outer loop slice with the pipeline design of the inner loop slice for the first time, achieving efficient coordinated allocation of computing tasks in terms of time and hardware resources.
[0027] The present invention fully utilizes the characteristics of data correlation in Givens rotation calculation through the step-by-step decreasing element update module design and the PE unit cascade pipeline architecture, breaking through the limitations of traditional fine-grained parallel methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The advantages of the above and additional aspects of the present invention will become apparent and easily understood in the description of the embodiments in conjunction with the following drawings, in which:
[0029] Figure 1 is a spatiotemporal diagram of a minimal-delay fine-grained parallel Givens rotation method for ultra-large-scale matrix QR decomposition according to an embodiment of the present invention;
[0030] Figure 2 It is a schematic diagram of data dependency in an outer loop and an inner loop of a super-large-scale matrix QR decomposition minimum delay fine-grained parallel Givens rotation method according to an embodiment of the present invention;
[0031] Figure 3 is a schematic flow chart of a minimum-delay fine-grained parallel Givens rotation method for ultra-large-scale matrix QR decomposition according to an embodiment of the present invention;
[0032] Figure 4 1 is a space-time diagram of the prior art of the minimum delay fine-grained parallel Givens rotation method for ultra-large-scale matrix QR decomposition according to an embodiment of the present invention. DETAILED DESCRIPTION
[0033] In order to more clearly understand the above-mentioned purpose, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present invention and the features in the embodiments can be combined with each other without conflict.
[0034] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited to the specific embodiments disclosed below.
[0035] like Figures 1 to 3 As shown, this embodiment provides a minimum delay fine-grained parallel Givens rotation method for ultra-large-scale matrix QR decomposition, the method comprising: an outer loop and an inner loop, the outer loop is responsible for the calculation of a row vector of the Q matrix and the R matrix, and the inner loop is responsible for the Givens rotation of the two row vectors.
[0036] The system structure includes a multi-level processing unit (PE unit) array, each PE unit contains an update factor calculation module and multiple element update modules. Through the collaborative pipeline design of the outer loop and the inner loop, this method significantly improves the efficiency of matrix QR decomposition, reduces the idle and waiting time of hardware resources, and realizes efficient parallel operation.
[0037] In the outer loop, the outer loop block is first divided, and the outer loop of the matrix is decomposed into several outer loop blocks, each of which contains multiple outer loop slices; each outer loop slice is calculated by a processing unit (PE unit).
[0038] In the QR decomposition process, the outer loop is responsible for updating each row vector of the Q matrix and the R matrix. In order to achieve parallel processing, the outer loop is decomposed into several outer loop blocks, each of which contains multiple outer loop slices. The calculation of each outer loop slice is completed by a PE unit.
[0039] The execution order of the outer loop blocks is serial. After the first outer loop block completes the calculation, its result is used as the input of the second outer loop block, and the two blocks are executed serially in sequence until the calculation of all outer loop blocks is completed.
[0040] Multiple outer loop slices in each outer loop block are cascaded to achieve pipeline parallel operation. The first-level PE unit receives the initial matrix data from the external memory, and immediately passes the intermediate results to the second-level PE unit after completing the partial update. The second-level PE unit continues the calculation. And so on, each level of PE unit passes the intermediate data to the next level after completing its own task, until the last level of PE unit completes the update and writes the result back to the memory. Through the pipeline design, the waiting time between PE units is significantly reduced, and the overall computing efficiency of the system is improved.
[0041] The structure of the PE unit is designed. Each PE unit consists of an update factor calculation module (FAC_CAL) and an element update module (GR_UPD). The update factor calculation module (FAC_CAL) is responsible for calculating the update factor required for Givens rotation; the element update module (GR_UPD) is responsible for updating each element of the row vector.
[0042] The number of element update modules in each level of PE units decreases gradually, specifically: the first level PE unit contains n element update modules; the second level PE unit contains n-1 modules; the pth level PE unit contains (n-p+1) modules; and the last level PE unit contains only 1 module.
[0043] This step-by-step decreasing design can not only save hardware resources, but also match the task volume of PE units at different levels to avoid resource waste.
[0044] Each PE unit is connected to adjacent units in a cascaded manner. The first-level PE unit reads the matrix data to be processed from the external memory and passes the intermediate results to the next-level unit after the update is completed. Subsequent PE units start calculations immediately after receiving the intermediate data from the previous-level unit until all row vectors in the entire outer loop block are updated.
[0045] Then in the inner loop, each PE unit further divides the outer loop slice it is responsible for into multiple inner loop slices, each of which contains the Givens rotation calculation task for two row vectors.
[0046] Each PE unit contains multiple element update modules, which are cascaded to realize pipeline parallel calculation of the inner loop slice. The specific operations are as follows: after the first-level element update module receives data from the update factor calculation module, it immediately starts the Givens rotation operation and updates some elements of the two row vectors; the subsequent element update modules receive the output data of the previous module step by step and continue the update operation until all row vectors are updated; the last-level element update module stores the updated row vectors in the on-chip memory (BRAM) to provide input for subsequent calculations.
[0047] The element update module in the PE unit adopts a multi-pipeline parallel calculation method. The n PE element update modules in the PE unit responsible for the i-th outer loop slice realize the multi-pipeline parallel of an inner loop block in the following way:
[0048] In the first inner loop block, the first-level element update module performs Givens rotation on the i-th vector and the (i+1)-th vector received by the current PE unit; in the s-th inner loop block, the first-level element update module performs Givens rotation on the (i+n×(s-1)+1)-th vector received by the current PE unit and the i-th vector stored in the on-chip memory in the (s-1)-th inner loop block; the k-th element update module (k≠1) performs Givens rotation on the (i+k)-th vector received by the current PE unit and the (k-1)-th element update module of the current PE unit. The i-th vector updated by the module performs a Givens rotation; the k-th PE unit passes the updated i-th vector to the (k+1)-th element update module of the current PE unit, and passes the (i+k)-th vector to the PE unit responsible for the (i+1)-th outer loop slice; the n-th element update module saves the updated i-th vector in the on-chip memory for the calculation of the next inner loop block; each inner loop block is executed serially, and after the calculation of the previous inner loop block is completed, the calculation of the next inner loop block is started until the calculation of all inner loop blocks is completed.
[0049] The calculation results of the inner loop slices are stored in the on-chip memory, which reduces the access frequency to the off-chip memory and optimizes the data transfer efficiency. In addition, the serial execution of the inner loop slices ensures that the next loop slice can be started immediately after the calculation of the current loop slice is completed.
[0050] Before the outer loop and the inner loop work together, the system first performs data initialization operations: the ultra-large-scale matrix is divided into blocks for processing, and the data of each block is stored in the external memory. The first-level PE unit reads the initial data from the external memory and starts the calculation.
[0051] In the collaborative work of the outer loop and the inner loop, the pipeline parallel design of the outer loop slice ensures the continuous data flow between PE units, and the pipeline parallel processing of the inner loop slice improves the computing efficiency within a single PE unit. The final result is written back to the external memory by the last-level PE unit for subsequent use.
[0052] like Figure 1As shown, it is a time-space diagram of the minimum delay fine-grained parallel Givens Rotation method in the present invention, which takes the existing fine-grained parallel model including three PE units as an example. Since the method in the present invention adopts a parallel computing mode when processing inner loop slices, the start condition of the PE unit responsible for the calculation of the (i+1)th outer loop slice is:
[0053] The PE unit responsible for the i-th outer loop slice calculates A(i+1,i+1) when performing Givens rotation on the i-th and (i+1)-th row vectors.
[0054] The PE unit responsible for the i-th outer loop slice calculates A(i+2, i+1) when performing Givens rotation on the (i+2)-th row vector and the i-th vector updated by the current-level PE unit.
[0055] Figure 4 The update factors and element update time of the existing conventional cycle. The next PE unit needs to wait until the previous operation is completed before it can start the operation. Figure 1 and Figure 4 It can be seen that after the pipeline parallel operation of the inner loop slice, the key data required for the PE unit to start can be generated in advance, which greatly shortens the waiting time for the next PE unit to start and reduces the idle rate of the PE unit. As the matrix size increases, the number of cycles of the PE array will also increase, and the role of this model in reducing delays and improving QR decomposition efficiency will be more significant.
[0056] In summary, the present invention proposes a minimum delay fine-grained parallel Givens rotation method for ultra-large-scale matrix QR decomposition, including: the method includes: an outer loop and an inner loop, the outer loop is responsible for the calculation of a row vector of the Q matrix and the R matrix, and the inner loop is responsible for the Givens rotation of two row vectors.
[0057] The outer loop of the matrix is decomposed into several outer loop blocks, each of which contains multiple outer loop slices; each outer loop slice is calculated by a PE unit.
[0058] Each PE unit is connected in a cascade manner, and processes the calculation tasks of the outer loop slice in parallel through a pipeline inside the outer loop block; each PE unit contains an update factor calculation module and several element update modules, which are responsible for the calculation tasks of the inner loop slice; the element update module completes the Givens rotation calculation of the inner loop slice through a pipeline parallel manner.
[0059] The steps in the present invention can be adjusted in order, combined or deleted according to actual needs.
[0060] The units in the device of the present invention can be combined, divided and deleted according to actual needs.
[0061] Although the present invention is disclosed in detail with reference to the accompanying drawings, it should be understood that these descriptions are merely exemplary and are not intended to limit the application of the present invention. The scope of protection of the present invention is defined by the appended claims and may include various modifications, alterations and equivalents made to the invention without departing from the scope and spirit of the present invention.
Claims
1. Ultra-large-scale matrix QR decomposition minimum delay fine-grained parallel Givens rotation method, characterized by: The method uses a computing architecture with an outer loop and an inner loop to perform matrix decomposition, and the method includes: the outer loop is responsible for calculating a row vector of a Q matrix and an R matrix, and the inner loop is responsible for a Givens rotation of two row vectors; When performing QR decomposition on a very large-scale matrix, the outer loop of the matrix is decomposed into several outer loop blocks, each of which contains multiple outer loop slices; each outer loop slice is calculated by a PE unit and is responsible for updating a row vector of the matrix; Each PE unit is connected sequentially in a cascade manner, and the calculation tasks of the outer loop slices are processed in parallel inside the outer loop block; the first-level PE unit reads the initial matrix row vector from the external memory and updates it. After receiving part of the intermediate results generated by the previous-level PE unit, each subsequent level of PE unit immediately starts the calculation, updates the row vector it is responsible for, and transmits the intermediate results to the next level of PE unit in real time until all the calculation tasks of the outer loop block are completed; Each PE unit further divides the outer loop slice into multiple inner loop slices. Each PE unit contains an update factor calculation module and several element update modules. Each inner loop slice is responsible for performing Givens rotation operations on two row vectors in the matrix. After the update factor calculation module calculates the Givens rotation factor, it directly passes the data to the element update module. The element update module processes part of the update task of the current row vector step by step, and passes the intermediate data to the next level PE unit in real time. The next level element update module receives the data generated by the previous level element update module step by step and starts the operation in real time to complete all the calculation tasks of the inner loop slice. After all PE units complete the above calculations, the updated results generated by the last level PE unit are written into the off-chip memory for storage.
2. The ultra-large-scale matrix QR decomposition minimum delay fine-grained parallel Givens rotation method according to claim 1, characterized in that: The number of element update modules in each PE unit decreases gradually, specifically: The first-level PE unit contains n element update modules; the second-level PE unit contains n-1 element update modules; the p-th level PE unit contains (n-p+1) element update modules; and the last-level PE unit contains only 1 element update module.
3. The ultra-large-scale matrix QR decomposition minimum delay fine-grained parallel Givens rotation method according to claim 2, characterized in that: The first-level PE unit reads the initial matrix data from the off-chip memory. The remaining PE units receive and start operations immediately after the intermediate calculation data is generated by the previous-level PE unit. The last-level PE unit writes the calculation results back to the off-chip memory to provide input data for subsequent outer loop blocks.
4. The ultra-large-scale matrix QR decomposition minimum delay fine-grained parallel Givens rotation method according to claim 1, characterized in that: The outer loop blocks are executed in serial mode. After all the outer loop slices of an outer loop block are calculated, the calculation of the next outer loop block is started; Inside the outer loop block, PE units execute the computational tasks of the outer loop slice in parallel in a cascaded manner.
5. The ultra-large-scale matrix QR decomposition minimum delay fine-grained parallel Givens rotation method according to claim 1, characterized in that: The n PE element update modules in the PE unit responsible for the i-th outer loop slice implement parallel computation of an inner loop block in the following way: In the first inner loop block, the first-level element update module performs Givens rotation on the i-th vector and the (i+1)-th vector received by the current PE unit; the updated i-th vector is passed to the second-level element update module of the PE unit, and the (i+1)-th vector is passed to the PE unit responsible for updating the (i+1)-th vector; In the s-th inner loop block, the first-level element update module performs Givens rotation on the (i+n×(s-1)+1)-th vector received by the current PE unit and the i-th vector stored in the on-chip memory in the (s-1)-th inner loop block; The k-th level element update module (k≠1) performs Givens rotation on the (i+k)-th vector received by the current PE unit and the i-th vector updated by the (k-1)-th level element update module of the current PE unit; The k-th PE unit passes the updated i-th vector to the (k+1)-th element update module of the current PE unit, and passes the (i+k)-th vector to the PE unit responsible for the (i+1)-th outer loop slice; The n-th level element update module saves the i-th vector generated by the update in the on-chip memory for calculation of the next inner loop block; Each inner loop block is executed serially. After the calculation of the previous inner loop block is completed, the calculation of the next inner loop block is started until the calculation of all inner loop blocks is completed.
6. The ultra-large-scale matrix QR decomposition minimum delay fine-grained parallel Givens rotation method according to claim 5, characterized in that: The intermediate vector calculated by the non-last-level PE unit is passed to the next-level PE unit through the data path; the intermediate vector calculated by the last-level PE unit is stored in the off-chip memory to provide input for the calculation of the subsequent outer loop block; The vector updated by the givens rotation performed by the non-last-level element update module in the inner loop is passed to the next-level element update module, and the i-th vector updated by the last-level element update module is stored in the on-chip memory.
Citation Information
Patent Citations
Parallelizing matrix factorization across hardware accelerators
CN108139887A
Allocating processing threads for matrix-matrix multiplication
CN114428936A
Numerical resolution method in matrix
JP2007264788A
Module and method for solving matrix triangular decomposition on basis of improved bitwise substitution
WO2017107337A1
Fast sparse neural networks
WO2021058578A1