Cholesky decomposition heterogeneous parallel optimization method and system based on SW architecture

By constructing a heterogeneous parallel optimization method for Cholesky decomposition on the Shenwei architecture, the problem of high computing resources and time consumption in the beam adjustment optimization process in the existing technology is solved, and efficient beam adjustment optimization and improvement of computing efficiency is achieved.

CN120179412AActive Publication Date: 2025-06-20QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Patent Information

Application Number
CN202510631260.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-06-20
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

The existing Cholesky decomposition algorithm requires a lot of computing resources and time in the beam adjustment optimization process, and fails to make full use of the parallel computing power of modern high-performance computing platforms, especially on Shenwei supercomputers.

Method used

A Cholesky decomposition heterogeneous parallel optimization method based on Shenwei architecture is proposed. By constructing a symmetric positive definite matrix of beam adjustment, and sub-block division of the matrix is ​​used to use a distributed parallel allocation scheme, task-level parallel acceleration is performed in combination with the MPI programming model, and secondary parallel acceleration optimization is performed on the core operations through the master-slave core acceleration parallel features.

Benefits of technology

Efficient optimization of beam adjustment is achieved, memory bandwidth bottlenecks and data access conflicts are minimized, computing efficiency is improved, and high overlap between computing tasks and communication tasks is achieved through technical means such as asynchronous communication and double buffering mechanisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179412A_ABST
    Figure CN120179412A_ABST
Patent Text Reader

Abstract

The invention provides a Cholesky decomposition heterogeneous parallel optimization method and system based on an SW architecture, and relates to the technical field of high-performance computing. The method comprises the following steps: performing sub-block division on a symmetric positive definite matrix based on a distributed parallel distribution scheme, and performing iteration to complete matrix decomposition; each sub-block is distributed to different processes through an MPI programming model, data exchange is carried out between the processes through asynchronous communication, and coarse-grained task-level parallel acceleration is carried out; performing two-stage parallel acceleration on four operations in Cholesky decomposition by utilizing the acceleration parallel characteristic of a master core and a slave core of the SW architecture; wherein for GEMM and SYRK operations, column vectors of a matrix are mapped to a slave core array, and columns are divided according to the number of slave cores; the calculation process is optimized through a double-buffering mechanism, vectorization operation and a loop expansion technology, and the parallel efficiency is improved; and for the TRSM operation, the TRSM operation is decomposed into a plurality of TRSV operations, the TRSV operations are allocated to the slave cores for parallel execution, and a circular reading and data broadcasting mode is adopted to reduce data dependence and realize efficient parallel calculation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of high-performance computing, and particularly relates to a Cholesky decomposition heterogeneous parallel optimization method and system based on the ShenWei architecture. Background Art

[0002] The statements in this part only provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] Cholesky decomposition, also known as the square root method, is most commonly used in engineering applications to solve linear systems of symmetric positive definite matrices and has wide applications in fields such as finite element analysis, power systems, signal processing, computer vision, financial engineering, medical imaging, and robotics. Its high efficiency and numerical stability make it an ideal choice for handling large-scale computational problems.

[0004] A symmetric positive definite matrix means that for a real symmetric matrix A is positive definite if and only if for all non-zero real coefficient vectors x, there is . Cholesky decomposition is an efficient matrix decomposition method in linear algebra, applicable to symmetric positive definite matrices, and is used to decompose a symmetric positive definite matrix A into the form of a lower triangular matrix and the transpose of its lower triangular matrix, that is, .

[0005] In the fields of computer vision and 3D reconstruction, bundle adjustment is a core optimization problem, aiming to optimize camera parameters and 3D point coordinates by minimizing the reprojection error. Based on the corresponding image points formed by multiple 3D points on multiple cameras, the orientation of the cameras and the positions of the 3D points can be estimated. Through bundle adjustment, the orientations of all cameras can be optimized, and the errors in the positions of the 3D points can be minimized as much as possible. Eventually, it will be transformed into solving a symmetric positive definite system. In the above process, Cholesky decomposition can play an important role, and solving this symmetric positive definite system is the key step in bundle adjustment and also the bottleneck in computing.

[0006] However, in the process of bundle adjustment optimization, the Cholesky decomposition algorithm still requires a large amount of computing resources and time. In addition, on modern high-performance computing platforms, the existing implementations of Cholesky decomposition do not fully utilize the parallel computing capabilities of the hardware. Currently, the implementation of the Cholesky decomposition algorithm on ShenWei supercomputers depends on multiple libraries. In the LAPACK library, it can only be executed serially on the main core and cannot utilize the parallelism of the slave cores. In the xMath library, parallelism of a single-core group can be carried out, but for large-scale problems, there are problems such as slow processing speed and insufficient memory in the core group. Although the ScaLAPACK library can handle large-scale problems through distributed memory on ShenWei supercomputers, it cannot utilize the parallelism of the slave cores. Summary of the Invention

[0007] To overcome the deficiencies of the above-mentioned prior art, the present invention provides a heterogeneous parallel optimization method and system for Cholesky decomposition based on the ShenWei architecture. A secondary parallel Cholesky decomposition algorithm is constructed based on the ShenWei supercomputer to achieve efficient optimization of bundle adjustment.

[0008] To achieve the above object, one or more embodiments of the present invention provide the following technical solutions: The first aspect of the present invention provides a heterogeneous parallel optimization method for Cholesky decomposition based on the ShenWei architecture; A heterogeneous parallel optimization method for Cholesky decomposition based on the ShenWei architecture includes: Construct a symmetric positive definite matrix for bundle adjustment; Based on a distributed parallel allocation scheme, divide the symmetric positive definite matrix for bundle adjustment into sub-blocks, and complete matrix decomposition by iterating the divided sub-blocks; Allocate the divided sub-blocks to different processes, and exchange data through asynchronous communication between processes to perform coarse-grained task-level parallel acceleration; Utilize the master-slave core acceleration parallel feature of the ShenWei architecture to perform secondary parallel acceleration optimization on the POTRF, TRSM, SYRK, and GEMM operations in Cholesky decomposition respectively; Among them, for the SYRK and GEMM operations in the master-slave core acceleration parallel, map the column vectors of the matrix to the slave core array, and divide the columns according to the number of slave cores; optimize the calculation process through a double-buffer mechanism, vectorization operation, and loop unrolling to improve parallel efficiency; For the TRSM operation in the master-slave core acceleration parallel, decompose it into multiple TRSV triangular matrix vector solution operations, allocate them to the slave cores for parallel execution, and adopt the method of circular reading and data broadcasting to reduce data dependence and achieve efficient parallel computing.

[0009] As a further technical solution, the process of constructing the symmetric positive definite matrix for bundle adjustment is as follows: The bundle adjustment method optimizes the three-dimensional structure and camera parameters by indirectly minimizing the projection error, where is the vector synthesized by all the true projection point coordinates, is the vector synthesized by all the estimated projection point coordinates. When all the parameters are synthesized into the parameter vector p, there exists a non-linear function such that ; Iteratively minimize the sum of the squares of the reprojection errors between all the true projection point coordinates and the estimated projection point coordinates; in each iteration, the error amount to be minimized is: ; Wherein, J is the Jacobian matrix, is the parameter error amount of the current iteration, and the parameter is gradually optimized by the Gauss-Newton direction in each iteration ; is obtained by the following equation: ; Wherein, E is the identity matrix, let , and A is a symmetric positive definite matrix; is the damping factor to make A a symmetric positive definite matrix.

[0010] As a further technical solution, the process of sub-block partitioning the symmetric positive definite matrix for bundle adjustment based on the distributed parallel allocation scheme is as follows: According to the number of processes participating in the calculation, the symmetric positive definite matrix is partitioned into p×p sub-blocks, and the size of each sub-block is as follows: Wherein, , ; When , ; When ,

[0011] Wherein, represents whether the matrix can be equally divided into p*p sub-matrices; , represents the remainder operation; is the number of rows or columns of the matrix; is the number of processes; is the size of each sub-matrix, .

[0012] As a further technical solution, the process of allocating the divided sub-blocks to different processes, and the processes exchanging data through asynchronous communication to perform coarse-grained task-level parallel acceleration is as follows: After the main process performs the matrix Cholesky decomposition operation, it broadcasts the result to other processes in the same column through the broadcast operation; When other processes perform the TRSM operation, they use non-blocking communication to send the calculation result to the process that needs to receive the triangular solution result of the current process to ensure seamless connection between calculation and communication; When the process performs the SYRK and GEMM operations, it transmits data through asynchronous communication to ensure high overlap between the calculation task and the communication task.

[0013] As a further technical solution, the process of optimizing the calculation process through the double-buffering mechanism, vectorization operation and loop unrolling to improve the parallel efficiency is as follows: When calculating matrix C_sub, first read a column of data from matrix A_sub for calculation, and at the same time read the next column of data in the background to achieve the parallelization of data access and calculation.

[0014] As a further technical solution, the vectorization operation and loop unrolling are as follows: utilize the SIMD instruction set of the ShenWei architecture to perform vectorization operation on the matrix, and simultaneously calculate the loaded double data based on the SW many-core processor; through the loop unrolling scheme, calculate multiple columns of data of each slave core simultaneously.

[0015] As a further technical solution, for the TRSM operation in the master-slave core acceleration parallelization, the process of decomposing it into multiple TRSV operations and distributing them to the slave cores for parallel execution is as follows: For the m×n matrix C_sub, divide it into 64 rows, and each slave core processes m / 64 rows; reduce data dependence through loop reading and data broadcasting.

[0016] The second aspect of the present invention provides a heterogeneous parallel optimization system for Cholesky decomposition based on the ShenWei architecture.

[0017] A heterogeneous parallel optimization system for Cholesky decomposition based on the ShenWei architecture includes: A symmetric positive definite matrix construction module, which is configured to: construct a symmetric positive definite matrix for bundle adjustment; A sub-block division module, which is configured to: perform sub-block division on the symmetric positive definite matrix for bundle adjustment based on a distributed parallel allocation scheme, and complete matrix decomposition by iterating the divided sub-blocks; An MPI parallel module, which is configured to: allocate the divided sub-blocks to different processes, and the processes exchange data through asynchronous communication to perform coarse-grained task-level parallel acceleration; A master-slave core parallel module, which is configured to: utilize the master-slave core acceleration parallel characteristics of the ShenWei architecture to perform secondary parallel acceleration optimization on the POTRF, TRSM, SYRK, and GEMM operations in Cholesky decomposition respectively; Among them, for the GEMM and SYRK operations in the master-slave core acceleration parallelization, map the column vectors of the matrix to the slave core array and divide the columns according to the number of slave cores; optimize the calculation process through a double-buffer mechanism, vectorization operation, and loop unrolling to improve the parallel efficiency; For the TRSM operation in the master-slave core acceleration parallelization, decompose it into multiple TRSV operations, distribute them to the slave cores for parallel execution, and adopt loop reading and data broadcasting to reduce data dependence to achieve efficient parallel computing.

[0018] The third aspect of the present invention provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, the steps in a heterogeneous parallel optimization method for Cholesky decomposition based on the ShenWei architecture as described in the first aspect of the present invention are implemented.

[0019] The fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, the steps in a heterogeneous parallel optimization method for Cholesky decomposition based on the ShenWei architecture as described in the first aspect of the present invention are implemented.

[0020] The above one or more technical solutions have the following beneficial effects: A heterogeneous parallel optimization method and system for Cholesky decomposition based on the ShenWei architecture of the present invention minimize the memory bandwidth bottleneck and data access conflicts by combining task scheduling and data access patterns, making the data access during the calculation more efficient. Coarse-grained task-level parallel acceleration is performed through the MPI programming model, and an asynchronous communication method is adopted to overlap the computing tasks and communication tasks. The core calculation of the Cholesky decomposition algorithm is secondarily parallel accelerated through master-slave acceleration parallel. In the secondary parallel method, for the characteristics of GEMM, TRSM, and SYRK, the DMA data transfer technology is utilized to maximize the acquisition of the data required for secondary parallel calculation into the slave core LDM, and one acquisition and multiple reuses are achieved. At the same time, for GEMM, TRSM, and SYRK, the slave core parallelization is completed through the double-buffer mode, SIMD vectorization, and loop unrolling methods.

[0021] The advantages of the additional aspects of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. Description of the Drawings

[0022] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0023] Figure 1 It is a flowchart of the method for the first embodiment.

[0024] Figure 2 It is a diagram of the distributed parallel allocation scheme of the four-process coefficient symmetric positive definite matrix A in the first embodiment.

[0025] Figure 3 It is a flowchart of the four-process master core first-level parallel algorithm in the first embodiment.

[0026] Figure 4It is the flowchart of the slave core secondary parallel algorithm in the first embodiment.

[0027] Figure 5 It is the diagram of the data partitioning method in the GEMM operation in the first embodiment.

[0028] Figure 6 It is the speedup ratio diagram of the method adopted in the first embodiment compared with the xMath math library.

[0029] Figure 7 It is the system structure diagram of the second embodiment. Detailed implementation manners

[0030] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0031] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention.

[0032] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0033] Embodiment 1 This embodiment discloses a heterogeneous parallel optimization method for Cholesky decomposition based on the ShenWei architecture; As Figure 1 shown, a heterogeneous parallel optimization method for Cholesky decomposition based on the ShenWei architecture includes: Step S1, constructing a symmetric positive definite matrix for bundle adjustment.

[0034] The bundle adjustment method optimizes the three-dimensional structure and camera parameters by indirectly minimizing the projection error where is the vector synthesized by all real projection point coordinates, is the vector synthesized by all estimated projection point coordinates. By synthesizing all parameters into the parameter vector p, there exists a non-linear function such that ; Using the LM (Levenberg-Marquardt) method to iterate the sum of squares of the reprojection errors between all real projection point coordinates x and the estimated projection point coordinates to be minimized, that is, is minimized; in each iteration, the error quantity to be minimized is: ; In the formula, J is the Jacobian matrix, is the parameter error amount for the current iteration, and through the Gauss-Newton direction in each iteration the parameters are gradually optimized; It is obtained by the following equation: ; In the formula, E is the identity matrix, let , and A is a symmetric positive definite matrix; is the damping factor to make A a symmetric positive definite matrix.

[0035] Step S2, based on the distributed parallel allocation scheme, divide the symmetric positive definite matrix into sub-blocks, and perform iterative calculations on the divided sub-blocks to complete matrix decomposition; According to the number of processes p participating in the calculation, divide the symmetric positive definite matrix A into p×p sub-blocks, and the size of each sub-block is as follows: where, , ; When at that time ; When at that time,

[0036] In the formula, represents whether the matrix can be equally divided into p*p sub-matrices; , represents the remainder operation; is the number of rows or columns of the matrix; is the number of processes; is the size of each sub-matrix, .

[0037] Since the matrix A is symmetric, only the sub-blocks of the lower triangular or upper triangular part of A need to be calculated. Taking 4 processes as an example, as Figure 2 shown, the process numbers of the j-th column and the p-j-1-th column form p processes except the main process, where 0 represents the 0th process of the first iteration, 0'represents the 0th process of the second iteration, and so on. Then, the algorithm completes the matrix decomposition through p iterations. In each iteration, all processes will perform different calculation tasks according to the dependency relationship. Only after the main process finishes, the processes in the same column will perform tasks depending on the value of the main process. When this column is completed, the remaining blocks will depend on the value of this column for update. The main process is responsible for the first task of each iteration start and coordination and scheduling. In the i-th iteration, all related tasks of the i-th column will be processed first. The blocks after the i-th column are updated according to the value of this column and the value of the i-1-th iteration. When the i-th iteration is completed, the processes in the i-th column get the final value. Until the p-th iteration completes all tasks, at this time the entire decomposition process is completed.

[0038] Furthermore, the divided sub-blocks are assigned to different processes through the MPI programming model, and data exchange is performed between processes through asynchronous communication to achieve coarse-grained task-level parallel acceleration; Use Cholesky decomposition to decompose the matrix into the form of a lower triangular matrix and the transpose of a lower triangular matrix, that is , where , that is: ; All tasks will be completed in four iterations of Cholesky decomposition. The first iteration includes: (1) First , to obtain

[0039] (2) After obtaining the value of , the blocks in the same column as are based on , , to obtain . The above operation is called the TRSM operation.

[0040] (3) After obtaining the value of , the diagonal blocks in the same row as are based on = , = , = to obtain . The above value is not the final value and needs to continue the next iteration, which is called the SYRK operation.

[0041] (4) For the remaining blocks, = , = , = to obtain . It is not the final value and needs to continue the iteration, which is called the GEMM operation.

[0042] After this iteration is completed, the first column obtains the final value, and then the next iteration starts with the second column.

[0043] Furthermore, the master process performs the POTRF operation on to obtain , where i represents the number of iterations, and then sends to other processes in the same column through the broadcast operation MPI_Bcast for execution = , namely the TRSM operation. After these processes execute the TRSM operation, the calculation results are sent to other processes using non-blocking communication to ensure seamless connection between calculation and communication, and then the processes can continue to execute. = - , namely the SYRK operation, and while executing the SYRK operation, the value of is sent to other processes to prepare for GEMM. After the SYRK operation is completed, the GEMM operation is directly executed. Data is transmitted through asynchronous communication to ensure a high degree of overlap between the calculation task and the communication task, without waiting for transmission. And while executing GEMM, the value of SYRK is continuously transmitted to the previous process, making the calculation task and the communication task highly overlapped, achieving a time balance between the calculation task and the communication task, eliminating waiting and idle time to the greatest extent, ensuring their seamless connection during parallel processing, and thus improving the overall efficiency.

[0044] Furthermore, the master-slave core acceleration parallel feature using the ShenWei architecture also includes transmitting matrix data from the main memory to the LDM of the slave core to reduce the main memory access overhead; realizing batch data transmission through the DMA technology to improve the data access efficiency.

[0045] Step S3, using the master-slave core acceleration parallel feature of the ShenWei architecture, respectively perform secondary parallel acceleration optimization on the POTRF, TRSM, SYRK, and GEMM operations in the Cholesky decomposition; Combined with Figure 3 and Figure 4 , master-slave acceleration parallelism is a two-level parallel method supported by the ShenWei supercomputer system. The master core mainly completes the calculation and communication of the part that cannot be parallelized by multiple cores. When the slave cores are performing task calculations, the master core waits. The core calculation problems in this invention are the four operations of POTRF, TRSM, SYRK, and GEMM in the Cholesky decomposition, where POTRF represents the Cholesky decomposition of the matrix, TRSM represents matrix triangular solution, SYRK represents matrix rank-k update, and GEMM represents matrix multiplication. Each slave core of the ShenWei many-core processor has a high-speed local local data storage space (Local Data Memory, LDM). DMA provides the function of batch data transmission between the main memory and the LDM, converting direct access to the memory into access to the local LDM. The efficiency of this method is much faster than directly accessing the main memory. Therefore, in the subsequent operations, the data is first read into the LDM and then the calculation task is executed. Step S4. For the GEMM and SYRK operations in the master-slave core acceleration parallelization, map the column vectors of the matrix to the slave core array and divide the columns according to the number of slave cores. Optimize the calculation process through the double-buffering mechanism, vectorization operation, and loop unrolling technology to improve the parallel efficiency. Combined with Figure 5 , the calculation process of GEMM is to calculate the product of matrices A_sub and B_sub to obtain an intermediate matrix, and then according to the weight coefficients and , sum the matrix C_sub with this intermediate matrix and accumulate the result to the corresponding position of the C_sub matrix. , but the GEMM in Cholesky decomposition is to calculate . The GEMM operation is a three-layer loop structure. To calculate each element of the matrix C_sub, it is necessary to read one row of the matrix A_sub and one column of the matrix B_sub. The entire process requires repeated reading of the matrices A_sub and B_sub, and its implementation efficiency is very low, but it has strong parallelism because each of its columns is independent.

[0046] Based on the above characteristics, in this embodiment, for a C_sub matrix, first map the column vectors of the matrix C_sub to the CPE array. To ensure load balancing, each slave core reads n / 64 column elements of C_sub into the LDM, and it is continuous n / 64 columns, that is, the matrix C_sub is divided into 64 columns (64 represents the number of slave cores used), and then each slave core corresponds to one of the columns. If the data is too much to be loaded into the LDM at one time, then it will be read in a loop, and each time the maximum number of columns that the slave core can put into the LDM will be read. In this way, the CPE array executes the GEMM operation in parallel to obtain high performance. For the matrix A_sub, in the i-th loop, only one slave core reads the i-th column of A_sub and then this slave core broadcasts this column to all slave cores, rather than each slave core reading the same column data of A_sub. For the matrix B_sub, then in the i-th loop, read the n / 64 data starting from the position until all columns of A_sub are looped through, and at this time the task is completed.

[0047] In terms of optimization, through the double-buffering mechanism, first, a column of A_sub needs to be read in to calculate C_sub, and while calculating C_sub, the next column of A_sub is continuously read; second, through vectorization operations, the main core of the SW many-core processor supports 256-bit SIMD extension instructions, and the slave cores support 512-bit SIMD extension instructions. For the double type, 8 double data can be loaded at a time, and then these 8 data are calculated simultaneously. When using, it is necessary to ensure 64-byte alignment; finally, through the loop unrolling scheme, multiple columns of data of each slave core are calculated simultaneously. For example, if each slave core processes x columns of data, then x vector registers are used to increase the parallelism of the program. Since each slave core has 32 vector registers, the vector registers used cannot exceed 32. When the last data cannot be aligned, it is processed separately. Step S5: For the TRSM operation in the master-slave core accelerated parallelism, it is decomposed into multiple TRSV operations and assigned to the slave cores for parallel execution. The loop reading and data broadcasting methods are adopted to reduce data dependencies and achieve efficient parallel computing.

[0048] Take the TRSV operation to illustrate the TRSM operation, that is, when B is a vector, it is the TRSV operation. For the TRSV operation, first calculate B[0,j] through the value of matrix A, and then update the subsequent elements B[0,k] (k>j) according to the value of B[0,j]. And TRSM can be regarded as the implementation of multiple TRSV operations. Therefore, multiple TRSVs are assigned to 64 slave cores for processing, thereby realizing the parallel TRSM operation. For a C_sub matrix, it is directly divided into 64 rows, and each slave core processes m / 64 rows. Subsequently, a scheme similar to GEMM is adopted.

[0049] Furthermore, the present invention further verifies the effectiveness of the method adopted by the present invention through test experiments. The test environment is the prototype of the Shenwei new generation supercomputer. Taking the Gaussian kernel matrix based on image pixel values in the field of image processing as a test case, the matrix scale is set to 20480*20480, and the test results are as Figure 6 shown. Figure 6 shows the speedup ratio compared with the high-performance math library xMath specially designed for the Shenwei many-core processor at a maximum of 32 core groups. It can be seen from Figure 6 that the maximum speedup ratio reaches 4.43 times.

[0050] Embodiment 2 This embodiment discloses a heterogeneous parallel optimization system for Cholesky decomposition based on the Shenwei architecture; As Figure 7 shown, a heterogeneous parallel optimization system for Cholesky decomposition based on the Shenwei architecture includes: A symmetric positive definite matrix construction module, which is configured to: construct a symmetric positive definite matrix for bundle adjustment; A sub-block division module, which is configured to: perform sub-block division on the symmetric positive definite matrix for bundle adjustment based on a distributed parallel allocation scheme, and perform iterative matrix decomposition on the divided sub-blocks; An MPI parallel module, which is configured to: allocate the divided sub-blocks to different processes, exchange data between the processes through asynchronous communication, and perform coarse-grained task-level parallel acceleration; A master-slave core parallel module, which is configured to: utilize the master-slave core acceleration parallel characteristics of the ShenWei architecture to perform secondary parallel acceleration optimization on the POTRF, TRSM, SYRK, and GEMM operations in the Cholesky decomposition respectively; Among them, for the GEMM and SYRK operations in the master-slave core acceleration parallel, map the column vectors of the matrix to the slave core array, and divide the columns according to the number of slave cores; optimize the calculation process through a double buffering mechanism, vectorization operation, and loop unrolling to improve the parallel efficiency; For the TRSM operation in the master-slave core acceleration parallel, decompose it into multiple TRSV operations, allocate them to the slave cores for parallel execution, and adopt the method of circular reading and data broadcasting to reduce data dependence and achieve efficient parallel computing.

[0051] Embodiment III The purpose of this embodiment is to provide a computer-readable storage medium.

[0052] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in a heterogeneous parallel optimization method for Cholesky decomposition based on the ShenWei architecture as described in Embodiment 1.

[0053] Embodiment IV The purpose of this embodiment is to provide an electronic device.

[0054] An electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in a heterogeneous parallel optimization method for Cholesky decomposition based on the ShenWei architecture as described in Embodiment 1.

[0055] The steps involved in the devices in Embodiments II, III, and IV above correspond to those in Method Embodiment 1. For specific implementation manners, reference may be made to the relevant description part of Embodiment 1. The term "computer-readable storage medium" should be understood to include a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.

[0056] Those skilled in the art should understand that the various modules or steps of the present invention described above can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0057] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications or deformations that can be made without creative efforts on the basis of the technical solutions of the present invention are still within the protection scope of the present invention.

Claims

1. A Cholesky decomposition heterogeneous parallel optimization method based on Shenwei architecture, characterized in that: include: Construct the symmetric positive definite matrix for bundle adjustment; The symmetric positive definite matrix of bundle adjustment is divided into sub-blocks based on a distributed parallel allocation scheme, and the matrix decomposition is completed by iteratively performing the divided sub-blocks. The divided sub-blocks are assigned to different processes, and the processes exchange data through asynchronous communication to perform coarse-grained task-level parallel acceleration; Using the master-slave core acceleration parallelism of the Shenwei architecture, two-level parallel acceleration optimization is performed on the POTRF, TRSM, SYRK and GEMM operations in the Cholesky decomposition; For the SYRK and GEMM operations in the master-slave core acceleration parallelism, the column vectors of the matrix are mapped to the slave core array, and the columns are divided according to the number of slave cores; the calculation process is optimized through the double buffer mechanism, vectorized operation and loop unrolling to improve the parallel efficiency; For the TRSM operation in the master-slave core acceleration parallelism, it is decomposed into multiple TRSV triangular matrix vector solution operations, which are assigned to the slave cores for parallel execution. The loop reading and data broadcasting methods are used to reduce data dependency and achieve efficient parallel computing.

2. The Cholesky decomposition heterogeneous parallel optimization method based on the Shenwei architecture as claimed in claim 1, characterized in that: The process of constructing the symmetric positive definite matrix of bundle adjustment is: The bundle adjustment method optimizes the 3D structure and camera parameters by indirectly minimizing the projection error, where is the vector synthesized from the coordinates of all real projection points, is the vector synthesized from the coordinates of all estimated projection points, and all parameters are synthesized into the parameter vector p. Then there exists a nonlinear function such that ; Iterate to minimize the sum of squares of the reprojection errors between all true projection point coordinates and the estimated projection point coordinates; in each iteration, the minimized error amount is: ; Where J is the Jacobian matrix, is the parameter error of the current iteration, which is calculated by the Gauss-Newton direction in each iteration. The parameters are gradually optimized; Calculated by the following equation: ; In the formula, E is the unit matrix, let , A is a symmetric positive definite matrix; is the damping factor that makes A a symmetric positive definite matrix.

3. The Cholesky decomposition heterogeneous parallel optimization method based on the Shenwei architecture as claimed in claim 1, characterized in that: The process of sub-blocking the symmetric positive definite matrix of the bundle adjustment based on the distributed parallel allocation scheme is as follows: According to the number of processes involved in the calculation, the symmetric positive definite matrix is ​​divided into p×p sub-blocks, each sub-block The sizes are as follows: in, , ; when hour ; when hour, In the formula, represents whether the matrix can be equally divided into p*p sub-matrices; , Represents the remainder operation; is the number of rows or columns of the matrix; is the number of processes; For the size of each submatrix, .

4. The Cholesky decomposition heterogeneous parallel optimization method based on Shenwei architecture as claimed in claim 1, characterized in that: The process of allocating the divided sub-blocks to different processes, exchanging data between the processes through asynchronous communication, and performing coarse-grained task-level parallel acceleration is as follows: After the main process performs the Cholesky decomposition operation on the matrix, it broadcasts the result to other processes in the same column through the broadcast operation; When other processes perform TRSM operations, they use non-blocking communication to send the calculation results to the processes that need to receive the triangulation solution results of the current process, ensuring seamless connection between calculation and communication; When executing SYRK and GEMM operations, processes transmit data through asynchronous communication to ensure a high degree of overlap between computing tasks and communication tasks.

5. The Cholesky decomposition heterogeneous parallel optimization method based on Shenwei architecture as claimed in claim 1, characterized in that: The process of optimizing the calculation process by double buffering mechanism, vectorized operation and loop unrolling to improve parallel efficiency is as follows: When calculating the matrix C_sub, first read a column of data from the matrix A_sub for calculation, and at the same time read the next column of data in the background to achieve parallelization of data access and calculation.

6. The Cholesky decomposition heterogeneous parallel optimization method based on Shenwei architecture as claimed in claim 1, characterized in that: The vectorized operation and loop expansion are as follows: using the SIMD instruction set of the Shenwei architecture to perform vectorized operations on the matrix, and based on the SW multi-core processor to simultaneously calculate the loaded double data; through the loop expansion scheme, multiple columns of data of each slave core are calculated simultaneously.

7. The Cholesky decomposition heterogeneous parallel optimization method based on Shenwei architecture as claimed in claim 1, characterized in that: For the TRSM operation in the master-slave core acceleration parallelism, the process of decomposing it into multiple TRSV operations and assigning them to the slave cores for parallel execution is as follows: For the m×n matrix C_sub, it is divided into 64 rows, and each slave core processes m / 64 rows; data dependency is reduced by loop reading and data broadcasting.

8. A Cholesky decomposition heterogeneous parallel optimization system based on the Shenwei architecture, characterized by: include: A symmetric positive definite matrix construction module, which is configured to: construct a symmetric positive definite matrix for bundle adjustment; A sub-block partitioning module is configured to: partition the symmetric positive definite matrix of the bundle adjustment into sub-blocks based on a distributed parallel allocation scheme, and iterate the partitioned sub-blocks to complete matrix decomposition; The MPI parallel module is configured to: assign the divided sub-blocks to different processes, exchange data between the processes through asynchronous communication, and perform coarse-grained task-level parallel acceleration; The master-slave core parallel module is configured to: utilize the master-slave core acceleration parallel characteristics of the Shenwei architecture to perform two-level parallel acceleration optimization on the POTRF, TRSM, SYRK and GEMM operations in the Cholesky decomposition; For GEMM and SYRK operations in master-slave core acceleration parallelism, the column vectors of the matrix are mapped to the slave core array, and the columns are divided according to the number of slave cores; the calculation process is optimized through double buffering mechanism, vectorized operation and loop unrolling to improve parallel efficiency; For the TRSM operation in the master-slave core acceleration parallelism, it is decomposed into multiple TRSV operations and assigned to the slave cores for parallel execution. The loop reading and data broadcasting methods are used to reduce data dependency and achieve efficient parallel computing.

9. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, the steps in a Cholesky decomposition heterogeneous parallel optimization method based on the Shenwei architecture as described in any one of claims 1 to 7 are implemented.

10. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the Cholesky decomposition heterogeneous parallel optimization method based on the Shenwei architecture are implemented as described in any one of claims 1-7.

Citation Information

Patent Citations

  • A GEMM (general matrix-matrix multiplication) high-performance realization method based on a domestic SW 26010 many-core CPU

    CN107168683A

  • High-performance parallel implementation method of K-means algorithm on domestic Sunway 26010 multi-core processor

    CN108509270A

  • Modal parallel computing method and system for heterogeneous many-core parallel computer

    CN111651208A

  • PCG precondition sub-method oriented to heterogeneous many-core architecture

    CN114611060A

Cited By

  • Memory architecture-oriented dual-precision general matrix multiplication optimization method and system

    CN121278223A

  • Memory architecture oriented double precision general matrix multiplication optimization method and system

    CN121278223B

  • Data processing device and method, electronic equipment and storage medium

    CN121807558A