CFD-based linear matrix solution acceleration method and device, and readable medium

By converting LDU format matrices into ELLB format and combining it with the PCG solver and AMG preprocessing steps on the GPU accelerator card, the problem of slow iterative solution of linear matrix systems in CFD is solved, and efficient matrix-vector multiplication calculation and stable solution process are achieved.

CN120687719APending Publication Date: 2025-09-23SHANGHAI NUCLEAR ENGINEERING RESEARCH & DESIGN INSTITUTE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510850525.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing iterative solution methods for linear matrix systems are slow in CFD, leading to performance bottlenecks, especially when matrix elements are frequently accessed.

Method used

The LDU format matrix is ​​converted to the ELLB format matrix, and the PCG solver and AMG preprocessing steps on the GPU accelerator card are used in combination with the CUDA kernel function to perform matrix-vector multiplication calculations to optimize the data transmission process.

Benefits of technology

The calculation speed of matrix-vector multiplication is significantly improved, memory usage is reduced, and exact solutions are obtained in fewer iterations, ensuring the stability and accuracy of the solution process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687719A_ABST
    Figure CN120687719A_ABST
Patent Text Reader

Abstract

The invention provides a CFD-based linear matrix solving acceleration method and device and a readable medium. The method comprises the steps that 1, a CPU reads grid information and generates an LDU format matrix; 2, the LDU format matrix is converted into an ELLB format matrix; 3, the ELLB format matrix and the known vector are transmitted into a GPU memory, and the GPU calls a PCG solver to iteratively calculate a solution vector according to the ELLB format matrix and the known vector; and step 4, converting the solution vector into an identifiable solution vector in the LDU format, and transmitting the solution vector in the LDU format back to the CPU by the GPU. According to the method, the LDU format matrix is converted into the ELLB format matrix, and the ELLB format and GPU acceleration are utilized, so that the calculation speed of matrix vector multiplication can be remarkably improved, and the overall solving time is shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention mainly relates to the field of computational fluid dynamics, and in particular to an acceleration method, device and readable medium for solving linear matrices based on CFD. Background Art

[0002] The design, optimization, and safe operation of nuclear power fuel assemblies rely heavily on computational fluid dynamics (CFD) technology. CFD simulations of the complex flow, heat transfer, and multi-physics coupling of nuclear power fuel assemblies not only improve the thermal and hydraulic performance of the assemblies but also provide critical data for safety assessments.

[0003] In computational fluid dynamics, matrices obtained by discretizing physical models using the finite volume method are stored in an LDU format. This format has the advantage of being space-efficient for sparse matrices, as most zero elements do not occupy additional storage space, especially when the non-zero elements are symmetrically distributed about the diagonal. However, while LDU matrices save space, they also introduce the problem of low access efficiency. This is particularly true during the iterative solution of linear matrix systems, where the large number of matrix-vector multiplication operations necessitates frequent access to matrix elements, creating a performance bottleneck.

[0004] Therefore, there is an urgent need for an acceleration method for linear matrix solution based on CFD to improve the large-scale solution speed of CFD software. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide an acceleration method, device and readable medium for solving linear matrices based on CFD, so as to solve the problem of slow solution speed of the existing iterative solution method of linear matrix system.

[0006] To solve the above technical problems, the present invention provides an accelerated method for solving linear matrices based on CFD, comprising:

[0007] Step 1: The CPU reads the grid information and generates an LDU format matrix;

[0008] Step 2: Convert the LDU format matrix into an ELLB format matrix;

[0009] Step 3: The ELLB format matrix and the known vector are transferred to the GPU memory, and the GPU calls the PCG solver to iteratively calculate the solution vector according to the ELLB format matrix and the known vector;

[0010] Step 4: Convert the solution vector back into a recognizable LDU format solution vector, and the GPU returns the LDU format solution vector to the CPU.

[0011] Optionally, converting the LDU format matrix into an ELLB format matrix includes:

[0012] Step 2.1: Traverse each row of the LDU format matrix and record the value of the non-zero element and the corresponding column index;

[0013] Step 2.2: Sort the non-zero elements according to the column index and store them in an ELLB format structure;

[0014] Step 2.3: Generate two arrays: one to store the values ​​of non-zero elements and the other to store the corresponding column indices;

[0015] Step 2.4: Generate a bit mask array, where each bit in the bit mask array corresponds to whether an element is valid.

[0016] Optionally, the GPU calling a PCG solver to iteratively calculate a solution vector according to the ELLB format matrix and the known vector includes:

[0017] Step 3.1: Setting the initial parameters of the PCG solver, wherein the initial parameters include a solution vector, a residual vector, a preprocessing vector, and a search direction vector;

[0018] Step 3.2: Perform matrix-vector multiplication calculation using the CUDA kernel function written on the GPU, and update the solution vector and the residual vector according to the calculation results;

[0019] Step 3.3: Determine whether the convergence condition is met based on the residual vector. If the condition is met, use the updated solution vector as the final solution vector. If the condition is not met, update the search direction vector and then enter step 3.2 for iterative calculation until the convergence condition is met.

[0020] Optionally, the method further includes: before step 3.2, converting the ELLB format matrix into a CSR format matrix; constructing an AMG preprocessor on the GPU based on the converted CSR format matrix; and replacing the original preprocessing steps in the PCG solver with the AMG preprocessing steps implemented by the AMG preprocessor.

[0021] Optionally, the method further includes: using multiple GPUs to calculate the solution vector, generating an initial grid before transferring the ELLB format matrix and the known vector into the memory of the GPU, and dividing the grid to ensure that each sub-grid is evenly distributed on each GPU.

[0022] Optionally, converting the solution vector back into a recognizable LDU format solution vector includes: traversing the solution vector and filling the values ​​of the solution vector into corresponding positions of the LDU format matrix.

[0023] Optionally, the method further includes: performing CFD simulation using the LDU format solution vector, and adjusting the grid division and the parameters of the PCG solver according to the simulation results.

[0024] Optionally, the CUDA API is used to transfer data between the CPU and the GPU.

[0025] To solve the above technical problems, the present invention provides an acceleration device for linear matrix solving based on CFD, comprising: a CPU and multiple GPUs, each GPU comprising a format conversion module and a calculation module; wherein the CPU is used to read grid information and generate an LDU format matrix, and receive an LDU format solution vector calculated by the GPU; the format conversion module is used to convert the LDU format matrix into an ELLB format matrix, transfer the ELLB format matrix and the known vector into the GPU memory, and convert the solution vector calculated by the calculation module back into a recognizable LDU format solution vector; the calculation module is used to call a PCG solver to iteratively calculate a solution vector based on the ELLB format matrix and the known vector.

[0026] In order to solve the above technical problems, a computer-readable medium storing computer program code is provided. When the computer program code is executed by a processor, the accelerated method for solving a linear matrix based on CFD as described in any one of the above items is implemented.

[0027] Compared with the prior art, the present invention has the following advantages:

[0028] The acceleration method, device and readable medium for CFD-based linear matrix solution of the present invention convert the LDU format matrix into the ELLB format matrix. Compared with the traditional LDU format, the ELLB format saves a lot of space and reduces memory usage. By using the ELLB format and GPU acceleration, the calculation speed of matrix-vector multiplication can be significantly improved, and the overall solution time can be reduced. By using the PCG solver and AMG preprocessing steps, an accurate solution can be obtained within a smaller number of iterations, ensuring the stability and accuracy of the solution process, and is suitable for various computational fluid dynamics problems. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The accompanying drawings are included to provide a further understanding of the present application. They are incorporated into and constitute a part of this application. The accompanying drawings illustrate embodiments of the present application and, together with this specification, serve to explain the principles of the present application. In the accompanying drawings:

[0030] Figure 1 4 is a flow chart of an accelerated method for solving a linear matrix based on CFD according to an embodiment of the present invention.

[0031] Figure 21 is a flowchart of converting an LDU format matrix into an ELLB format matrix according to an embodiment of the present invention.

[0032] Figure 3 4 is a flowchart of calculating a solution vector of an ELLB format matrix and a known vector according to an embodiment of the present invention.

[0033] Figure 4 It is a system block diagram of an acceleration device for linear matrix solving based on CFD according to an embodiment of the present invention. DETAILED DESCRIPTION

[0034] To more clearly illustrate the technical solutions of the embodiments of this application, the following is a brief introduction to the drawings required for describing the embodiments. Obviously, the drawings described below are merely examples or embodiments of this application. Those skilled in the art can apply this application to other similar scenarios based on these drawings without inventive effort. Unless otherwise apparent from the context or otherwise noted, the same reference numerals in the figures represent the same structure or operation.

[0035] Flowcharts are used in this application to illustrate the operations performed by systems according to embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the various steps may be processed in reverse order or simultaneously. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.

[0036] Figure 1 FIG. 1 is a flow chart of an accelerated method for solving a linear matrix based on CFD according to an embodiment of the present invention. Figure 1 As shown, the acceleration method 100 for linear matrix solution based on CFD includes:

[0037] Step S1: The CPU reads the grid information and generates an LDU format matrix.

[0038] The CFD software on the CPU reads the mesh information and generates a matrix in LDU format (abbreviated as LDU format matrix). The advantage of LDU format is that for sparse matrices, most of the zero elements do not occupy additional storage space, making it more space-efficient, especially when the non-zero elements are symmetrically distributed about the diagonal. However, while LDU format matrices save space, they also introduce the problem of low access efficiency. This is especially true during the iterative solution of linear matrix systems, where a large number of matrix-vector multiplication operations require frequent access to matrix elements, creating a performance bottleneck.

[0039] Step S2: Convert the LDU format matrix into an ELLB format matrix.

[0040] This application converts the LDU format matrix into the ELLB format matrix, and the ELLB format matrix refers to the matrix in the ELLB format. The LDU format matrix is ​​a matrix after the lower triangular-diagonal-upper triangular decomposition, which usually includes three parts: the lower triangular matrix L, the diagonal matrix D, and the upper triangular matrix U. The ELLB (ELLPACK-Bit) format is a sparse matrix storage format, which is mainly optimized for GPUs and can efficiently handle structured grid problems, such as matrices in finite difference or finite element methods. The characteristic of the ELLB format is that non-zero elements are stored in rows, each row has a fixed number of elements, and the insufficient part is padded with a filling value (such as zero), and a bit mask is used to identify the position of valid elements, so that the storage space can be reduced while maintaining efficient access.

[0041] Figure 2 FIG. 1 is a flow chart of converting an LDU format matrix into an ELLB format matrix according to an embodiment of the present invention. Figure 2 As shown, the conversion of the LDU format matrix to the ELLB format matrix includes:

[0042] Step S2.1: traverse each row of the LDU format matrix and record the values ​​of non-zero elements and the corresponding column index;

[0043] Step S2.2: Sort these non-zero elements by column index and store them in a structure in ELLB format.

[0044] For each row, the nonzero elements are sequentially placed into the corresponding rows of the value array and column index array. If the number of nonzero elements in a row is less than K, the remaining positions are filled with zeros or other fill values. K is the maximum number of nonzero elements allowed in each row.

[0045] Step S2.3: Generate two arrays: one to store the values ​​of non-zero elements, and the other to store the corresponding column indices.

[0046] Step S2.4: Generate a bit mask array, where each bit in the bit mask array corresponds to whether an element is valid.

[0047] In one embodiment of the present invention, a format conversion function is written in a GPU accelerator card, and the LDU format matrix is ​​converted into an ELLB format matrix through the format conversion function. Furthermore, the format conversion function can be implemented based on the CUDA platform embedded in the GPU accelerator card. CUDA (Compute Unified Device Architecture) is a programming interface provided by Nvidia for writing programs for its GPU. In CUDA, you can express the calculation you want to run on the GPU in a form similar to a C / C++ function. This function is called a kernel. The kernel performs parallel calculations on the digital vectors provided to it as function parameters. The working principle of CUDA is that in GPU-accelerated applications, serial workloads run on the CPU, while the computationally intensive parts run in parallel on thousands of cores of the GPU. Through the above steps, the LDU format matrix can be efficiently converted into an ELLB format matrix, making full use of the parallel computing capabilities of the GPU.

[0048] In other embodiments, a format conversion function may also be written in the CPU to convert the LDU format matrix into the ELLB format matrix through the format conversion function. This application does not limit the location of the format conversion function.

[0049] Step S3: The ELLB format matrix and known vector are passed into the GPU memory, and the GPU calls the PCG solver to iteratively calculate the solution vector based on the ELLB format matrix and known vector.

[0050] In the linear equation system Ax=b, A represents the ELLB format matrix, b is the known vector, and x is the solution vector to be solved. In this application, the GPU calls the preconditioned conjugate gradient solver (PCG) to iteratively calculate the solution vector of the ELLB format matrix A and the known vector b. Figure 3 As shown, the iterative calculation of the solution vector of the linear matrix system includes:

[0051] Step S3.1: Set the initial parameters of the PCG solver, which include the solution vector, residual vector, preprocessing vector and search direction vector.

[0052] Copy the matrix A and vector b to the GPU memory, and allocate GPU memory to vectors such as the solution vector x, residual vector r, preprocessing vector z, and search direction vector p.

[0053] Set the initial solution vector x to 0 or a guess value; the residual vector r = bA*x;

[0054] The preprocessing vector z=M-1*r, the Jacobi preprocessing uses the inverse of the diagonal elements of the matrix A as the preprocessing matrix M; the search direction vector p=z.

[0055] Step S3.2: Perform matrix-vector multiplication using the CUDA kernel function written on the GPU, and update the solution vector and residual vector based on the calculation results;

[0056] For example, cuSPARSE is used to calculate the matrix-vector multiplication of the matrix A and the search direction vector p:

[0057] A_p=A*p。

[0058] CUSPARSE contains a series of basic linear algebra subroutines for processing sparse matrices and is part of the CUDA library.

[0059] Use cuBLAS's axpy function to update the solution vector and residual vector:

[0060] x=x+α*p

[0061] r=r-α*A_p

[0062] Among them, α is a scalar. The cuBLAS library under CUDA is a library that contains some basic matrix operation functions, which can improve the efficiency of matrix operations.

[0063] Step S3.3: Determine whether the residual vector meets the convergence condition. If so, proceed to step S3.4. If not, update the search direction vector and then proceed to step 3.2 for iterative calculation until the convergence condition is met.

[0064] Step S3.4: Use the updated solution vector as the final solution vector.

[0065] In one embodiment, the norm of the residual vector r is calculated using the cuBLAS nrm2 function. If the norm is less than a threshold, the iteration is terminated and the updated solution vector x is used as the final solution vector. If the norm is greater than or equal to the threshold, the search direction vector is updated: p = z + β * p, where β is the ratio of the current residual vector to the previous residual vector.

[0066] Algebraic multigrid (AMG) is an effective preconditioning technique for solving large-scale sparse linear systems, and is particularly suitable for matrices obtained by discretizing elliptic partial differential equations.

[0067] Optionally, before step S3.2, it also includes converting the ELLB format matrix into a CSR format matrix; using the converted CSR format matrix, constructing an AMG preprocessor, constructing the AMG preprocessor includes constructing an AMG hierarchy, the hierarchy includes coarse grid selection, interpolation operator and coarse grid operator; integrating the AMG preprocessing step into the PCG solver to replace the original preprocessing step.

[0068] The AMG preprocessing step is typically a V-loop or W-loop. In each iterative solution, the AMG preprocessing step is called to solve the residual equation. In an iterative solver (such as PCG), the implementation of the preprocessing step M-1 is a single AMG loop. The following is the program code for using AMG preprocessing in a PCG solver according to one embodiment of the present invention:

[0069]

[0070] Here, amg_vcycle(amg,r,z) means: amg_vcycle() is the V-cycle in the AMG preprocessing step, amg is the AMG preprocessor, r is the residual vector, and z is the preprocessing vector. This line of code replaces the original preprocessing vector z in the PCG solver with the V-cycle in the AMG preprocessing step. As can be seen above, this application uses the AMG preprocessing step to replace the original preprocessing step (such as Jacobi preprocessing), and the subsequent PCG steps remain unchanged.

[0071] By converting ELLB format matrices to CSR format and then using the AMG preconditioner built on the GPU, the convergence speed of the iterative solver can be significantly improved. Although format conversion and AMG construction have additional overhead, for matrices with poor condition numbers (such as anisotropic problems), AMG preconditioning can greatly reduce the number of iterations, thereby achieving an overall speedup.

[0072] During the iterative solution process, matrix-vector multiplication (SpMV) still uses the ELLB format for optimal performance. However, matrix-vector multiplications involved in the AMG preprocessing step (such as A*z_smooth) can use the CSR format. Combining the AMG preprocessor with the ELLB format can achieve a 5-50x speedup on the GPU, making it particularly suitable for scientific computing problems with billions of degrees of freedom.

[0073] Optionally, the method further includes: before transferring the ELLB format matrix and the known vector into the GPU memory, generating an initial grid and dividing the grid to ensure that each sub-grid is evenly distributed on each GPU.

[0074] Based on the mesh information, a global mesh is generated on the CPU and divided into subgrids equal to the number of GPUs. Each subgrid is carefully designed to contain a similar number of points to ensure balanced computational load. The data for each subgrid (including mesh coordinates, cell information, boundary conditions, etc.) is then distributed to the corresponding GPU. By evenly distributing the mesh across multiple GPUs, combined with the AMG preprocessor and the ELLB storage format, the efficiency of solving large-scale linear systems can be significantly improved.

[0075] Step S4: converting the solution vector back into a recognizable LDU format solution vector, and the GPU transmitting the LDU format solution vector back to the CPU.

[0076] Traverse the solution vector and fill the solution vector value into the corresponding position of the LDU format matrix to ensure data consistency and accuracy during the conversion process.

[0077] Optionally, the CUDA API is used to transfer data between the CPU and GPU. This includes: allocating memory on the GPU using cudaMalloc, copying data from the CPU to the GPU using cudaMemcpy, executing calculations on the GPU by launching CUDA kernel functions, copying results from the GPU back to the CPU using cudaMemcpy, and releasing GPU memory using cudaFree. In CUDA programming, the cudaMalloc function is the basic API for allocating space in GPU memory. cudaMemcpy is a core function used to transfer data between host (CPU) memory and device (GPU) memory. The cudaFree() function is used to release memory space previously allocated by the cudaMalloc() function.

[0078] By utilizing the CUDA API to transfer data between the CPU and GPU, the data transfer process is optimized, the communication overhead between the CPU and GPU is reduced, and the continuity and efficiency of the computing process are ensured.

[0079] In some embodiments, the acceleration method for CFD-based linear matrix solving further includes: performing CFD simulation using an LDU format solution vector, and adjusting meshing and PCG solver parameters according to the simulation results.

[0080] The accelerated method for linear matrix solution based on CFD of the present invention converts the LDU format matrix into the ELLB format matrix. Compared with the traditional LDU format, the ELLB format saves a lot of space and reduces memory usage. By using the ELLB format and GPU acceleration, the calculation speed of matrix-vector multiplication can be significantly improved, and the overall solution time can be reduced. By using the PCG solver and the AMG preprocessing step, an accurate solution can be obtained within a smaller number of iterations, ensuring the stability and accuracy of the solution process, and is suitable for various computational fluid dynamics problems.

[0081] The present invention also provides an acceleration device for solving linear matrix based on CFD. Figure 4 FIG. 1 is a system block diagram of an acceleration device for linear matrix solving based on CFD according to an embodiment of the present invention. Figure 4 As shown, the acceleration device 400 for linear matrix solution based on CFD includes: a CPU and multiple GPUs. Each GPU includes a format conversion module, a memory module, and a computing module, and each computing module includes a PCG solver.

[0082] The CPU reads the grid information and generates a global grid based on it. This grid is then divided into subgrids equal to the number of GPUs, ensuring that each subgrid contains a similar number of points to ensure balanced computational load. The data for each subgrid (including grid coordinates, cell information, boundary conditions, etc.) is distributed to the corresponding GPU. The default storage format for each subgrid is LDU format, with each subgrid corresponding to an LDU-format matrix.

[0083] The format conversion module of each GPU is used to convert the LDU format matrix into the ELLB format matrix and transfer the ELLB format matrix and known vectors into the GPU memory.

[0084] The computing module of each GPU calls the built-in PCG solver to iteratively calculate the solution vector of the ELLB format matrix and the known vector.

[0085] The format conversion module of each GPU is further used to convert the solution vector back into a recognizable LDU format solution vector and feed the LDU format solution vector back to the CPU.

[0086] Optionally, data transmission between the CPU and the GPU is implemented using a CUDA API.

[0087] The present application also includes a computer-readable medium storing computer program code, which, when executed by a processor, implements the aforementioned accelerated method for CFD-based linear matrix solution.

[0088] When the accelerated method for linear matrix solving based on CFD is implemented as a computer program, it can also be stored in a computer-readable storage medium as an article of manufacture. For example, a computer-readable storage medium can include, but is not limited to, a magnetic storage device (e.g., a hard disk, a floppy disk, a magnetic stripe), an optical disk (e.g., a compact disk (CD), a digital versatile disk (DVD)), a smart card, and a flash memory device (e.g., an electrically erasable programmable read-only memory (EPROM), a card, a stick, a key drive). In addition, the various storage media described herein can represent one or more devices and / or other machine-readable media for storing information. The term "machine-readable medium" can include, but is not limited to, wireless channels and various other media (and / or storage media) that can store, contain, and / or carry code and / or instructions and / or data.

[0089] The basic concepts have been described above. It will be apparent to those skilled in the art that the above disclosures are merely illustrative and do not constitute limitations on this application. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and amendments to this application. Such modifications, improvements, and amendments are suggested in this application and remain within the spirit and scope of the exemplary embodiments of this application.

[0090] At the same time, this application uses specific terms to describe the embodiments of this application. For example, "one embodiment," "an embodiment," and / or "some embodiments" refer to a certain feature, structure, or characteristic related to at least one embodiment of this application. Therefore, it should be emphasized and noted that "one embodiment," "an embodiment," or "an alternative embodiment" mentioned twice or multiple times in different locations in this specification does not necessarily refer to the same embodiment. In addition, certain features, structures, or characteristics in one or more embodiments of this application may be appropriately combined.

[0091] Some aspects of the present application can be performed entirely by hardware, entirely by software (including firmware, resident software, microcode, etc.), or by a combination of hardware and software. The above hardware or software can be referred to as "data blocks", "modules", "engines", "units", "components" or "systems". The processor can be one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DAPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors or combinations thereof. In addition, various aspects of the present application may be expressed as computer products located in one or more computer-readable media, which include computer-readable program code. For example, computer-readable media may include, but are not limited to, magnetic storage devices (e.g., hard disks, floppy disks, tapes...), optical disks (e.g., compact disks CDs, digital versatile disks DVDs...), smart cards, and flash memory devices (e.g., cards, sticks, key drives...).

[0092] A computer-readable medium may include a propagated data signal embodying computer program code, for example, in baseband or as part of a carrier wave. The propagated signal may be in a variety of forms, including electromagnetic, optical, etc., or a suitable combination thereof. A computer-readable medium may be any computer-readable medium other than a computer-readable storage medium that can be connected to an instruction execution system, apparatus, or device to communicate, propagate, or transmit the program for use. The program code on the computer-readable medium may be transmitted via any suitable medium, including radio, cable, fiber optic cable, radio frequency signal, or similar medium, or any combination of the above.

[0093] Similarly, it should be noted that, in order to simplify the presentation of this disclosure and thereby facilitate understanding of one or more embodiments of the invention, the foregoing descriptions of the embodiments of this application sometimes combine multiple features into a single embodiment, figure, or description thereof. However, this disclosure method does not imply that the subject matter of this application requires more features than those mentioned. In fact, an embodiment may have fewer features than all of the features of a single embodiment disclosed above.

[0094] As used herein, unless the context clearly indicates otherwise, the terms "a," "an," "an," and / or "the" are not intended to refer to the singular but may include the plural. Generally speaking, the terms "include" and "comprise" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include other steps or elements.

[0095] Unless otherwise specified, the relative arrangement of the parts and steps, numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present application. Meanwhile, it should be understood that, for ease of description, the sizes of the various parts shown in the accompanying drawings are not drawn according to actual proportional relationships. Technology, methods and equipment known to those of ordinary skill in the relevant art may not be discussed in detail, but in appropriate cases, the technology, methods and equipment should be considered as a part of the specification. In all examples shown and discussed here, any specific value should be interpreted as being merely exemplary, rather than as a limitation. Therefore, other examples of exemplary embodiments can have different values. It should be noted that similar numbers and letters represent similar items in the following drawings, and therefore, once an item is defined in an accompanying drawing, it does not need to be further discussed in subsequent drawings.

[0096] In some embodiments, numbers are used to describe the quantity of components and attributes. It should be understood that such numbers used in the description of the embodiments are modified by the modifiers "about", "approximately" or "substantially" in some examples. Unless otherwise stated, "about", "approximately" or "substantially" indicate that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification are approximate values, which may change according to the required characteristics of individual embodiments. In some embodiments, the numerical parameters should take into account the specified significant digits and adopt the general method of retaining digits. Although the numerical domains and parameters used to confirm the breadth of their range in some embodiments of the present application are approximate values, in specific embodiments, the settings of such numerical values ​​are as accurate as possible within the feasible range.

[0097] Although the present application has been described with reference to the current specific embodiments, ordinary technicians in this technical field should recognize that the above embodiments are only used to illustrate the present application, and various equivalent changes or substitutions can be made without departing from the spirit of the present application. Therefore, as long as the changes and modifications to the above embodiments are within the scope of the essential spirit of the present application, they will fall within the scope of the present application.

Claims

1. An accelerated method for linear matrix solution based on CFD, characterized in that: include: Step 1: The CPU reads the grid information and generates an LDU format matrix; Step 2: Convert the LDU format matrix into an ELLB format matrix; Step 3: The ELLB format matrix and the known vector are transferred to the memory of the GPU, and the GPU calls the PCG solver to iteratively calculate the solution vector according to the ELLB format matrix and the known vector; Step 4: Convert the solution vector back into a recognizable LDU format solution vector, and the GPU returns the LDU format solution vector to the CPU.

2. The CFD-based linear matrix solving acceleration method according to claim 1, wherein: Converting the LDU format matrix into an ELLB format matrix includes: Step 2.1: Traverse each row of the LDU format matrix and record the value of the non-zero element and the corresponding column index; Step 2.2: Sort the non-zero elements according to the column index and store them in an ELLB format structure; Step 2.3: Generate two arrays: one to store the values ​​of non-zero elements and the other to store the corresponding column indices; Step 2.4: Generate a bit mask array, where each bit in the bit mask array corresponds to whether an element is valid.

3. The CFD-based linear matrix solving acceleration method according to claim 1, wherein: The GPU calls the PCG solver to iteratively calculate the solution vector according to the ELLB format matrix and the known vector, including: Step 3.1: Setting the initial parameters of the PCG solver, wherein the initial parameters include a solution vector, a residual vector, a preprocessing vector, and a search direction vector; Step 3.2: Perform matrix-vector multiplication calculation using the CUDA kernel function written on the GPU, and update the solution vector and the residual vector according to the calculation results; Step 3.3: Determine whether the convergence condition is met based on the residual vector. If the condition is met, use the updated solution vector as the final solution vector. If the condition is not met, update the search direction vector and then enter step 3.2 for iterative calculation until the convergence condition is met.

4. The CFD-based linear matrix solving acceleration method according to claim 3, wherein: Also includes: Before step 3.2, convert the ELLB format matrix into a CSR format matrix; build an AMG preprocessor on the GPU based on the converted CSR format matrix; and replace the original preprocessing steps in the PCG solver with the AMG preprocessing steps implemented by the AMG preprocessor.

5. The CFD-based linear matrix solving acceleration method according to claim 1, wherein: Also includes: Multiple GPUs are used to calculate the solution vector. Before the ELLB format matrix and the known vector are transferred to the memory of the GPU, an initial grid is generated and the grid is divided to ensure that each subgrid is evenly distributed on each GPU.

6. The CFD-based linear matrix solving acceleration method according to claim 5, wherein: It also includes: using the LDU format solution vector to perform CFD simulation, and adjusting the grid division and the parameters of the PCG solver according to the simulation results.

7. The CFD-based linear matrix solving acceleration method according to claim 1, wherein: Converting the solution vector back to a recognizable LDU format solution vector includes: The solution vector is traversed, and the values ​​of the solution vector are filled into corresponding positions of the LDU format matrix.

8. The CFD-based linear matrix solving acceleration method according to claim 1, wherein: The CUDA API is used to transfer data between the CPU and GPU.

9. An acceleration device for linear matrix solving based on CFD, characterized in that: include: A CPU and multiple GPUs, each GPU including a memory module, a format conversion module, and a computing module, wherein the computing module has a PCG solver; Among them, the CPU is used to read the grid information and generate an LDU format matrix, and receive the LDU format solution vector calculated by the GPU; the format conversion module is used to convert the LDU format matrix into an ELLB format matrix, transfer the ELLB format matrix and the known vector into the memory module of the GPU, and convert the solution vector calculated by the calculation module back to a recognizable LDU format solution vector; the calculation module is used to call the PCG solver to iteratively calculate the solution vector based on the ELLB format matrix and the known vector.

10. A computer-readable medium storing computer program codes, wherein when the computer program codes are executed by a processor, the computer-readable medium implements the CFD-based linear matrix solving acceleration method according to any one of claims 1 to 8.