Hy-GA Parallel Solving Method for Solving Seepage Equation Based on Domestic CPU + DCU Architecture

Through the Hy-GA parallel solution method based on the domestic CPU+DCU architecture, combined with algebraic multi-grid and geometric multi-grid, mixed multi-grid pre-processing is designed, which solves the problems of low parallel solution efficiency and poor convergence effect of seepage mechanical equations in numerical simulation of oil and gas reservoirs, and achieves more efficient calculations and more accurate simulation results.

CN119416693BActive Publication Date: 2025-07-08UNIV OF SCI & TECH BEIJING
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411458842.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-18
Publication Date
2025-07-08
Estimated Expiration
2044-10-18

AI Technical Summary

Technical Problem

The prior art has low parallel solution efficiency of seepage mechanical equations in numerical simulation of oil and gas reservoirs, especially in complex grid structures and heterogeneous media, which is difficult to meet the needs of efficient calculations.

Method used

The Hy-GA parallel solution method based on the domestic CPU+DCU architecture is adopted, and combined with the characteristics of algebraic multi-grid and geometric multi-grid, a hybrid multi-grid pre-processing method is designed. Using the computing advantages of domestic CPU+DCU, heterogeneous parallel computing is performed through the MPI/OpenMP model, and tasks are assigned to the CPU and DCU to execute.

Benefits of technology

It improves the parallel solution efficiency of seepage mechanical equations, can handle more complex physical models and finer grid division, improves calculation accuracy and stability, and solves the problem of insufficient traditional computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119416693B_ABST
    Figure CN119416693B_ABST
Patent Text Reader

Abstract

The present invention discloses a Hy‑GA parallel solution method for solving the seepage equation based on the domestic CPU+DCU architecture, which belongs to the field of high-performance computing technology; based on the computing advantages of the domestic CPU+DCU architecture, the present invention combines the characteristics of algebraic multigrids and geometric multigrids to design an adaptive hybrid multigrid to solve the seepage mechanics equation; the present invention uses the domestic supercomputing architecture for large-scale parallel computing, which can decompose the simulation task for parallel processing, allowing the use of finer grid divisions and more complex physical models, which not only improves the accuracy of the simulation results, but also solves the computing bottleneck problem caused by insufficient traditional computing resources. In addition, the hybrid multigrid preconditioning used in the present invention uses the geometric multigrid aggregation coarsening method in constructing the coarse grid level, and the construction coefficient matrix is ​​based on the Galerkin operator of algebraic information; while maintaining the physical properties as much as possible, the complexity of the construction coefficient matrix is ​​reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of high-performance computing technology, and particularly relates to a Hy-GA parallel solution method for solving the seepage equation based on a domestic CPU+DCU architecture. Background Technique

[0002] The seepage mechanics equation has wide applications in engineering fields such as numerical simulation of oil and gas reservoirs. Numerical simulation of oil and gas reservoirs refers to using a computer to establish a mathematical model of the oil and gas reservoir simulation. Under the control of boundary conditions, the partial differential equation system is converted into discrete grid points and solved by an iterative method. As Figure 1 shown, its main components include a complete oil reservoir numerical model, which mainly includes a partial differential equation system of seepage mechanics and definite solution conditions (boundary conditions, initial conditions). These conditions play a restrictive role in the subsequent discretization into grid points, and finally give the oil-water distribution at a certain moment to predict the dynamic state of the oil reservoir.

[0003] Solving the seepage mechanics equation system is the core of numerical simulation. Some of the equations included are as follows: Component mass conservation equation: Seepage motion equation: Energy conservation equation: State equation: (1) Component ratio relationship: (2) Saturation relationship: (3) Surface tension P c (S j ) = P1 - P j = f(S j ), (4) Multiphase equilibrium equation: f ij = f i1 , i = 1:n c , j = 1:n p and well pressure control equation: and so on.

[0004] Based on the above partial differential equation system, after discrete processing, it is solved by a three-step iterative method, and the process is as Figure 2As shown in the figure, it mainly includes the following steps: (1) Time step selection: Start a new time step, i = 0; (2) Newton iteration: Construct the residual and the Jacobi matrix. Check for convergence. If it converges, proceed to the next time step; if it does not converge, continue with the maximum Newton iteration; (3) Linear solution: Solve the linear equation. Check for linear convergence. If it converges linearly, continue with the non-linear iteration; if it does not converge, perform a step size redo; (4) Step size redo: If the maximum Newton iteration does not converge, redo the time step size. Among them, the linear solver is the main bottleneck in parallel computing, accounting for more than 80% of the workload in the entire solution process of the seepage mechanics equations. In reservoir numerical simulation, the massive and fine grid simulation of the reservoir results in a large coefficient matrix size, poor matrix properties, poor convergence, and irregular sparsity. The above problems affect the improvement of the parallel solution efficiency of the coupled non-linear equations. Therefore, we need a numerical solution method with strong convergence and good stability.

[0005] For large-scale numerical simulations, the multigrid (MG) algorithm has been proven to be one of the most efficient preconditioning acceleration algorithms for solving linear equations. The basic idea of the multigrid (MG) is to restrict the original linear system to a coarse grid, determine an approximate solution on the coarse grid, and then interpolate it back to the fine grid to correct the initial solution.

[0006] The multigrid preconditioning acceleration algorithm is divided into two types: one is the geometric multigrid (GMG), and the other is the algebraic multigrid (AMG). The GMG algorithm usually requires discretization and physical-level grid division based on known geometric shapes, which is a true grid division; the AMG does not require the physical and geometric background of the problem being solved and constructs a virtual multigrid only using the algebraic information in the coefficient matrix of the equations. The difference between GMG and AMG lies in the construction of the coarse grid hierarchy and the coarse grid coefficient matrix A C , and the specific process and differences are as Figure 3 shown. Starting from the large-scale sparse linear equations obtained after discretization, it mainly includes the following content:

[0007] Initialization processing: Perform grid coarsening, restriction operator, and interpolation operator processing.

[0008] Pre-smoothing processing: Perform smoothing processing to reduce high-frequency errors.

[0009] Residual restriction: Restrict the residual on the fine grid to the coarse grid and construct a correction equation.

[0010] Coarse grid solution: Solve the correction on the coarse grid and then perform pre-smoothing processing.

[0011] Grid correction: Add the correction to the initial solution and perform grid correction.

[0012] Post - smoothing: Perform smoothing again to reduce low - frequency errors.

[0013] Convergence verification: Converge when the error reaches the given tolerance; otherwise, continue the iteration until the convergence condition is met.

[0014] The corresponding relationship between GMG and AMG is as follows:

[0015]

[0016] In numerical simulation, they each have their advantages and disadvantages in accelerating the linear solution algorithm: Algebraic multigrid starts directly from the coefficient matrix of the equation, abstracting the grid into a matrix form (as Figure 4 shown). The following are the common methods for selecting coarse grid points in AMG:

[0017] Node i is strongly dependent on j, or equivalently, j strongly influences i if and only if: where α ∈ (0, 1); generally, α is taken as 0.25 - 0.5. For the determination of strong connection relationships, different literatures vary in details. For example, the determination method in "Rajesh Gandham a, Kenneth Esler b, Yongpeng Zhang b. A GPU accelerated aggregation algebraic multigrid method. Computers and Mathematics with Applications. 2014" is:

[0018]

[0019] The coarse grid is generated by the connectivity in the matrix nodes to form a strong connection matrix, which is applicable to a wider range of scenarios. However, for more complex grid structures in higher dimensions, the connectivity of grid nodes increases, and the cost of the setup phase surges. For geometric multigrid, it is necessary to generate the coarse grid for each layer through the initial grid according to the problem background and traverse each grid point to calculate the corresponding coefficient matrix, which has the best effect in the known problem geometric background but has a large amount of calculation and more complex operations.

[0020] Therefore, if traditional GMG or AMG is used alone for pre - processing, when the grid structure is complex or irregular, the efficiency of GMG may decline; when dealing with highly heterogeneous or highly anisotropic media, the convergence speed of AMG is slow. The hybrid pre - processing method of GMG and AMG (Hy - GA) can combine the advantages of both to obtain better results.

[0021] In the field of efficient solution of seepage mechanics equations, it is often necessary to combine the geometric and physical backgrounds of the problems while improving the parallel solution efficiency. In view of this, the present invention combines the characteristics of both and proposes a Hy-GA parallel solution method for solving seepage equations based on the domestic CPU+DCU architecture. At the same time, taking advantage of the computing advantages of the domestic CPU+DCU architecture, a hybrid multiple grid heterogeneous parallel scheme is designed. Summary of the Invention

[0022] The object of the present invention is to propose a Hy-GA parallel solution method for solving seepage equations based on the domestic CPU+DCU architecture to solve the problems of poor convergence effect in reservoir numerical simulation and low parallel solution efficiency of seepage mechanics equations.

[0023] To achieve the above object, the present invention adopts the following technical solutions:

[0024] The Hy-GA parallel solution method for solving seepage equations based on the domestic CPU+DCU architecture, taking advantage of the computing advantages of the domestic CPU+DCU architecture and combining the characteristics of algebraic multigrid and geometric multigrid, designs a conforming hybrid multigrid to solve the seepage mechanics equations, which specifically includes the following contents:

[0025] S1. Initialization stage: Read the coefficient matrix and vector data, and initialize the algorithm parameters and grid hierarchy on the CPU side;

[0026] S2. Setting stage: Construct a multigrid hierarchy, and move the data stored in the memory to the video memory of the DCU through the memory copy function Memcpy between the CPU and the DCU; in this stage, all calculation operations are completed in the DCU;

[0027] S3. Solving stage: Use the multigrid hierarchy constructed in S2 to perform iterative solution between multigrids; in the solving stage, repeat the operations of pre-smoothing, residual restriction, solving on the coarsest grid layer, interpolation correction, and post-smoothing;

[0028] S4. Output stage: When the number of iterations in the solving stage is not less than the preset maximum number of iterations or the current residual is less than the preset tolerance, the solving stage ends, and the iterative information and execution time of each layer are output.

[0029] Preferably, the initialization operation in S1 includes the following contents: Initialize the MPI environment and the DCU environment, and determine the work area and the number of threads of each process, specifically: divide the fine grid into blocks, and allocate each block to an MPI process; within each MPI process, use OpenMP and the DCU to divide each block of the grid to multiple threads and the DCU for processing.

[0030] Preferably, S2 specifically includes the following contents:

[0031] S2.1. Construct a grid hierarchy: Utilize the geometric properties of geometric multigrid and construct a coarse grid hierarchy using the smooth aggregation method, denoted as There are a total of n layers of grids, where Ω 1 The first layer is the initial fine grid;

[0032] S2.2. Construct transfer operators: Construct an interpolation operator using the geometric relationship between the coarse and fine grids, denoted as Construct a restriction operator using the algebraic properties between the coarse and fine grids, denoted as

[0033] S2.3. Construct the coefficient matrix: Utilize the algebraic properties of algebraic multigrid and construct the coefficient matrix based on the Galerkin operator, denoted as

[0034] Preferably, the specific content of S3 is as follows:

[0035] S3.1. Obtain the initial solution: Construct the initial equation AX = B, perform iteration on the initial equation to obtain the initial solution X0, and at the same time calculate the resulting residual r0 = B - AX0, and transfer it to Ω 2 in the coarse level through the restriction operator;

[0036] S3.2. Construct the coarse grid equation: The coefficient matrix on the coarse grid is denoted as Derive the residual equation according to this formula, and the specific formula is expressed as:

[0037]

[0038] After obtaining the residual, transfer it to the next layer of grid through the restriction operator to get A k+1 e k+1 = r k+1 , until the n coarsest grid layer;

[0039] S3.3. Coarse grid correction: After reaching Ω n , directly solve to obtain e n , and interpolate back to the fine grid through the interpolation operator to correct the initial solution. The specific formula is expressed as:

[0040]

[0041] Preferably, non-blocking communication is used as the communication mode between the multiple grids of the constructed multiple grid hierarchy to reduce communication overhead.

[0042] Preferably, the front smoothing and rear smoothing operations specifically refer to: applying a smoothing operator at each grid level; distributing part of the smoothing calculation tasks to the DCU and using OpenMP to parallelize the remaining part to ensure synchronous communication for updating boundary data.

[0043] Preferably, the residual restriction specifically refers to transferring the residuals of the fine grid to the coarse grid through a restriction operator; the interpolation correction specifically refers to transferring the solution of the coarse grid back to the fine grid through an interpolation operator.

[0044] Preferably, in the solution phase, the global error is calculated through a global reduction operation, and it is determined whether to continue the iteration based on the calculated error.

[0045] Compared with the prior art, the present invention provides a Hy-GA parallel solution method for solving the seepage equation based on a domestic CPU+DCU architecture, having the following beneficial effects:

[0046] (1) The present invention proposes a Hy-GA parallel solution method for solving the seepage equation based on a domestic CPU+DCU architecture. Based on the computing advantages of the domestic CPU+DCU architecture and combining the characteristics of algebraic multigrid and geometric multigrid, a well-posed hybrid multigrid is designed to solve the seepage mechanics equation; the present invention uses the domestic supercomputer architecture for large-scale parallel computing, can decompose and parallelize simulation tasks, allows for the use of finer grid division and more complex physical models, not only improves the accuracy of simulation results, but also solves the computational bottleneck problem caused by insufficient traditional computing resources.

[0047] (2) The hybrid multigrid preconditioner used in the present invention uses the geometric multigrid aggregation coarsening method in constructing the coarse grid level, while the coefficient matrix is constructed based on the Galerkin operator of algebraic information; while maintaining physical properties as much as possible, the complexity of constructing the coefficient matrix is reduced.

[0048] In summary, the present invention combines the advantages of algebraic multigrid (AMG) and geometric multigrid (GMG), not only maintains a certain degree of self-adaptability, but also reduces the cost consumption of constructing operators. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 It is a schematic diagram of the mathematical model for numerical simulation of oil and gas reservoirs mentioned in the background art of the present invention;

[0050] Figure 2 It is a flowchart of the time iteration algorithm mentioned in the background art of the present invention;

[0051] Figure 3 It is a diagram of the entire calculation process (CPU serial) of multigrid mentioned in the background art of the present invention;

[0052] Figure 4 Schematic diagram of the basic idea principle of algebraic multigrid (AMG) mentioned in the background art of the present invention;

[0053] Figure 5 Schematic diagram of the multigrid V-cycle iteration strategy mentioned in the background art of the present invention;

[0054] Figure 6 Schematic diagram of the MPI / OpenMP parallel model mentioned in the embodiments of the present invention;

[0055] Figure 7 Overall flowchart of the Hy-GA parallel solution method for solving the seepage equation based on the domestic CPU+DCU architecture mentioned in the embodiments of the present invention;

[0056] Figure 8 Schematic diagram of the supercomputer structure mentioned in the embodiments of the present invention;

[0057] Figure 9 Schematic diagram of the domestic DCU+CPU hybrid architecture mentioned in the embodiments of the present invention;

[0058] Figure 10 Flowchart of creating an aggregation mapping mentioned in the embodiments of the present invention. Detailed implementation manners

[0059] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.

[0060] Embodiment 1:

[0061] The present invention proposes a Hy-GA parallel solution method for solving the seepage equation based on the domestic CPU+DCU architecture. Compared with the existing design, the present invention has the following advantages:

[0062] Complementary advantages of hybrid preprocessing:

[0063] (1) Initial approximation: Geometric multigrid (GMG) can provide a better initial approximation, which helps to accelerate the convergence of algebraic multigrid (AMG).

[0064] (2) Coarsening process: The coarsening processes of geometric multigrid (GMG) and algebraic multigrid (AMG) can be used complementarily. For example, geometric multigrid (GMG) can be used for coarsening several times first, and then algebraic multigrid (AMG) can be used for processing on a coarser grid, which can better handle complex geometric structures and heterogeneities.

[0065] (3) Smoothing: Geometric multigrid (GMG) usually performs smoothing on a finer grid, while algebraic multigrid (AMG) can perform efficient smoothing on a coarser grid. Their combined use can improve the overall smoothing effect.

[0066] (4) Adapting to complex physical phenomena: Numerical simulation of oil reservoirs often involves complex phenomena such as multiphase flow and multicomponent transport. Hybrid preprocessing methods can better adapt to these complex physical phenomena and improve the stability and accuracy of simulation.

[0067] At the same time, on the basis of keeping the basic principles of reservoir numerical simulation unchanged, a heterogeneous parallel multi-grid preprocessing method based on the MPI / OpenMP model is used to give full play to the following advantages of heterogeneous parallel computing:

[0068] (1) Improved computing power: DCU has significant advantages in processing large-scale parallel computing tasks, especially in matrix operations, vector operations, etc., and can provide much higher computing power than CPU. Distributing the preprocessing of geometric multigrid (GMG) and algebraic multigrid (AMG) to DCU can significantly improve the computing efficiency of the solution stage.

[0069] (2) Maximizing resource utilization: By taking advantage of the different strengths of the CPU and DCU, hardware resources can be maximized. The CPU is good at processing complex logic control and serial calculations, while the DCU is good at large-scale data parallel calculations. Assigning appropriate tasks to the corresponding hardware can improve overall efficiency.

[0070] (3) Task parallelization: Many steps of geometric multigrid (GMG) and algebraic multigrid (AMG) (such as grid generation, coarsening, smoothing, interpolation, etc.) can be parallelized. Heterogeneous parallel computing can distribute these tasks to multiple computing units for parallel execution, thereby significantly reducing the computing time and making the Jacobian coefficient matrix solution more efficient, thereby improving the efficiency of reservoir numerical simulation.

[0071] In order to more clearly describe the specific steps of the Hy-GA parallel solution method for solving the percolation equation based on the domestic CPU+DCU architecture proposed in the present invention, some terms are explained as follows:

[0072] (1) Smoothing operator: also called preconditioner, used to eliminate high-frequency error components on each layer of coarse grid. Commonly used smoothing operators are Gaussian and Jacobi relaxation iterative methods.

[0073] (2) Transfer operator: including interpolation operator and restriction operator. The interpolation matrix is ​​used to transfer the solution or residual information from the coarse grid to the fine grid; the restriction matrix is ​​used to transfer the solution or residual information of the fine grid to the coarse grid.

[0074] (3) Mesh coarsening: Generate a coarser mesh hierarchy with fewer grid points and larger intervals from the initial original fine mesh through certain geometric relationships or strong connectivity.

[0075] (4) V-cycle: An iterative loop strategy for multigrid. As Figure 5 shown, the present invention adopts a two-level coarse mesh structure. In the initialization stage, two-level coarse meshes are constructed, the V-cycle strategy is called, and the coarse mesh equations are solved recursively once until the coarsest mesh level.

[0076] (5) CPU-DCU architecture: As Figure 8 、 Figure 9 shown, it is a domestic CPU+DCU architecture; both the DCU and the CPU have their own memories (of course, there are also some heterogeneous architectures similar to the DCU, and their storage architecture sharing the memory space with the CPU, such as integrated graphics cards. However, current supercomputer architectures, such as Summit, Frontier, and domestic DCU series supercomputers, are basically designed with separate memories for GPUs / DCUs and CPUs). Based on this storage and supercomputer hardware architecture, which is different from Shenwei, the general idea for heterogeneous accelerated parallelism is:

[0077] 1) Copy data from the CPU memory to the DCU memory (referred to as Host to Device or H2D);

[0078] 2) Start the kernel (kernel function) to perform calculations on the DCU;

[0079] 3) Copy the data back from the DCU memory to the CPU-end memory (Device to Host or D2H);

[0080] 4) Perform necessary calculations or communications on the CPU side.

[0081] There will surely be overhead when the CPU and the DCU communicate. At the same time, optimizing the calculations for compute-intensive tasks is also the key to improving the DCU's computing performance. The present invention uses the following methods for improvement:

[0082] 1) Appropriately split the data;

[0083] 2) Use different streams for transmission and calculation;

[0084] 3) Branch optimization: The DCU has relatively weak control ability for logic, so branches should be minimized;

[0085] 4) For matrix calculations, the parts that take more time are vector addition and subtraction operations, scalar-vector multiplication operations, and matrix-vector multiplication operations; the parallel scheme studied adopts the LightSpMV method.

[0086] (6) MPI / OpenMP parallel model: As Figure 6 shown, MPI provides a mechanism for inter - process communication and is suitable for parallel programs that need to transfer data between multiple computing nodes; while OpenMP is a parallel programming standard for shared - memory systems and provides a communication mechanism between threads; MPI extends from single - core to multiple nodes, with each region placed on a computing node: For the initial grid, regional division can be carried out, and each process processes a grid region to complete a solution process; in this process, the smoothing process of each layer of coarse grid is completed in parallel by OpenMP threads, and finally the results are summarized into one process.

[0087] Please refer to Figure 7 , for the specific process of the Hy - GA parallel solution method for solving the seepage equation based on the domestic CPU + DCU architecture proposed by the present invention, the specific content is as follows:

[0088] Setup stage:

[0089] (1) Construct the grid hierarchy: Utilize the geometric characteristics of geometric multigrid (GMG) and use the smooth aggregation method to construct the coarse - grid hierarchy. There are a total of n layers of grids, where Ω 1 is the initial fine grid of the first layer.

[0090] (2) Transfer operator construction: Construct the interpolation operator using the geometric relationship between the coarse and fine grids Construct the restriction operator using algebraic properties

[0091] (3) Coefficient matrix construction: Utilize the algebraic properties of AMG and construct the coefficient matrix based on the Galerkin operator

[0092] Solve:

[0093] (1) Obtain the initial solution: Perform several simple iterations (smoothing process) on the initial equation AX = B to obtain the initial solution X0, and at the same time calculate the resulting residual r0 = B - AX0, which is transferred to the Ω 2 coarse level by the restriction operator.

[0094] (2) Construct the coarse - grid equation: The coefficient matrix on the coarse grid The residual equation can be derived from the formula After obtaining the residual, it is then transferred to the next - layer grid through the restriction operator to obtain A k+1 e k+1 = r k+1 , until the Ω n coarsest grid layer.

[0095] (3) Coarse - grid correction: Reach Ωn , directly solve to obtain e n , through the interpolation operator, interpolate back to the fine grid to correct the initial solution,

[0096] Heterogeneous parallel design scheme of MG:

[0097] Mesh partitioning and allocation: Partition the fine grid into blocks, and allocate each block to an MPI process. Inside each MPI process, use OpenMP and DCU to further partition each block of the grid to multiple threads and DCUs for processing. As Figure 6 shown.

[0098] Communication mode: Use non-blocking communication (such as MPI_Isend and MPI_Irecv) to reduce communication overhead. Ensure the correct transmission and update of boundary data.

[0099] Details of parallel heterogeneous implementation:

[0100] Initialization:

[0101] Initialize the MPI environment (MPI_Init).

[0102] Initialize the DCU environment (such as OpenCL, CUDA).

[0103] Determine the working area and number of threads for each process.

[0104] Smoothing step:

[0105] Apply the smoothing operator at each grid level. Allocate part of the smoothing calculation tasks to the DCU, and use OpenMP to parallelize the remaining part. Ensure synchronous communication (using MPI) to update the boundary data.

[0106] Restriction and interpolation:

[0107] Through the restriction operator, transfer the residuals of the fine grid to the coarse grid. It can be accelerated and implemented on the DCU.

[0108] Through the interpolation operator, transfer the solution of the coarse grid back to the fine grid. It can be accelerated and implemented on the DCU.

[0109] Convergence judgment:

[0110] Calculate the global error through a global reduction operation (such as MPI_Rce). Determine whether to continue the iteration based on the error.

[0111] In traditional reservoir numerical simulation, the seepage equation often involves complex physical phenomena such as multiphase flow, heat conduction, and geomechanics. A large amount of geological data, logging data, and production data need to be processed, and it is difficult for traditional computing platforms to efficiently handle these cases. With the domestic supercomputers entering the E-level era, the powerful computing capabilities of domestic DCUs can significantly shorten the simulation time, improve the computing efficiency, and have strong data processing and storage capabilities, enabling rapid storage, retrieval, and processing of big data. At the same time, the domestic supercomputer architecture supports large-scale parallel computing, which can decompose simulation tasks for parallel processing, allowing the use of finer grid division and more complex physical models, improving both the accuracy of simulation results and solving the computing bottleneck problem caused by insufficient traditional computing resources.

[0112] In general numerical simulations, such as CFD, electromagnetic field simulation, and structural mechanics, the multigrid preconditioning technique is widely used. However, most of them use algebraic multigrid (AMG). The reason is that it does not need to consider the complex geometric background of the problem and can quickly solve the approximate solution of the problem, with a wider range of applicability. But this method may ignore the impact of changes in geometric information stored on the grid on the accuracy of the solution when constructing the grid hierarchy according to the problem background. Therefore, in the present invention, the hybrid multigrid preconditioner used constructs the coarse grid hierarchy using the geometric multigrid aggregation coarsening method and constructs the coefficient matrix based on the Galerkin operator of algebraic information, reducing the complexity of constructing the coefficient matrix while maintaining physical characteristics as much as possible.

[0113] GMG AMG HyGA (Hybrid Multigrid Preconditioning) Memory requirement Low High Low Operator complexity Less costly More costly Less costly Matrix quality High Less control High at finer levels User-friendliness Less friendly More friendly Intermediate Irregular domains Not robust Robust Robust

[0114] As shown above, it can be seen that the well-posed hybrid multigrid (Hy-GA) proposed by the present invention combines the advantages of GMG and AMG, not only maintaining a certain degree of self-adaptability but also reducing the cost of constructing the operator.

[0115] Example 2:

[0116] Based on Example 1 but with differences, please refer to Figure 7 - 8 , the execution process of adopting the well-posed hybrid multigrid (Hy-GA) proposed by the present invention specifically includes the following contents:

[0117] Header files and libraries:

[0118] #include<mpi.h>

[0119] #include<omp.h>

[0120] #include <vector>

[0121] #include <iostream>

[0122] #include "dcu_utils.h" / / DCU utility library, providing DCU initialization, data transmission, and calculation functions

[0123] Define the MG structure:

[0124] Contains the data stored in each layer of the grid (coefficient matrix, right - hand side b, etc.), AMG hierarchy and GMG hierarchy, as well as the processing methods of smoothing operators, interpolation methods, and methods for calculating residuals.

[0125]

[0126]

[0127]

[0128]

[0129] Define the method for constructing the coarse - grid hierarchy in an aggregated manner:

[0130] (Assume criterion = smoothed aggregation coarsening algorithm)

[0131] createCoarseLevelSystem(const O& fineOperator) / / Pass in the initial fine grid and aggregation strategy criterion

[0132] {

[0133] P = criterion_.getProlongationDampingFactor(); / / P is the interpolation matrix

[0134] / *

[0135] ① Use the matrix of the fine operator to create a MatrixGraph, and then construct a PropertiesGraph based on this graph. These graphs are used to store the topological structure of the matrix, including the properties of vertices and edges.

[0136] MG::MatrixGraph<const typename O::matrix_type> MatrixGraph;

[0137] MG::PropertiesGraph<MatrixGraph, Dune::Amg::VertexProperties, Dune::Amg::EdgeProperties, Dune::IdentityMap, Dune::IdentityMap> PropertiesGraph

[0138] / *

[0139] ② Creation of the aggregation mapping: The aggregation mapping (aggregates Map) is constructed through the property graph and the decision criterion. The construction process returns four different types of aggregate quantities: ordinary aggregates, isolated aggregates, single aggregates, and skipped aggregates. To optimize storage and calculation, a renumberer is used to renumber the aggregates.

[0140] MG::tie(noAggregates, isoAggregates, oneAggregates, skippedAggregates) = aggregatesMap_->buildAggregates(fineOperator.getmat(), PropertiesGraph; MatrixGraph, criterion_, true);

[0141] ③ Calculation of the coefficient matrix of the Galerkin operator: (A″ = p r AP)

[0142] GerlerkinProductBuilder.calculate(P, fineOperator.getmat(), *aggregatesMap_, *matrix)

[0143] }

[0144] As described above, it is only a preferred specific implementation manner of the present invention. However, the protection scope of the present invention is not limited thereto. For those of ordinary skill in the art in this technical field, without departing from the principle described in the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.< / iostream> < / vector>

Claims

1. A Hy-GA parallel solution method for solving the seepage equation based on a domestic CPU+DCU architecture, characterized in that, Based on the computing advantages of the domestic CPU+DCU architecture, combined with the characteristics of algebraic multigrid and geometric multigrid, a conforming hybrid multigrid is designed to solve the seepage mechanics equation, which specifically includes the following contents: S1. Initialization stage: Read the coefficient matrix and vector data, and initialize the algorithm parameters and grid hierarchy on the CPU side; S2. Setup stage: Construct the multigrid hierarchy, and move the data stored in memory to the video memory of the DCU through the memory copy function Memcpy between the CPU and the DCU; in this stage, all computing operations are completed in the DCU; specifically, it includes the following contents: S2.

1. Construct a grid hierarchy: Utilize the geometric properties of geometric multigrid and construct a coarse grid hierarchy using the smooth aggregation method, denoted as There are a total of n grids, where Ω 1 The first layer is the initial fine grid; S2.

2. Construct the transfer operator: construct the interpolation operator by using the geometric relationship between the coarse and fine grids, denoted as Construct the restriction operator by using the algebraic properties between the coarse and fine grids, denoted as S2.

3. Construct the coefficient matrix: Utilize the algebraic properties of algebraic multigrid to construct the coefficient matrix based on the Galerkin operator, denoted as S3. Solving stage: Use the multigrid hierarchy constructed in S2 to perform iterative solution between multigrids; in the solving stage, pre-smoothing, residual restriction, coarsest grid layer solution, interpolation correction, and post-smoothing operations are repeated; specifically, it includes the following contents: S3.

1. Obtain the initial solution: Construct the initial equation AX = B, iterate the initial equation to obtain the initial solution X0, and at the same time calculate the resulting residual r0 = B - AX0, which is transferred to Ω through the restriction operator 2 in the coarse level; S3.

2. Construct the coarse grid equation: The coefficient matrix on the coarse grid is denoted as Based on this formula, the residual equation is derived, and the specific formula is expressed as: After obtaining the residual, it is then transferred to the next level of the grid through a restriction operator to obtain A k+1 e k+1 = r k+1 , until Ω n the coarsest grid level; S3.

3. Coarse grid correction: reaching Ω n After that, directly solve to obtain e n , and through the interpolation operator, interpolate back to the fine grid to correct the initial solution. The specific formula is expressed as: S4. Output stage: When the number of iterations in the solving stage is not less than the preset maximum number of iterations or the current residual is less than the preset tolerance, the solving stage ends, and the iteration information and execution time of each layer are output.

2. The Hy-GA parallel solution method for solving the seepage equation based on the domestic CPU+DCU architecture according to claim 1, wherein, The initialization operation described in S1 includes the following contents: Initialize the MPI environment and the DCU environment, and determine the working area and number of threads of each process, specifically: divide the fine grid into blocks, and each block is assigned to an MPI process; within each MPI process, each block of the grid is further divided among multiple threads and DCUs for processing using OpenMP and the DCU.

3. The Hy-GA parallel solution method for solving the seepage equation based on the domestic CPU+DCU architecture according to claim 1, characterized in that, For non-blocking communication is used as the communication mode between multigrids of the constructed multigrid hierarchy to reduce communication overhead.

4. The Hy-GA parallel solution method for solving the seepage equation based on the domestic CPU + DCU architecture according to claim 1, characterized in that, The pre-smoothing and post-smoothing operations specifically refer to: applying a smoothing operator at each grid level; assigning part of the smoothing calculation tasks to the DCU, and using OpenMP to parallelize the remaining part to ensure synchronous communication to update boundary data.

5. The Hy-GA parallel solution method for solving the seepage equation based on the domestic CPU+DCU architecture according to claim 1, characterized in that, The residual restriction specifically refers to passing the residual of the fine grid to the coarse grid through a restriction operator; the interpolation correction specifically refers to passing the solution of the coarse grid back to the fine grid through an interpolation operator.

6. The Hy-GA parallel solution method for solving the seepage equation based on the domestic CPU+DCU architecture according to claim 1, characterized in that, In the solving stage, the global error is calculated through a global reduction operation, and it is judged whether to continue the iteration according to the calculated error.

Citation Information

Patent Citations

  • SW architecture-oriented fluid dynamics multi-grid solver parallel optimization method

    CN115906684A

  • Thermal fluid high-order DNS large-scale parallel computing method based on domestic supercomputing

    CN118690679A