A computer system for collaborative acceleration of fluid simulation by a DPU and a CPU
By configuring the DPU and CPU on the fluid simulation computing node to cooperate, the fluid simulation algorithm is accelerated, and the problem of memory read and write speed bottleneck in fluid simulation is solved, improving computing efficiency and reducing hardware costs.
Patent Information
- Application Number
- CN202510364416.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-26
AI Technical Summary
The prior art is difficult to effectively solve the bottleneck problem of memory read and write speed in fluid simulation, and there are problems such as high hardware cost and difficult to port algorithm replication.
By configuring the DPU and CPU to cooperate on the computing node, the fluid simulation algorithm is accelerated, and the DPU directly reads memory data for residual calculation and convergence judgment of iteration results, reducing the CPU's waiting time and I/O burden.
It improves the computing efficiency of fluid simulation, reduces the waiting time of the CPU core, significantly reduces the delay of memory reading and writing, and realizes the acceleration of fluid simulation.
Smart Images

Figure CN119883381B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of large-scale parallel solution in computational fluid dynamics, and particularly to a computer system for collaborative acceleration of fluid simulation by a DPU and a CPU. Background Art
[0002] Fluid simulation is a technology for simulating fluid flow by computer, which is widely used in many fields such as aerospace, automotive, and energy. In fluid simulation, the CPU undertakes the core computing tasks, and its multi-core processing ability and multi-thread characteristics are crucial for processing large-scale parallel computing tasks, especially in fluid dynamics problems involving a large number of iterations and complex calculations. However, with the continuous improvement of the complexity and accuracy requirements of the simulation model, the demand for computing resources has increased sharply. As the complexity and accuracy requirements of the simulation model continue to increase, the demand for computing resources has increased sharply. And the improvement of CPU performance is restricted by the lithography process, and the development speed has slowed down, making it difficult for the CPU to meet the high-performance fluid simulation requirements.
[0003] The performance bottlenecks in fluid simulation are mainly reflected in two aspects: one is the floating-point computing performance of the CPU, which is crucial for performing numerical calculations; the other is the memory read / write speed, because fluid simulation needs to frequently read and write data from memory. For example, matrix iterative solution is usually limited by the floating-point performance of the CPU, while the residual calculation of physical fields such as pressure and flow velocity is more limited by the memory read / write speed. Different hardware bottlenecks of different algorithms have given rise to heterogeneous algorithms of different computing units.
[0004] In the prior art, a computer system for accelerating fluid simulation, such as a computer network system and a computer system with the publication number CN118714020B, optimizes the network topology to improve the communication bottleneck between CPUs, thereby improving the parallel computing efficiency, but there is a bottleneck in the memory read / write speed.
[0005] Another example is the publication number CN118245039B, a method for transplanting and optimizing parallel algorithms based on domestic accelerators. By heterogeneous of CPU and GPU, parallel algorithms are developed to accelerate fluid simulation, and the floating-point computing performance of the GPU is utilized to improve the parallel computing speed, and the overall computing performance has been significantly improved. However, there are also some disadvantages, such as high hardware costs, replication of CPU parallel algorithms, and difficulty in transplanting to the GPU.
[0006] The above two different types of prior art do not solve the bottleneck problem of the memory read / write speed in fluid simulation, and the applicable scope is limited. Summary of the Invention
[0007] In view of the deficiencies in the prior art, the purpose of the present invention is to provide a computer system for collaborative acceleration of fluid simulation by a DPU and a CPU, which is used to achieve the effects of increasing the memory access bandwidth, reducing the waiting time of CPU cores, and thus improving the fluid simulation efficiency.
[0008] To achieve the above object, a computer system for collaborative acceleration of fluid simulation by a DPU and a CPU is proposed, which includes one or more than two computing nodes. The computing nodes communicate with each other through a network. Each computing node is configured with the same number of CPUs and DPUs.
[0009] Each computing node has one or more than two CPUs. The multiple CPUs communicate with each other through an internal connection. The CPU is responsible for the iterative solution of the fluid simulation control equation and writes the iterative result into the memory.
[0010] Each computing node has one or more than two DPUs. The DPU adopts a direct memory access controller and directly reads the memory data without passing through the CPU memory controller. It calculates the residual of the iterative result and / or verifies whether it meets the convergence criterion. When the current iterative result meets the convergence criterion, the CPU completes the ongoing iterative solution, and the DPU calculates the residual of this iterative result for post-processing and result output.
[0011] Furthermore, the computer system uses the collaborative acceleration of the DPU and the CPU to solve the fluid simulation algorithm. The specific steps are as follows:
[0012] S1: Allocate the computational domain to each computing node. The computational domain grid and physical quantities are stored in the memory. The CPUs in each computing node set the initial conditions and boundary conditions of the computational domain.
[0013] S2: Each computing node performs parallel fluid simulation calculations according to the task allocation. Specifically,
[0014] The CPU advances the time step based on the initial conditions or the calculation results of the previous time step. In each time step, it iteratively solves the physical quantities and stores them in the memory. The DPU reads the boundary conditions and iterative results in the memory to calculate the residual and determines whether the residual meets the convergence criterion. The CPU reads the convergence judgment result of the DPU. If it does not meet the criterion, the CPU continues the iterative solution. If it meets the criterion, the CPU completes the ongoing iterative solution, and the DPU calculates the residual of this iterative result.
[0015] S3: The DPU performs post-processing and result output, and determines whether this time step is the last time step. If not, it enters the next time step. In each time step, S2 and S3 are repeated until the last time step. The DPU writes the solved physical quantities into the hard disk.
[0016] Further, the specific steps in step S1 include:
[0017] S11: Divide the computational domain into several blocks with relatively balanced numbers of grids according to the grid density and the geometric characteristics of the computational domain;
[0018] S12: Mark the cells at the interfaces of several blocks so that the calculation results of the block nodes are equal for different computing nodes during iterative calculations;
[0019] S13: Configure the parallel computing environment so that each block is independently processed in parallel on different computing nodes, and each computing node can access the physical quantities in the computational domains of other computing nodes through the network;
[0020] S14: Specify initial values for the physical quantities in the computational domain, set boundary conditions, and apply the boundary conditions to the boundary grids of the computational domain.
[0021] Further, based on the initial values or the calculation results of the previous time step, perform time step advancement. The specific steps for parallel fluid simulation calculations based on physical quantity computing nodes in each time step are as follows:
[0022] ①. The CPU performs iterative solutions of the momentum conservation equation and updates the flow velocity Ui. The DPU calculates the residuals of the momentum conservation equation and performs convergence judgment. Repeat this process until the residuals calculated by the DPU in the Nth iteration satisfy the convergence judgment. The CPU performs N + 1 iterative solutions, obtains the iterative results of the N + 1th iteration, and updates the flow velocity Ui for subsequent steps. At the same time, the DPU calculates the residuals of the iterative results of the N + 1th iteration as the final residual record;
[0023] Or if the residuals never satisfy the convergence judgment and reach the maximum number of iterations, obtain the iterative results of the maximum number of iterations and update the flow velocity Ui for subsequent steps. At the same time, the DPU calculates the residuals of the iterative results of the maximum number of iterations as the final residual record;
[0024] ②. The CPU performs iterative solutions of the pressure Poisson equation and updates the pressure p. The DPU calculates the residuals of the pressure Poisson equation and performs convergence judgment. Repeat this process until the residuals calculated by the DPU in the Nth iteration satisfy the convergence judgment. The CPU performs N + 1 iterative solutions, obtains the iterative results of the N + 1th iteration, and updates the pressure p for subsequent steps. At the same time, the DPU calculates the residuals of the iterative results of the N + 1th iteration as the final residual record;
[0025] Or if the residuals never satisfy the convergence judgment and reach the maximum number of iterations, obtain the iterative results of the maximum number of iterations and update the flow velocity Ui for subsequent steps. At the same time, the DPU calculates the residuals of the iterative results of the maximum number of iterations as the final residual record;
[0026] ③. Use the pressure p to correct the flow velocity Ui. The DPU calculates the residual of the mass conservation equation and makes a convergence judgment. If the convergence judgment is not satisfied, return to step ① to continue the iterative calculation until the convergence criterion is met or the maximum number of inner loops between steps ① - ③ is reached;
[0027] ④. Solve for other physical quantities. Use the corresponding equations to solve them in the same way as in step ①. At the same time, solve the unsolved physical quantities through the already solved physical quantities until all physical quantities are solved;
[0028] ⑤. After all physical quantities are solved, the DPU calculates the residual of the continuity equation through the solved physical quantities and determines whether the residual satisfies the convergence criterion of the outer loop between steps ① - ⑤; If not, return to step ① to continue the iterative calculation until the convergence criterion of the outer loop between steps ① - ⑤ is met or the maximum number of outer loops between steps ① - ⑤ is reached, and then enter step S3.
[0029] Further, the specific steps of post - processing and result output include:
[0030] After reaching the convergence criterion, use the DPU to perform auxiliary calculations such as result monitoring and real - time post - processing of the simulation calculation;
[0031] Through the DPU direct memory access technology, directly initiate a data write request to the CPU memory controller to complete the post - processing and write the results to the memory;
[0032] Or use the PCIe passthrough technology to directly or after compression write the calculation results and post - processing results to the hard disk.
[0033] Further, in step ④, other physical quantities include the turbulent viscosity coefficient vt, the turbulent kinetic energy k, and the specific dissipation rate ω. The specific steps for solving them include:
[0034] [1]. The CPU performs iterative solution of the turbulent kinetic energy k equation and updates the turbulent kinetic energy k. The DPU calculates the residual of the momentum conservation equation and makes a convergence judgment. Repeat this process until the N - th residual calculated by the DPU satisfies the convergence judgment. The CPU performs the (N + 1) - th iterative solution, obtains the (N + 1) - th iterative result, and updates the turbulent kinetic energy k for subsequent steps,
[0035] Or the CPU reaches the maximum number of iterations, obtains the iterative result of the maximum number of iterations, and updates the turbulent kinetic energy k for subsequent steps;
[0036] [2]. The CPU calculates the specific dissipation rate ω through the turbulent kinetic energy k, and then calculates the turbulent viscosity coefficient vt.
[0037] The technical effects and advantages of the present invention:
[0038] 1. By offloading the residual calculation task in fluid simulation to the DPU, the present invention effectively reduces the waiting time of the CPU cores, thereby accelerating the fluid simulation calculation.
[0039] 2. The asynchronous parallel method allows the CPU to continue the iteration while the DPU calculates the residual. After the residual meets the convergence condition, the current iteration is continued and the calculation result is retained. According to experience, adding one more iteration can further reduce the residual and improve the accuracy.
[0040] 3. Through the DMA technology, the present invention enables the DPU to directly initiate a data write request to the CPU memory controller, offloading the result output task of the CPU. This process bypasses the intervention of the CPU, allowing the computing resources and storage resources to transfer independently, significantly reducing the I / O burden and data transfer latency of the CPU, thereby accelerating the fluid simulation.
[0041] 4. The DPU is connected to the computer motherboard using a PCIe slot, which is compatible with current common servers, facilitating the upgrade and transformation of existing servers. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is a schematic diagram of the computer system principle;
[0043] Figure 2 is a schematic diagram of the main process of the method for accelerating fluid simulation in the computer system;
[0044] Figure 3 is a schematic diagram of the hardware of the computer system embodiment;
[0045] Figure 4 is a schematic diagram of the PIMPLE algorithm process for accelerating fluid simulation in the computer system;
[0046] Reference numerals: 1, CPU; 2, CPU core; 3, CPU memory controller; 4, PCIe controller; 5, memory; 6, hard disk; 7, DPU; 8, memory direct access controller; 9, network controller; 10, computing node; 11, DAC high-speed cable. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following, in combination with the accompanying drawings and preferred embodiments, details the specific implementation manner, structure, features, and their effects of the present invention as follows.
[0048] Refer to Figure 1As shown in the figure, the present invention provides a computer system for collaborative acceleration of fluid simulation by DPU and CPU, including one or more than two computing nodes, which communicate with each other through a network. Each computing node is configured with the same number of CPUs and DPUs. Each computing node has one or more than two CPUs, and the multiple CPUs communicate with each other through an internal connection. The CPU is responsible for the iterative solution of the fluid simulation control equation and writes the iterative result into the memory. Each computing node has one or more than two DPUs. The DPU all adopts a direct memory access controller, directly reads the memory data without passing through the CPU memory controller, calculates the residual of the iterative result and / or verifies whether the convergence criterion is satisfied. When the current iterative result satisfies the convergence criterion, the CPU completes the ongoing iterative solution, and the DPU calculates the residual of the iterative result and performs post-processing and result output.
[0049] For a specific embodiment, refer to Figure 3 As shown in the figure, the computer system includes 2 ASUS RS720A-E11-RS12 rack-mounted servers, each of which is called 1 computing node. Each computing node contains 2 AMD 7773X CPUs, and the CPU integrates a memory controller and a PCIe controller; each CPU is connected to 1 DPU and 1 network controller through the PCIe controller respectively; the motherboard of each computing node integrates 1 DMA controller and is connected to 1 hard disk; each CPU is connected to 12 memory modules; the network controllers on the two computing nodes are connected to each other in pairs through DAC high-speed cables.
[0050] Table 1 Hardware Selection
[0051]
[0052] Refer to Figure 2 and Figure 4 As shown in the figure, on this computer system, a method for collaborative acceleration of the PIMPLE algorithm by DPU and CPU for fluid simulation is provided. This method is based on the numerical solution of incompressible flow of the PIMPLE algorithm. The numerical solution of incompressible flow is an important topic in fluid mechanics, and its key algorithms include pressure-velocity coupling, such as SIMPLE, PISO, and PIMPLE. Among them, the PIMPLE method combines the iterative correction of SIMPLE and the multi-step pressure correction of PISO, and is applicable to the solution of steady-state and unsteady incompressible flows, with a wide range of adaptability and high solution accuracy. Based on the hardware system of the embodiment, the specific steps are as follows:
[0053] S1: Allocate the computational domain to two computing nodes. The computational domain grid and physical quantities are stored in the memories of the two computing nodes respectively. The CPU sets the initial conditions and boundary conditions of the computational domain. Specifically:
[0054] According to the grid density and the geometric characteristics of the computational domain, the computational domain is divided into 48 blocks. In this embodiment, the number of blocks is equal to the number of memories, and the number of grids in the blocks is relatively balanced, so that the computational load of each computing node is balanced. The units on the block interface are marked to ensure that the calculation results are equal when the block interface is iterated at different computing nodes.
[0055] Configure the parallel computing environment so that each block can be processed independently and in parallel on different computing nodes, and each computing node can access the physical quantities of the computing domain in other computing nodes through the network; specify the initial value for each physical quantity in the computing domain, and set the corresponding boundary conditions. The boundary conditions are assigned to the boundary grid using fixed values or gradient values.
[0056] S2: Two computing nodes perform parallel calculations of fluid simulation according to task allocation. The CPU starts from the initial conditions or the calculation results of the previous time section and iteratively solves the physical quantities in the calculation domain. The physical quantities in this embodiment include flow velocity Ui, pressure p, turbulent viscosity coefficient vt, turbulent kinetic energy k and specific dissipation rate ω. The above physical quantities are coupled. This application uses the PIMPLE algorithm to decouple and predict, solve and correct the physical quantities in turn. After each iteration is completed, the iteration results are written to the memory. The DPU reads the boundary conditions in the memory and the iteration results of the physical quantities in the calculation domain, calculates the residual of the physical quantity, and determines whether the residual meets the convergence criterion. In this embodiment, the convergence criterion is set to 10 -6 When the DPU calculates the residual, the CPU continues to perform iterative solution and writes the iterative result into the memory; the CPU reads the convergence judgment result of the DPU, and if it has not converged, it continues to perform iterative solution, and if it has converged, it goes to step S3.
[0057] Specifically: the time step is advanced based on the initial value or the calculation result of the previous time part, and the fluid simulation parallel calculation is performed based on the physical quantity calculation node in each time step:
[0058] ①. The CPU performs iterative solution of the momentum conservation equation and updates the flow velocity Ui. The DPU calculates the residual of the momentum conservation equation and makes a convergence judgment. This cycle repeats until the Nth residual calculated by the DPU satisfies the convergence judgment. The CPU performs N+1th iterative solution, obtains the N+1th iterative result and updates the flow velocity Ui for subsequent steps. At the same time, the DPU calculates the residual of the N+1th iterative result as the final residual record. For example: the CPU performs the first iterative solution, and then the CPU performs the second iterative solution. At the same time, the DPU calculates the residual of the first iterative solution for convergence judgment. After the DPU judges that the residual of the result of the first iteration satisfies the convergence, the CPU completes the second iteration and uses the result of the second iteration as the final result for subsequent steps. At the same time, the DPU calculates the residual of the second iterative solution and directly uses it as the final residual record without performing a convergence judgment.
[0059] Or the residual has not satisfied the convergence judgment and reaches the maximum number of iterations, obtaining the iterative result of the maximum number of iterations and updating the flow velocity Ui for subsequent steps. At the same time, the DPU calculates the residual of the iterative result of the maximum number of iterations as the final residual record;
[0060] ②. The CPU performs iterative solution of the pressure Poisson equation and updates the pressure p , the DPU calculates the residual of the pressure Poisson equation and performs convergence judgment. This process repeats until the Nth residual calculated by the DPU satisfies the convergence judgment. Then the CPU performs N + 1 iterative solutions, obtains the iterative result of N + 1 times and updates the pressure p for subsequent steps. At the same time, the DPU calculates the residual of the iterative result of N + 1 times as the final residual record;
[0061] Or the residual has not satisfied the convergence judgment and reaches the maximum number of iterations, obtaining the iterative result of the maximum number of iterations and updating the updated flow velocity Ui for subsequent steps. At the same time, the DPU calculates the residual of the iterative result of the maximum number of iterations as the final residual record;
[0062] ③. Use the pressure p to correct the flow velocity Ui. The DPU calculates the residual of the mass conservation equation and performs convergence judgment. If it does not satisfy the convergence judgment, return to step ① to continue iterative calculation until the convergence criterion is met or the maximum number of times of the inner loop between steps ① - ③ (set to 3 times in this embodiment) is reached;
[0063] ④. Solve other physical quantities, using the method of step ① for the corresponding equations for solution, and at the same time solve the unsolved physical quantities through the already solved physical quantities until all physical quantities are solved; In step ④, other physical quantities include turbulent viscosity coefficient vt, turbulent kinetic energy k, and specific dissipation rate ω. The specific steps for solution include:
[0064] [1]. The CPU performs iterative solution of the turbulent kinetic energy k equation and updates the turbulent kinetic energy k. The DPU calculates the residual of the momentum conservation equation and performs convergence judgment. This process repeats until the Nth residual calculated by the DPU satisfies the convergence judgment. Then the CPU performs N + 1 iterative solutions, obtains the iterative result of N + 1 times and updates the turbulent kinetic energy k for subsequent steps,
[0065] or the CPU reaches the maximum number of iterations, obtaining the iterative result of the maximum number of iterations and updating the turbulent kinetic energy k for subsequent steps;
[0066] [2]. The CPU calculates the specific dissipation rate ω through the turbulent kinetic energy k, and then calculates the turbulent viscosity coefficient vt.
[0067] ⑤ After all physical quantities are solved, the DPU calculates the residual of the continuity equation using the solved physical quantities and determines whether the residual meets the convergence criterion for the outer loop between steps ① - ⑤; if not, return to step ① to continue the iterative calculation until the convergence criterion for the outer loop between steps ① - ⑤ is met or the maximum number of times for the outer loop between steps ① - ⑤ (set to 20 times in this embodiment) is reached, and then enter step S3.
[0068] S3: The DPU performs post - processing and result output, and determines whether this time step is the last time step. If not, it enters the next time step. In each time step, S2 and S3 are repeated until the last time step. The DPU writes the solved physical quantities to the hard disk. The specific steps for post - processing and result output are as follows:
[0069] After reaching the convergence criterion, use the results of the simulation calculation performed by the DPU for auxiliary calculations such as result monitoring and real - time post - processing; through the direct memory access technology of the DPU, directly initiate a data write request to the CPU memory controller to complete post - processing and write the results to memory; or use the PCIe passthrough technology to directly or after compression write the calculation results and post - processing results to the hard disk.
[0070] Specifically in this embodiment, the DPU outputs the numerical calculation results of flow field quantities such as the velocity field U, pressure field p, and turbulent kinetic energy k, and draws contour maps of the flow field, pressure distribution, turbulent intensity, etc. to evaluate the fluid behavior and calculation accuracy; repeat S2 and S3 until the last time step is completed. After completing the last time step, the DPU compresses the physical quantities and writes them to the hard disk.
[0071] As described above, it is only a preferred embodiment of the present invention and does not impose any form of limitation on the present invention. Although the present invention has been disclosed as above with a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art, without departing from the scope of the technical solution of the present invention, can make some changes or modifications to the above - disclosed technical content to obtain equivalent embodiments with equivalent changes. However, as long as it does not depart from the content of the technical solution of the present invention, any brief modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A computer system for accelerating fluid simulation by cooperating between a DPU and a CPU, characterized in that: It includes one or more computing nodes, which communicate with each other through the network. Each computing node is configured with the same number of CPUs and DPUs. Each computing node has one or more CPUs, and multiple CPUs communicate with each other through internal connections. The CPU is responsible for iteratively solving the fluid simulation control equations and writing the iterative results into the memory; Each computing node has one or more DPUs. The DPUs use a direct memory access controller to directly read memory data without going through the CPU memory controller, and calculate the residual of the iteration result and / or verify whether it meets the convergence criterion. When the current iteration result meets the convergence criterion, the CPU completes the ongoing iteration solution, and the DPU calculates the residual of the iteration result, performs post-processing and outputs the result. The computer system uses DPU and CPU to accelerate the solution of fluid simulation algorithm. The specific steps are as follows: S1: The computational domain is allocated to each computing node. The computational domain grid and physical quantities are stored in the memory. The CPU in each computing node sets the initial conditions and boundary conditions of the computational domain. S2: Each computing node performs parallel computation of fluid simulation according to task allocation. Specifically: The CPU advances the time step based on the initial conditions or the calculation results of the previous time step, iteratively solves the physical quantity in each time step and stores it in the memory. The DPU reads the boundary conditions and iterative results in the memory to calculate the residual and judge whether the residual meets the convergence criterion; the CPU reads the convergence judgment result of the DPU, if not, the CPU continues to iterate the solution; if satisfied, the CPU completes the ongoing iterative solution, and the DPU calculates the residual of the iterative result; S3: DPU performs post-processing and result output, and determines whether this time step is the last time step. If not, it enters the next time step, repeating S2 and S3 in each time step until the last time step. DPU writes the solved physical quantity to the hard disk.
2. The computer system for accelerating fluid simulation by cooperation between DPU and CPU according to claim 1, characterized in that: The specific steps in step S1 include: S11: Divide the computational domain into several blocks with relatively balanced numbers of grids according to the grid density and geometric characteristics of the computational domain; S12: Marking several units of the block interface so that the calculation results of the block nodes of different calculation nodes are equal during iterative calculation; S13: configure a parallel computing environment so that each block can be processed independently and in parallel on different computing nodes, and each computing node can access the physical quantities of the computing domain in other computing nodes through the network; S14: Specify initial values for each physical quantity in the computational domain, set boundary conditions, and apply the boundary conditions to the boundary grid of the computational domain.
3. The computer system for accelerating fluid simulation by cooperation between DPU and CPU according to claim 2, characterized in that: The time step is advanced based on the initial value or the calculation result of the previous time part. The specific steps for parallel calculation of fluid simulation based on physical quantity calculation nodes in each time step are: ①. The CPU iteratively solves the momentum conservation equation and updates the flow velocity Ui. The DPU calculates the residual of the momentum conservation equation and makes a convergence judgment. This cycle repeats until the Nth residual calculated by the DPU satisfies the convergence judgment. The CPU iterates and solves the equation N+1 times, obtains the N+1 iteration result, and updates the flow velocity Ui for subsequent steps. At the same time, the DPU calculates the residual of the N+1 iteration result as the final residual record. Or the residual does not satisfy the convergence judgment and reaches the maximum number of iterations, the iteration result of the maximum number of iterations is obtained and the flow rate Ui is updated for subsequent steps. At the same time, the DPU calculates the residual of the iteration result of the maximum number of iterations as the final residual record; ② The CPU iteratively solves the pressure Poisson equation and updates the pressure p. The DPU calculates the residual of the pressure Poisson equation and makes a convergence judgment. This cycle repeats until the Nth residual calculated by the DPU satisfies the convergence judgment. The CPU iterates and solves the N+1th time, obtains the N+1th iteration result and updates the pressure p for subsequent steps. At the same time, the DPU calculates the residual of the N+1th iteration result as the final residual record. Or the residual does not satisfy the convergence judgment and reaches the maximum number of iterations, the iteration result of the maximum number of iterations is obtained and the flow rate Ui is updated for subsequent steps. At the same time, the DPU calculates the residual of the iteration result of the maximum number of iterations as the final residual record; ③. Use pressure p to correct flow rate Ui. DPU calculates the residual of mass conservation equation and makes convergence judgment. If convergence judgment is not met, return to step ① and continue iterative calculation until convergence criterion is met or the maximum number of inner loops between step ① and step ③ is reached; ④. Solve other physical quantities by using the corresponding equations in step ①. At the same time, use the solved physical quantities to solve the unsolved physical quantities until all physical quantities are solved; ⑤. After all physical quantities are solved, the DPU calculates the residual of the continuity equation through the solved physical quantities and determines whether the residual meets the convergence criterion of the outer loop between steps ① and ⑤; if not, return to step ① to continue iterative calculation until the convergence criterion of the outer loop between steps ① and ⑤ is met or the maximum number of outer loops between steps ① and ⑤ is reached, and then enter step S3.
4. The computer system for accelerating fluid simulation by cooperation between DPU and CPU according to claim 3, characterized in that: The specific steps of post-processing and result output include: After reaching the convergence criterion, the DPU is used to perform auxiliary calculations such as simulation calculation result monitoring and real-time post-processing; Through DPU direct memory access technology, a data write request is directly initiated to the CPU memory controller, and after completion of the processing, the result is written to the memory; Or use PCIe pass-through technology to write the calculation results and post-processing results to the hard disk directly or after compression.
5. The computer system for accelerating fluid simulation by cooperation between DPU and CPU according to claim 3, characterized in that: In step ④, other physical quantities include the turbulent viscosity coefficient vt, the turbulent kinetic energy k and the specific dissipation rate ω. The specific solution steps include: [1] The CPU iteratively solves the turbulent kinetic energy k equation and updates the turbulent kinetic energy k. The DPU calculates the residual of the momentum conservation equation and makes a convergence judgment. This cycle repeats until the Nth residual calculated by the DPU satisfies the convergence judgment. The CPU iterates and solves the equation N+1 times, obtains the N+1th iteration result, and updates the turbulent kinetic energy k for subsequent steps. Or the CPU reaches the maximum number of iterations, obtain the iteration result of the maximum number of iterations and update the turbulent kinetic energy k for subsequent steps; [2] The CPU calculates the specific dissipation rate ω through the turbulent kinetic energy k, and then calculates the turbulent viscosity coefficient vt.
Citation Information
Patent Citations
A parallel algorithm transplantation optimization method based on domestic accelerator
CN118245039B
A computer network system and a computer system
CN118714020B
System and method for stabilizing and accelerating iterative numerical simulation
US20240126943A1
Identification of sub-graphs from a directed acyclic graph of operations on input data
US20240370301A1