Multiphase flow simulation method and system based on heterogeneous mixing precision acceleration
Through the heterogeneous hybrid accuracy acceleration framework, sensitive areas in multiphase flow simulation are dynamically identified, combined with GPU parallel computing and Kahan error compensation, the contradiction between accuracy and efficiency in multiphase flow simulation is solved, and efficient and reliable multiphase flow interface dynamic simulation is achieved.
Patent Information
- Application Number
- CN202510740554.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
Smart Images

Figure CN120257893A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of multiphase flow simulation, and particularly to a multiphase flow simulation method and system based on heterogeneous mixed-precision acceleration. Background Art
[0002] The simulation of multiphase flow interface dynamics has important application values in fields such as chemical reactor design and energy equipment optimization. For example, in scenarios such as oil-water separation and nuclear reactor cooling, the accurate simulation of droplet collision and bubble coalescence behaviors is the key to ensuring the efficiency and safety of equipment. However, achieving high precision and efficient calculation with limited hardware resources poses a great challenge to existing methods.
[0003] Based on the advantages of mesoscopic scale modeling and natural parallelism, the lattice Boltzmann method (LBM) has become the mainstream method for multiphase flow interface dynamics simulation. However, existing LBM frameworks generally adopt a static storage strategy and a global double-precision design. This approach of "global double-precision in sensitive regions" lacks real-time perception and feedback of local physical features such as gradients and curvatures, and cannot dynamically adjust the precision distribution according to the transient characteristics of the flow field. This static strategy is prone to error diffusion or numerical oscillation when dealing with interface mutations and flow couplings, thereby affecting the reliability of engineering-level simulations. Therefore, existing multiphase flow simulation methods still cannot well balance the video memory requirements and calculation precision. Summary of the Invention
[0004] To solve the above problems, the present disclosure proposes a multiphase flow simulation method and system based on heterogeneous mixed-precision acceleration, providing a feasible solution that balances efficiency and precision for industrial-level multiphase flow interface dynamics simulation through dynamic threshold identification and multi-level collaborative optimization.
[0005] According to some embodiments, the present disclosure adopts the following technical solutions: A multiphase flow simulation method based on heterogeneous mixed-precision acceleration, comprising: Obtaining information of a target object to be simulated; Using the information of the target object, establishing a hydrodynamics model through the lattice Boltzmann method, discretizing the continuous macroscopic behavior of the fluid into tiny grids and time steps, and simulating the macroscopic behavior of the fluid by the collision and migration of a number of particles in the grid-shaped fluid region; Based on an initial density distribution function, performing time-step iteration on the hydrodynamics model, and in each time step, performing parallel simulation of the macroscopic behavior of the fluid for all particles until the condition for stopping the iteration is satisfied, obtaining a multiphase flow simulation result; Among them, the parallel simulation of the macroscopic behavior of the fluid adopts a heterogeneous mixed-precision acceleration framework. Through the gradient-curvature collaborative criterion, sensitive regions in the fluid region are dynamically identified. Based on the identification results, mixed-precision calculations of double precision and single precision are performed, so as to perform parallel calculations on the collision process in the multiphase flow simulation.
[0006] According to some embodiments, the present disclosure adopts the following technical solutions: A multiphase flow simulation system based on heterogeneous mixed-precision acceleration, comprising: An acquisition module, configured to: acquire information of a target object to be simulated; A modeling module, configured to: use the information of the target object to establish a hydrodynamics model by means of the lattice Boltzmann method, discretize the continuous macroscopic behavior of the fluid into tiny grids and time steps, and simulate the macroscopic behavior of the fluid by the collision and migration of a number of particles in the grid-like fluid region; A simulation module, configured to: perform time-step iteration on the hydrodynamics model based on an initial density distribution function, perform parallel simulation of the macroscopic behavior of the fluid on all particles in each time step until the condition for stopping the iteration is met, and obtain a multiphase flow simulation result; Among them, the parallel simulation of the macroscopic behavior of the fluid adopts a heterogeneous mixed-precision acceleration framework. Through the gradient-curvature collaborative criterion, sensitive regions in the fluid region are dynamically identified. Based on the identification results, mixed-precision calculations of double precision and single precision are performed, so as to perform parallel calculations on the collision process in the multiphase flow simulation.
[0007] According to some embodiments, the present disclosure adopts the following technical solutions: A computer program product, comprising a computer program, where when the computer program is executed by a processor, the described multiphase flow simulation method based on heterogeneous mixed-precision acceleration is implemented.
[0008] According to some embodiments, the present disclosure adopts the following technical solutions: A non-transitory computer-readable storage medium, where the non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the described multiphase flow simulation method based on heterogeneous mixed-precision acceleration is implemented.
[0009] According to some embodiments, the present disclosure adopts the following technical solutions: An electronic device, comprising: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory, so that the electronic device executes and implements the described multiphase flow simulation method based on heterogeneous mixed-precision acceleration.
[0010] Compared with the prior art, the beneficial effects of the present disclosure are as follows: The present disclosure provides a multiphase flow simulation method and system based on heterogeneous mixed-precision acceleration. Through dynamic threshold identification and multi-level collaborative optimization, a feasible solution that takes into account both efficiency and accuracy is provided for industrial-level multiphase flow interface dynamics simulation. Specifically: 1) Adaptive dynamic precision allocation: Identify sensitive regions through the gradient-curvature collaborative criterion. Maintain double-precision FP64 calculation in high-gradient regions and use single-precision FP32 in low-gradient regions. Combine Kahan cumulative error compensation to stabilize the mass conservation error within 0.1%.
[0011] 2) Full-stack optimization of storage-communication-computation: Through reconstructing the compressed storage architecture and double-buffered communication, organize mixed-precision data more effectively and cooperate with subsequent communication processing. In the case of 141.4 million grids in the D3Q19 discrete model, the video memory occupancy is reduced by about 40%, the communication efficiency is increased by 77%, and the maximum performance acceleration benefit can reach 30 times. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The schematic diagrams in the specification forming a part of the present disclosure are used to provide a further understanding of the present disclosure. The schematic embodiments and descriptions thereof of the present disclosure are used to explain the present disclosure and do not constitute an improper limitation to the present disclosure.
[0013] Figure 1 It is a flowchart of the multiphase flow simulation method for Embodiment 1. Figure 2 It is a framework diagram of heterogeneous mixed-precision acceleration for Embodiment 1. Figure 3 It is a flowchart of generating a marker array for Embodiment 1. Figure 4 It is a schematic diagram of marker-driven compression mapping for Embodiment 1.
[0014] Figure 5 It is a schematic diagram of the reconstruction of the storage layout for Embodiment 1.
[0015] Figure 6 It is an overall acceleration ratio diagram of GDHP-SAF for Embodiment 1 in the grid scale range from 64³ to 521³.
[0016] Figure 7 It is an acceleration ratio diagram of the collision kernel (excluding communication) for Embodiment 1 across various resolutions. Figure 8 It is a comparison diagram of the video memory usage for Embodiment 1. Figure 9 It is a diagram of the influence of the threshold coefficient on the mass conservation error for Embodiment 1. Figure 10 It is a diagram of the droplet interface morphology under the original Palabos version for Embodiment 1. Figure 11 It is the droplet interface morphology diagram under the GDHP-SAF framework of Example 1. Specific implementation manners
[0017] The present disclosure will be further described below in conjunction with the accompanying drawings and embodiments.
[0018] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further descriptions of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present disclosure belongs.
[0019] Example 1 In one embodiment of the present disclosure, a multiphase flow simulation method based on heterogeneous mixed-precision acceleration is provided. As Figure 1 shown, it includes: Step S1: Obtain the information of the target object to be simulated; Step S2: Utilize the information of the target object to establish a hydrodynamics model through the lattice Boltzmann method, discretize the continuous macroscopic behavior of the fluid into tiny grids and time steps, and simulate the macroscopic behavior of the fluid by the collision and migration of several particles in the grid-shaped fluid region; Step S3: Based on the initial density distribution function, perform time-step iteration on the hydrodynamics model, and perform parallel simulation of the macroscopic behavior of the fluid for all particles in each time step until the condition for stopping the iteration is met, and obtain the multiphase flow simulation result; Among them, the parallel simulation of the macroscopic behavior of the fluid adopts a heterogeneous mixed-precision acceleration framework, dynamically identifies sensitive regions in the fluid region through a gradient-curvature collaborative criterion, and based on the identification result, performs mixed-precision calculation of double precision and single precision, so as to perform parallel calculation on the collision process in the multiphase flow simulation.
[0020] As an embodiment, a multiphase flow simulation method based on heterogeneous mixed-precision acceleration of the present disclosure provides a feasible solution that takes into account both efficiency and accuracy for industrial-level multiphase flow interface dynamics simulation. The specific implementation process will be described below.
[0021] Early single-phase flow research has confirmed the feasibility of single-precision FP32 for single-phase flow; however, in the simulation of liquid-gas interfaces, single-precision FP32 is prone to introduce cumulative interface tension errors, resulting in non-physical oscillations. In recent years, the research trend in the field of scientific computing has gradually shifted towards the development of mixed-precision algorithms.
[0022] Current research on mixed precision mainly focuses on single-phase flow scenarios, usually allocating FP64 / FP32 computing resources with a fixed partitioning strategy, which cannot be dynamically switched according to the transient interface evolution. This method has two defects in multiphase flow simulations: First, the fixed precision is difficult to meet the strict requirements of double-precision calculations in the transient high-gradient interface region, resulting in an interface tension error exceeding 1.5% and continuous accumulation; Second, if relying solely on FP32 in low-gradient regions, the global mass conservation deviation will exceed the industrial standard (<1%
[18] ) due to the continuous accumulation of truncation errors.
[0023] More critically, existing methods lack real-time perception and feedback of local physical characteristics such as gradients and curvatures, and cannot dynamically adjust the precision distribution according to the transient characteristics of the flow field. This static strategy is prone to error diffusion or numerical oscillation when dealing with interface mutations and flow coupling, thereby affecting the reliability of engineering-level simulations; Therefore, there is an urgent need to develop an adaptive mixed-precision scheme based on dynamic perception of physical field characteristics to synergistically optimize video memory efficiency and numerical stability.
[0024] To solve the above problems, this embodiment proposes a gradient-driven heterogeneous mixed-precision acceleration framework (GDHP-SAF) for parallel collision calculations in multiphase flow simulations. This framework uses dynamic threshold identification to perform double-precision (FP64) calculations in high-gradient regions to ensure accurate capture of interface tension, while using single-precision (FP32) calculations in low-gradient regions. Combining with Kahan compensation technology, it effectively reduces video memory occupancy and truncation errors. This framework extends dynamic precision perception and three-level mixed-precision pipelining on the Palabos platform, and reconstructs the compressed storage and double-buffer communication protocol, reducing video memory occupancy by 40% and improving performance by about 30 times in a 141 million grid case.
[0025] First, a brief description of multiphase flow simulations based on the lattice Boltzmann method (LBM): The lattice Boltzmann method (LBM) is a fluid simulation method based on a microscopic kinetic model that simulates macroscopic fluid behavior through a discretized velocity space and simple collision-migration rules. The following are the core steps of LBM simulation: (1) Discretize the velocity space (select a discrete model) LBM uses discrete velocity directions to describe particle motion and discretizes the macroscopic behavior of fluids through the DnQm discrete model with n dimensions and m velocities. For example, D2Q9 has 9 velocity directions (including stationary particles) in a 2D space, and D3Q19 has 19 velocity directions in a 3D space. Each direction has corresponding weight coefficients and velocity vectors. This embodiment uses D3Q19.
[0026] (2) Initialize the distribution function Define the distribution function , representing the particle density in the grid x at time step t in direction i. Usually, the initial distribution function is set to the equilibrium distribution.
[0027] (3)Collision step Particles undergo local interactions at the grid nodes, and calculate the distribution function after collision: (1) Among them, is the particle density in the grid x at the t-th time step in direction i, that is, the distribution function before collision, is the multi-relaxation-time (MRT) collision operator.
[0028] (4)Streaming step Move the distribution function after collision along the velocity direction to the adjacent node, and update the distribution function for the next time step, which is expressed by the formula: (2) Among them, are the 19 discrete velocity directions of the D3Q19 model, is the time step size.
[0029] This step explicitly updates the distribution function without solving partial differential equations.
[0030] (5)Calculate macroscopic quantities Calculate macroscopic variables through the zero-order and first-order moments of the distribution function, which is expressed by the formula:
[0031]
[0032] Among them, is the density, is the macroscopic velocity.
[0033] (6)Iterative advancement Repeat the collision-streaming-calculate macroscopic quantities steps until reaching the steady state or the required number of time steps.
[0034] In the above simulation process, this embodiment uses the heterogeneous mixed-precision acceleration framework (GDHP-SAF) to accelerate the collision step on the GPU. The GPU is good at parallel processing of highly localized and data-independent computational tasks, and in the collision step, the distribution function on each grid is updated independently, only depending on the distribution function and local macroscopic quantities (i.e., density , macroscopic velocity ), meeting the requirements of "highly localized and data-independent", all grids can be synchronously and parallelly computed without cross-grid dependencies. Therefore, using a GPU, each GPU thread processes a grid block composed of several grids, and uses shared memory (Shared Memory) to cache frequently accessed data (such as distribution functions and velocity directions and local macroscopic variables), and through memory coalescing optimization, random access to global memory is avoided.
[0035] Based on the above content, the structure of the heterogeneous mixed-precision acceleration framework (GDHP-SAF) is described in detail.
[0036] Figure 2 is the architecture diagram of GDHP-SAF, which is divided into three steps: generating a marker array, optimizing storage, and mixed-precision calculation. Among them, generating the marker array and optimizing storage run on the CPU side, while mixed-precision calculation runs on the GPU side. The following is a separate description.
[0037] 1. Generating a marker array Based on gradient-curvature analysis, real-time perception and feedback of local physical characteristics such as gradients and curvatures, and according to the transient characteristics of the flow field in the fluid region, the precision distribution is dynamically adjusted to generate a marker array, as Figure 3 shown specifically as: (1) Calculation of the gradient extreme value within a block The fluid region is divided into multiple grid blocks, and each grid block contains several grids. Collision calculations are performed in units of grid blocks. Therefore, here, the gradient extreme value within the block is first calculated through an in-block reduction operation, denoted by , where is a grid block, and is the concentration gradient of each grid.
[0038] (2) Generation of dynamic thresholds Based on the gradient extreme value within the block and combined with the curvature field of the current block, a dynamic threshold is dynamically generated for each grid in the block, and is expressed by the formula:
[0039] In the formula, is the gradient extreme value within the block, is the curvature gradient (reflecting the discrete spatial change rate of curvature), and the parameters α and β are calibrated based on grid sensitivity analysis and the Young-Laplace pressure balance criterion to ensure that the dynamic threshold can both suppress gradient noise and avoid over-triggering double-precision calculations.
[0040] (3) Generation of the marker array Through the concentration gradient of each grid ( ), curvature gradient ( ) is compared with the threshold to obtain the 0 / 1 mark of each grid, which forms the mark array flag[N] (N is the total number of grids).
[0041] Specifically, when the i-th grid or , flag[i]=1, otherwise flag[i]=0.
[0042] Based on the above content, the pseudo code of the tag array is generated, as shown in Table 1: Table 1 Label array generation function
[0043] 2. Optimize storage In order to organize mixed-precision data more effectively and cooperate with subsequent communication and computing processing, this embodiment proposes a two-stage optimized storage strategy: mark-driven compression mapping and storage layout reconstruction, specifically: like Figure 4 As shown in the figure, in the mark-driven compression mapping, the mark array flag[N] (N is the total number of grids) generated in the previous step is used to divide the grid data arr[N*19] to be calculated into two groups: double precision (D_arr) and single precision (F_arr). When the mark flag[i] of the i-th grid is 1, the 19 distribution function data of the grid are stored in the array D_arr in double precision FP64 format, and the index is maintained by D_map[i]; if flag[i]=0, it is stored in the array Far_arr in single precision FP32 format and identified by F_map[i].
[0044] This method can alleviate the pressure on video memory in large-scale multiphase flow simulations and ensure that large-scale calculations can be successfully completed within a single GPU card.
[0045] Figure 5 The different storage methods of AoS (array structure) and SoA (array of structure) are compared. In AoS mode, all distribution functions of a single grid are arranged continuously according to physical properties. However, when global data in a specific direction (such as f0) needs to be read in batches, it is easy to cause strided access, thereby reducing cache efficiency. The SoA layout stores the distribution functions in the same direction in a centralized manner. This design significantly improves the pre-fetch efficiency of the GPU cache, and the bandwidth utilization rate jumps to 68%. In addition, the continuous data blocks generated by SoA can be directly encapsulated as communication messages, simplifying the subsequent packaging and transmission processes; therefore, this embodiment converts the array structure AoS into the structure array SoA, and stores the grid data in the same direction in the simulation in a centralized manner to complete the storage layout reconstruction.
[0046] 3. Mixed Precision Computing Based on storage optimization, GDHP-SAF enables deep parallelism between communication and computing on the GPU through asynchronous double buffering, multiple CUDA streams, and task scheduling.
[0047] Specifically, multiple CUDA streams are used to handle simultaneously: one stream is responsible for the collision calculation of the current time step, that is, iteratively calculating the distribution function values of each time step through formula (1), and another stream executes the data transfer of the previous time step in parallel, thereby reducing the blocking of communication to computing. Among them, combined with dynamic precision allocation, the stream responsible for collision calculation is further divided into a double-precision pipeline and a single-precision pipeline. Together with the stream for data transfer, they form a three-level pipeline consisting of double-precision, single-precision, and communication.
[0048] Since the data blocks output in SoA format are already differentiated into double-precision and single-precision in terms of type, the communication pipeline can directly perform message packing and send it to the GPU for underlying double-precision and single-precision collision calculations, that is, iteratively calculating the distribution function values of each time step through formula (1), avoiding frequent conversions. Experimental results show that this mechanism reduces the communication delay from 12.8 milliseconds to 2.9 milliseconds, and the communication efficiency is increased by approximately 77%.
[0049] In the collision calculations based on the double-precision pipeline and the single-precision pipeline, through task scheduling, that is, instruction-level restructuring and hardware resource scheduling, the computing power of the GPU is fully released. Specifically: First, fuse the fused multiply-add (FFMA) and load global (LDG) instructions to reduce the number of instructions and register occupancy in the single-precision path, and improve the thread block scheduling efficiency of the streaming multiprocessor (SM). Second, adopt the loop unrolling strategy to improve the instruction-level parallelism (ILP), significantly increasing the computing power utilization rate of the single-precision kernel function.
[0050] Finally, dynamically allocate SM resources according to the characteristics of the Ampere architecture: the high-gradient region preferentially executes FP64 calculations to ensure the accuracy of the interfacial tension, and the low-gradient region maximizes the FP32 data transfer efficiency through a preemptive communication channel, ultimately achieving load balancing between computing and communication.
[0051] In addition to the above improvements in task scheduling to release the computing power of the GPU, the rounding error in single-precision calculations is also handled.
[0052] First, explain the sources of rounding errors in single-precision calculations. In a computer, there are precision limitations in the storage and calculation of floating-point numbers (especially single-precision FP32). Firstly, it is the representation with a finite number of digits. A single-precision floating-point number uses 32 bits (1 bit for the sign, 8 bits for the exponent, and 23 bits for the mantissa), and it cannot precisely represent all real numbers. Secondly, there is the binary conversion error. For example, the decimal number 0.1 is an infinite repeating decimal (0.0001100110011...) in binary and must be truncated or rounded. Finally, there is the arithmetic operation truncation. In addition, subtraction, multiplication, and division operations, the part exceeding the mantissa bits will be discarded or approximated, resulting in tiny errors.
[0053] To reduce the impact of rounding errors on the accuracy of single-precision calculations, in the single-precision collision calculation of the single-precision pipeline, the Kahan error compensation mechanism is introduced. By temporarily storing the rounding errors of single-precision calculations and dynamically correcting them in subsequent calculations, the error accumulation is suppressed. As shown in Table 2, specifically: (1)Initialize the accumulated sum (sum) and the compensation amount (compensation) used to temporarily store the rounding error; (2)Process the distribution function values in different directions of the grid item by item : Through the distribution function value Remove the compensation amount compensation of the previous round to calculate the intermediate variable y; Define the sum of the accumulated sum sum and the intermediate variable as the temporary accumulated sum; Extract the rounding error compensation of this calculation by subtracting the accumulated sum sum and the intermediate variable y from the temporary accumulated sum temp_sum respectively; Overwrite the original value of the distribution function value with the accumulated sum sum to ensure that the error is dynamically corrected during iteration.
[0054] Table 2 Kahan error compensation function
[0055] This compensation mechanism temporarily stores the rounding error in a dedicated variable (i.e., the compensation amount compensation in Table 2) and dynamically adjusts the distribution function value in the next iteration cycle ( ), so that the mass conservation error is stabilized within 0.1%. Compared with the static strategy, this design effectively suppresses the error diffusion phenomenon in the interface evolution while reducing the video memory occupancy.
[0056] In this embodiment, the performance of the GDHP-SAF framework in terms of performance and video memory optimization, numerical accuracy, and physical fidelity was evaluated through experiments. The experimental environment was based on an NVIDIA A100 GPU (40 GB video memory), and the typical droplet collision problem on Palabos was used as the test object to test the actual effect of the framework. All experiments used the D3Q19 discrete model, and the average value of three independent trials (standard deviation < 2%) was taken for each group of results.
[0057] A. Performance and Video Memory Optimization Figure 6 And Figure 7 The acceleration ratio difference between GDHP-SAF and the previous full double-precision model within the executable grid scale was compared, Figure 6 showing the overall acceleration ratio of GDHP-SAF in the grid scale range from 64³ to 521³, demonstrating nearly linear weak scalability. At 420³, it achieved an acceleration ratio of 25.2 times, which was better than the 17.7 times achieved by the previous full double-precision model. At 521³, the implementation of the full FP64 model exceeded the 40 GB memory capacity of the A100 GPU, while GDHP-SAF reduced the memory usage to 37.2 GB through adaptive precision switching. Compared with the CPU reference, this achieved an overall acceleration of 30.5 times.
[0058] Figure 7 The acceleration ratio of the collision kernel (excluding communication) across various resolutions was shown. GDHP-SAF was always superior to the complete FP 64 benchmark model, and the kernel acceleration range was from 1.8 times to 2.5 times. At 256³, the acceleration increased from 446 times (FP 64) to 887 times (GDHP-SAF), with a gain of 1.99 times. This advantage increased with the increase in scale: at 420³, GDHP-SAF reached 936 times, while FP 64 was 524 times. At 521³, it maintained a kernel acceleration ratio of 965 times, while the FP 64 model could not execute due to memory limitations. These results emphasized the scalability and robustness of GDHP-SAF in a GPU environment with limited memory.
[0059] Figure 8 The comparison of video memory usage was shown. Under the 420³ grid, GDHP-SAF reduced the video memory requirement from 32.2 GB to 19.5 GB through dynamic compression storage, a reduction of approximately 39.4%. When the grid scale was less than 256³, the video memory difference between mixed precision and full double precision was less than 9%, which was mainly due to the dominance of fixed storage overheads (such as index tables, marker arrays). It can be seen that the larger the simulation scale, the more obvious the video memory advantage of GDHP-SAF, especially suitable for industrial scenarios with more than hundreds of millions of grid scales.
[0060] B. Numerical Accuracy Verification To verify the dynamic precision allocation, this embodiment analyzes the influence of the threshold coefficient α on the mass conservation error. Figure 9 The influence of the threshold coefficient α on the mass conservation error is shown. In the droplet collision case of the 521³ grid, when α = 0.25, the mass conservation error is 1.0%, meeting the industrial standard (<1%
[18] ); after increasing α to 0.30, the error rises to 1.1%, still significantly better than 1.5% of the static mixed precision scheme [3]. Comparing different grid scales, it is found that when α increases from 0.25 to 0.30, the error of the 121³ grid rises from 1.5% to 1.7% (relative increase of 13.3%), while the 521³ grid only rises from 1.0% to 1.1% (relative increase of 10%), indicating that higher resolution helps to reduce the parameter sensitivity and improve the robustness of dynamic allocation.
[0061] It should be noted that existing mixed precision studies have achieved an error <1% in single-phase flow (Poiseuille flow), but their tests did not cover complex multiphase flow scenarios. GDHP-SAF, through dynamic precision allocation, stably controls the mass conservation error within 1% in multiphase scenarios such as droplet collision, expanding the applicable range of the mixed precision strategy. In addition, the threshold coefficient β is used to identify the curvature-sensitive region. For the current droplet collision problem, the interfacial tension error is mainly determined by the gradient component, and the influence of the curvature effect can be ignored.
[0062] C. Verification of Physical Rationality As Figure 10 、 Figure 11 shown, the droplet interfaces in the original Palabos version and the GDHP-SAF framework both exhibit smooth and continuous evolution, without interface breakage or non-physical oscillations. Based on the precision and performance results, GDHP-SAF, by means of dynamic precision allocation, maintains double-precision calculations in high-gradient regions to accurately capture the interfacial tension, and uses single-precision combined with compensation techniques in low-gradient regions, reducing the video memory occupancy by about 40% while maintaining an interfacial shape fidelity comparable to that of the full double-precision model. The overall experiment verifies the engineering practicability of this framework from multiple perspectives of precision, efficiency, and physical rationality.
[0063] This embodiment proposes a gradient-driven heterogeneous mixed precision acceleration framework (GDHP-SAF). Through multiple optimization means such as gradient-curvature determination, Kahan compensation, and SoA storage, in the 141.4 million grid droplet collision test, compared with the full double-precision scheme, GDHP-SAF saves about 40% of the video memory while maintaining industrial-grade precision, and achieves a performance improvement of up to 30 times, significantly reducing the GPU resource occupancy of high-resolution multiphase lattice Boltzmann simulations and fully demonstrating its robustness and efficiency in complex interfacial dynamics.
[0064] Example 2 In one embodiment of the present disclosure, a multiphase flow simulation system based on heterogeneous mixed-precision acceleration is provided, including: An acquisition module, configured to: acquire target object information to be simulated; A modeling module, configured to: use the target object information to establish a hydrodynamic model by the lattice Boltzmann method, discretize the continuous macroscopic behavior of the fluid into tiny grids and time steps, and simulate the macroscopic behavior of the fluid by the collision and migration of several particles in the grid-based fluid region; A simulation module, configured to: perform time-step iteration on the hydrodynamic model based on the initial density distribution function, and perform parallel simulation of the macroscopic behavior of the fluid for all particles in each time step until the iteration stop condition is satisfied, to obtain a multiphase flow simulation result; Wherein, the parallel simulation of the macroscopic behavior of the fluid is to adopt a heterogeneous mixed-precision acceleration framework, dynamically identify sensitive regions in the fluid region through a gradient-curvature collaborative criterion, and perform mixed-precision calculation of double precision and single precision based on the identification result, so as to perform parallel calculation on the collision process in multiphase flow simulation.
[0065] Example 3 In one embodiment of the present disclosure, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, it implements the multiphase flow simulation method based on heterogeneous mixed-precision acceleration as described above.
[0066] Example 4 In one embodiment of the present disclosure, a non-transitory computer-readable storage medium is provided, and the non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the multiphase flow simulation method based on heterogeneous mixed-precision acceleration as described above is implemented.
[0067] Example 5 In one embodiment of the present disclosure, an electronic device is provided, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory, so that the electronic device executes and implements the multiphase flow simulation method based on heterogeneous mixed-precision acceleration as described above.
[0068] Although the specific embodiments of the present disclosure are described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that, based on the technical solutions of the present disclosure, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present disclosure.
Claims
1. A multiphase flow simulation method based on heterogeneous mixed-precision acceleration, characterized in that Including: Obtain the information of the target object to be simulated; Using the information of the target object, establish a hydrodynamic model through the lattice Boltzmann method, discretize the continuous macroscopic behavior of the fluid into tiny grids and time steps, and simulate the macroscopic behavior of the fluid by the collision and migration of several particles in the grid-based fluid region; Based on the initial density distribution function, perform time-step iteration on the hydrodynamic model. In each time step, perform parallel simulation of the macroscopic behavior of the fluid for all particles until the iteration stop condition is met, and obtain the multiphase flow simulation result; Among them, the parallel simulation of the macroscopic behavior of the fluid is to adopt a heterogeneous mixed-precision acceleration framework, dynamically identify sensitive regions in the fluid region through the gradient-curvature collaborative criterion, and based on the identification result, perform mixed-precision calculation of double precision and single precision, so as to perform parallel calculation on the collision process in the multiphase flow simulation.
2. The multiphase flow simulation method based on heterogeneous mixed-precision acceleration according to claim 1, wherein, The heterogeneous mixed-precision acceleration framework includes an application layer, a heterogeneous layer, and an optimization layer; Among them, the application layer dynamically identifies sensitive regions in the fluid region through the gradient-curvature collaborative criterion, generates a marker array as the identification result; The heterogeneous layer, based on the identification result, adopts a three-stage pipeline composed of double precision, single precision, and communication, and coordinates communication and calculation with a double-buffer protocol to finally simulate the collision process; The optimization layer performs full-stack optimization of storage-communication-computation through reconstructing the compressed storage architecture and double-buffer communication.
3. The multiphase flow simulation method based on heterogeneous mixed-precision acceleration according to claim 2, wherein The dynamic identification of sensitive regions in the fluid region through the gradient-curvature collaborative criterion and generating a marker array is specifically: Use intra-block reduction to obtain the intra-block gradient extreme value; Based on the intra-block gradient extreme value, combined with the curvature field of the current block, dynamically generate a threshold for each grid in the block; Through the comparison of the gradient, curvature, and threshold of each grid, obtain the 0 / 1 mark of each grid, and form a marker array.
4. The multiphase flow simulation method based on heterogeneous mixed-precision acceleration according to claim 3, characterized in that, The dynamic generation of a threshold for each grid in the block is expressed by the formula: Among them, is the gradient extreme value within the block, is the curvature mutation intensity, and α and β are two parameters.
5. The multiphase flow simulation method based on heterogeneous mixed-precision acceleration according to claim 3, wherein The three-stage pipeline composed of double precision, single precision, and communication is specifically: The double-precision pipeline and the single-precision pipeline respectively select the corresponding double-precision floating-point format and single-precision floating-point format according to the 0 / 1 mark of the grid, and through double-precision calculation and single-precision calculation, combined with Kahan cumulative error compensation, perform parallel collision calculation for the current time step; The communication pipeline is based on the double-buffer protocol, constructs a double-precision buffer area and a single-precision buffer area, parallelly executes data transmission for the previous time step, and saves the transmitted data to the corresponding buffer area according to the 0 / 1 mark of the grid.
6. The multiphase flow simulation method based on heterogeneous mixed-precision acceleration according to claim 2, wherein, The reconstructed compressed storage architecture includes marker-driven compression mapping and storage layout reconstruction; Among them, the marker-driven compression mapping is to compress the grid data in a grid grouping manner according to the 0 / 1 mark of the grid to form an array structure; The storage layout reconstruction is to convert the array structure into a structure array and centrally store the grid data in the same direction in the simulation.
7. A multiphase flow simulation system based on heterogeneous mixed-precision acceleration, characterized in that, Including: An acquisition module, configured to: obtain the information of the target object to be simulated; A modeling module, configured to: utilize target object information to establish a hydrodynamics model through the lattice Boltzmann method, discretize the continuous macroscopic behavior of the fluid into tiny grids and time steps, and simulate the macroscopic behavior of the fluid by the collision and migration of a number of particles in the grid-based fluid region; A simulation module, configured to: perform time-step iteration on the hydrodynamics model based on the initial density distribution function, and perform parallel simulation of the macroscopic behavior of the fluid for all particles in each time step until the condition for stopping the iteration is met, to obtain a multiphase flow simulation result; Wherein, the parallel simulation of the macroscopic behavior of the fluid adopts a heterogeneous mixed-precision acceleration framework, dynamically identifies sensitive regions in the fluid region through a gradient-curvature collaborative criterion, and performs mixed-precision calculation of double precision and single precision based on the identification result, so as to perform parallel calculation on the collision process in the multiphase flow simulation.
8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements a multiphase flow simulation method based on heterogeneous mixed-precision acceleration according to any one of claims 1-6.
9. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, it implements a multiphase flow simulation method based on heterogeneous mixed-precision acceleration according to any one of claims 1-6.
10. An electronic device, characterized in that, Including: A processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory, so that the electronic device executes and implements a multiphase flow simulation method based on heterogeneous mixed-precision acceleration according to any one of claims 1-6.
Citation Information
Patent Citations
Multi-relaxation lattice Boltzmann model-based underground water flowing simulation acceleration method
CN107515987A
Super-computing Internet multi-target workflow optimization method and system based on ant colony algorithm
CN117670005A
Liquid chromatogram flow velocity real-time detection and optimal control method based on multi-point sensing
CN120065760A
Lattice boltzmann solver enforcing total energy conservation
JP2020123325A