An OpenFoam data reorganization method, medium and system

CN122614586BActive Publication Date: 2026-09-22青岛国实科技集团有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611031490.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-13
Publication Date
2026-09-22
Estimated Expiration
2046-07-13

AI Technical Summary

Technical Problem

[0003]有鉴于此,本发明提供一种OpenFoam数据重组方法、介质及系统,能够解决现有技术中存在OpenFOAM大规模并行仿真数据重组阶段因串行执行模式导致超算算力无法被有效利用的技术问题

Benefits of technology

[0027]本发明通过将有效仿真时间步转化为以单时间步为粒度的完全独立并行任务,并引入基于深度学习的并行核数自适应评估模型动态确定最优并行核数,结合资源安全系数函数对并行度进行动态修正,再借助动态任务池管控机制与三重隔离机制实现无冲突并行写入,将数据重组阶段从全局强串行段重构为高度并行化的多核协同工作流,使得Amdahl定律中串行部分占比趋近于0,系统加速比可随并行进程数线性提升,超算平台的高并发算力潜力得以充分释放。本发明解决了OpenFOAM大规模并行仿真数据重组阶段因串行执行模式导致超算算力无法被有效利用的技术问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122614586B_ABST
    Figure CN122614586B_ABST
Patent Text Reader

Abstract

The application provides an OpenFoam data reorganization method, medium and system, belongs to the technical field of OpenFoam data reorganization, and extracts effective simulation time step parameters through a scanning processor block subdirectory, dynamically determines optimal parallel core numbers by using a parallel core number self-adaptive evaluation model based on deep learning in combination with a resource safety coefficient function, creates a temporary log directory to centrally manage running states, adopts a dynamic task pool management and control mechanism to convert each time step into an independent parallel task, realizes conflict-free parallel writing by means of a triple isolation mechanism, keeps computing resources full by real-time task replenishment, and finally waits for all tasks to be completed and outputs a reorganization state log, so that the technical problem that computing power under the assistance of supercomputing artificial intelligence cannot be effectively utilized due to a serial execution mode in the OpenFOAM large-scale parallel simulation data reorganization stage is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of OpenFoam data reconstruction technology, and specifically relates to an OpenFoam data reconstruction method, medium and system. Background Technology

[0002] OpenFOAM is an open-source computational fluid dynamics simulation platform widely used in industry and academia. Its parallel simulation process decomposes the computational domain into multiple processor blocks, with each MPI process independently completing the local numerical solution. After the numerical solution phase, the reconstructPar tool must be called to merge the local field data of each processor block into a complete global dataset for subsequent visualization and analysis. On large-scale supercomputing platforms such as Sunway, the mesh size of CFD simulation tasks can reach tens of millions or even hundreds of millions of elements, with the corresponding number of effective simulation time steps often reaching hundreds to thousands. In existing technologies, because the reconstructPar tool adopts a global serial execution mode, it processes all effective simulation time steps sequentially. In large-scale simulation examples, when the number of effective simulation time steps is large and the data volume of a single step is significant, the time consumed in the serial reconstructing phase increases linearly with the number of time steps. During this phase, a large number of computing cores on the supercomputing node are idle and waiting. According to Amdahl's law, when the proportion of the serial part approaches 1, the overall speedup approaches 1, and the high-concurrency computing potential of the supercomputing platform is completely locked by the serial post-processing. In other words, existing technologies suffer from a technical problem where supercomputing power cannot be effectively utilized due to the serial execution mode during the OpenFOAM large-scale parallel simulation data reassembly stage. Summary of the Invention

[0003] In view of this, the present invention provides an OpenFoam data reconstruction method, medium and system, which can solve the technical problem in the prior art that the supercomputing power cannot be effectively utilized due to the serial execution mode in the OpenFOAM large-scale parallel simulation data reconstruction stage.

[0004] The present invention is implemented as follows: The first aspect of the present invention provides an OpenFoam data reconstruction method, comprising the following steps:

[0005] Scan all processor block subdirectories under the simulation root directory, remove constant directories and cache files, count the total number of effective simulation time steps, extract the minimum and maximum time steps, and complete the adaptive calculation of core parameters.

[0006] Based on the total number of effective simulation time steps, the optimal number of parallel cores is calculated through the adaptive evaluation model of parallel cores, and the running configuration information is output to complete the parameter verification and status output before running.

[0007] Create a temporary log directory to centrally store the running logs and error messages of each parallel reassembly task;

[0008] The system iterates through all valid simulation time steps, transforms each time step into an independent scheduling and computing unit based on a dynamic task pool management mechanism, and achieves conflict-free parallel writing by relying on a triple isolation mechanism. Each process independently completes the reorganization of the grid and flow field data for the corresponding time step.

[0009] The system monitors the number of concurrent tasks in the dynamic task pool in real time. When the number of concurrent tasks is lower than the optimal number of parallel cores threshold, it automatically pulls subsequent pending time step tasks to supplement the task pool.

[0010] Wait for all background parallel reorganization tasks to complete, release computing resources, and output a reorganization completion status log.

[0011] Specifically, the processor block subdirectory refers to the subdirectory that is independently generated and stores local partitioned grid data and flow field data by each MPI process during the OpenFOAM parallel simulation process. The naming rule is processor0 to processorN, where N is the total number of parallel processes minus 1.

[0012] Specifically, the effective simulation time step refers to the set of all named subdirectories obtained after scanning the processor0 subdirectory using the directory traversal algorithm, excluding the constant directory and the initial time step directory 0. The total number of effective simulation time steps is denoted as […]. The minimum time step is denoted as The maximum time step is recorded as .

[0013] The parallel core number adaptive evaluation model is based on deep learning. The input layer receives a 7-dimensional feature vector, which is determined by the total number of effective simulation time steps. Average data volume per time step Total number of processor block subdirectories Available memory capacity of nodes Maximum number of parallel cores per node Current bandwidth load of parallel file system Average single-core recombination time in historical operations constitute.

[0014] The structure of the parallel core number adaptive evaluation model is as follows: after the input layer, a batch normalization layer is connected; after the batch normalization layer, a first fully connected layer, a first Dropout layer, a second fully connected layer, a second Dropout layer, a third fully connected layer, and an output layer are connected in sequence to output the optimal parallel core number prediction value. Furthermore, a residual skip connection is introduced between the first fully connected layer and the second fully connected layer.

[0015] The training of the parallel core number adaptive evaluation model uses the Adam optimizer and the mean squared error loss function. When the validation set loss continuously reaches the non-decreasing round number threshold, the learning rate decay is triggered, and the decay coefficient is the decay coefficient value. The model weight corresponding to the minimum validation set loss is saved.

[0016] Among them, the optimal parallel core number prediction value Through the resource security factor function Make corrections to the resource security factor function. ,in and These are the weighting coefficients. This represents the current free memory capacity of the node.

[0017] Wherein, the resource security coefficient function Based on upper and lower thresholds, the task pool is divided into three operating modes: full load mode, load reduction mode, and conservative mode. Each of the three operating modes corresponds to a different final number of parallel cores. Calculation rules and task submission interval control strategies.

[0018] The core scheduling logic of the dynamic task pool management mechanism is as follows: the system iterates through the list of effective simulation time steps, counts the number of currently running background tasks in real time, and when the number of running background tasks reaches the final number of parallel cores... The time-scheduling process enters a waiting state, and immediately submits the next time-step reorganization task after any background task is completed. The reorganization task is submitted asynchronously in the background using the `SETID` method.

[0019] The triple isolation mechanism consists of three layers: directory isolation, time-step fragmentation isolation, and spatial fragmentation isolation, which eliminates the risk of file lock contention and data overwriting during concurrent writing by multiple processes from both physical and logical perspectives.

[0020] The directory isolation specifically means that each parallel reassembly process only performs read and write operations on the output directory corresponding to its own allocated time step, and the output directories of different processes are completely independent at the file system level and have no overlap.

[0021] Specifically, the time step segmentation and isolation refers to the allocation of all effective simulation time steps to different processes according to the principle of independent processing of each step. Each process processes and only processes all field data of one time step, and the processes do not share any time step read / write handles.

[0022] Specifically, the space fragmentation isolation refers to the fact that when each process reads the processor's block subdirectory, it reads the local data corresponding to the time steps of processor0 to processorN in the order of OpenFOAM's native reorganization logic. All file handles are process-private, and there is no cross-process file handle sharing.

[0023] The upper bound of the acceleration performance of the parallel reassembly compared to the traditional serial reassembly is determined by Amdahl's law, assuming the serial portion accounts for a certain percentage. The number of parallel processes is Then the system's maximum speedup ratio is This method will account for the proportion of the serial portion. Approaching 0, the speedup increases linearly with the number of parallel processes.

[0024] The threshold for the number of training rounds without decreasing is 20 rounds, the decay coefficient is 0.5, the number of training rounds is 200-500 rounds, the batch size is 64-128, the upper threshold is 0.8, the lower threshold is 0.5, the reduction ratio of the number of parallel cores in the load reduction mode is 0.75, the reduction ratio of the number of parallel cores in the conservative mode is 0.5, the task submission interval is 0.1-0.5s, and the interval between adjacent task submissions in the load reduction mode is extended to 0.3-0.5s.

[0025] A second aspect of the present invention provides a computer-readable storage medium storing program instructions that, when executed in a computer, perform the aforementioned OpenFoam data reconstruction method.

[0026] A third aspect of the present invention provides an OpenFoam data reconstruction system comprising the aforementioned computer-readable storage medium, wherein the system is a computer, the computer-readable storage medium is disposed within the system, and the system is provided with a microprocessor for executing program instructions stored in the computer-readable storage medium.

[0027] This invention transforms effective simulation time steps into fully independent parallel tasks at the single-time-step granularity, introduces a deep learning-based adaptive evaluation model for the number of parallel cores to dynamically determine the optimal number of parallel cores, dynamically adjusts the parallelism using a resource safety coefficient function, and achieves conflict-free parallel writing through a dynamic task pool management mechanism and a triple isolation mechanism. This refactors the data reassembly stage from a globally strongly serial segment into a highly parallelized multi-core collaborative workflow, making the serial component in Amdahl's Law approach zero. The system speedup increases linearly with the number of parallel processes, fully releasing the high-concurrency computing power potential of the supercomputing platform. This invention solves the technical problem of ineffective utilization of supercomputing power due to the serial execution mode in the data reassembly stage of OpenFOAM large-scale parallel simulation. Attached Figure Description

[0028] Figure 1 This is a flowchart of the method of the present invention.

[0029] Figure 2 This is a flowchart illustrating the reconstructPar process.

[0030] Figure 3 This is a schematic diagram of a distributed splicing architecture based on time-sliced ​​pre-parsing.

[0031] Figure 4 This is a flowchart illustrating the scheduling logic of the dynamic task pool.

[0032] Figure 5 This is a schematic diagram of the overall architecture of the parallel data reassembly system based on the new generation Sunway supercomputer of this invention.

[0033] Figure 6 This is a comparison chart of actual efficiency for reorganization of grid instances with tens of millions of operations per second. Detailed Implementation

[0034] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below.

[0035] like Figure 1 The diagram shown is a flowchart of an OpenFoam data reconstruction method provided in the first aspect of this invention. This method includes the following steps:

[0036] S01. Scan all processor block subdirectories under the simulation root directory, remove constant directories and cache files, count the total number of effective simulation time steps, extract the minimum and maximum time steps, and complete the adaptive calculation of core parameters.

[0037] S02. Based on the total number of effective simulation time steps, calculate the optimal number of parallel cores through the parallel core number adaptive evaluation model, output the running configuration information, and complete the parameter verification and status output before running.

[0038] S03. Create a dedicated temporary log directory to centrally store the running logs and error messages of each parallel reassembly task, ensuring that the running status of the entire process is traceable.

[0039] S04. Loop through all valid simulation time steps. Based on the dynamic task pool management mechanism, each time step is transformed into an independent scheduling and computing unit. Relying on the triple isolation mechanism, conflict-free parallel writing is achieved. Each process independently completes the reorganization of the grid and flow field data for the corresponding time step and redirects the output to the log directory.

[0040] S05. Monitor the number of concurrent tasks in the dynamic task pool in real time. When the number of concurrent tasks is lower than the optimal number of parallel cores threshold, automatically pull subsequent pending time step tasks to supplement the task pool to keep computing resources running at full capacity until all time steps are reassembled.

[0041] S06. Wait for all background parallel reorganization tasks to complete, release computing resources, output the reorganization completion status log, and end the full-process data reorganization job.

[0042] The processor block subdirectory refers to the subdirectory that is independently generated and stores local partitioned grid data and flow field data by each MPI process during the OpenFOAM parallel simulation process. The naming rule is processor0, processor1 to processorN, where N is the total number of parallel processes minus 1. Each subdirectory is organized into subfolders according to the simulation time step. Each time step subfolder stores the physical field numerical data such as velocity, pressure, and temperature at the corresponding time.

[0043] The effective simulation time step refers to the set of all named subdirectories obtained after scanning the processor0 subdirectory using the directory traversal algorithm, excluding the constant directory and the initial time step directory 0. The total number of effective simulation time steps is denoted as […]. The minimum time step is denoted as The maximum time step is denoted as The above three parameters together constitute the core basis for subsequent parallel task scheduling.

[0044] The parallel core number adaptive evaluation model is based on deep learning, and its specific structure is as follows: the input layer receives a 7-dimensional feature vector, which is composed of the total number of effective simulation time steps. Average data volume per time step (Unit: MB) Total number of processor block subdirectories Available memory capacity of nodes (Unit: GB) Maximum number of parallel cores per node Current bandwidth load of parallel file system (Value range 0-1), average single-core recombination time in historical operations (Unit: s) Structure: The input layer is followed by a batch normalization layer to standardize the 7-dimensional feature vector and eliminate dimensional differences; the batch normalization layer is followed by a first fully connected layer with 256 neurons and a modified linear unit activation function; the first fully connected layer is followed by a first dropout layer with a dropout rate of 0.2–0.3 to suppress overfitting; the first dropout layer is followed by a second fully connected layer with 128 neurons and a modified linear unit activation function; the second fully connected layer is followed by a second dropout layer with a dropout rate of 0.2–0.3; the second dropout layer is followed by a third fully connected layer with 64 neurons and a modified linear unit activation function; the third fully connected layer is followed by an output layer with 1 neuron and a linear activation function, outputting the optimal parallel kernel number prediction value. A residual skip connection is introduced between the first and second fully connected layers. The input of the first fully connected layer is linearly projected and added to the output of the second fully connected layer to alleviate the gradient vanishing problem. Fully connected weight matrices are used for inter-layer feature transfer between neurons within each fully connected layer. Matrix multiplication is used to complete information flow between feature maps of different layers. Neuron dropout operations in the Dropout layer are performed independently during each forward propagation using a random mask. The mean and variance of the batch normalization layer are independently calculated within each training batch. A moving average statistic is used during inference. The model allocates GPU memory blocks in layer order. The weight matrix, bias vector, and intermediate activation values ​​of each layer occupy independent GPU memory partitions. CUDA computation streams are scheduled sequentially according to layer order. Batch normalization layers and fully connected layers share the same CUDA stream to reduce stream switching overhead. Dropout layers perform random mask generation on independent CUDA streams to achieve asynchronous overlap with forward computation. The training dataset establishment steps for the parallel core number adaptive evaluation model are as follows: Historical job records of CFD examples of different scales are collected on the Sunway supercomputing platform. Each record contains the above 7-dimensional input features and the corresponding measured optimal parallel core number label. The optimal parallel core number label is determined by assigning a value of 2 to 10 ... The recombination job is repeatedly executed within a range of different core counts, and the core count corresponding to the shortest total recombination time is used for labeling. The training set has no less than 5000 samples, the validation set has no less than 1000 samples, and the test set has no less than 500 samples. The dataset is divided in a 7:2:1 ratio. The training steps of the parallel core count adaptive evaluation model are as follows: the Adam optimizer is used, and the initial learning rate is set to... The loss function uses mean squared error loss, with 200-500 training epochs and batch sizes of 64-128 per epoch. Learning rate decay is triggered when the validation set loss does not decrease for 20 consecutive epochs, with a decay coefficient of 0.5. Finally, the model weights corresponding to the minimum validation set loss are saved. The adaptive parallel core number evaluation model transforms heterogeneous resource constraints such as node memory capacity, file system bandwidth load, and historical latency—which are difficult to model uniformly using analytical formulas—through nonlinear mapping of multidimensional operating environment characteristics into accurate predictions of the optimal number of parallel cores. This avoids overestimation or underestimation using fixed empirical formulas under different computational scales, thereby maximizing computational resource utilization while ensuring memory safety and reducing the risk of resource idleness or task crashes due to improper core number configuration.

[0045] The adaptive evaluation model for the number of parallel cores outputs the optimal predicted value for the number of parallel cores. Then, through the resource security factor function The predicted value of the optimal parallel core count Make corrections to the resource security factor function. The calculation formula is: ;in and These are the weighting coefficients. This represents the current free memory capacity of the node (in GB). This represents the current bandwidth load rate of the parallel file system. and The values ​​were obtained through multiple rounds of experiments and regression analysis on historical job data under different load conditions on the Sunway supercomputing platform. The experimental range was [missing information]. Ten points were sampled evenly between 0 and 1. Ten points were uniformly sampled between 0.1 and 1, resulting in 100 sets of experimental conditions. Each set of conditions was repeated three times, and the average was taken. The objective function was determined by minimizing the sum of the actual reorganization crash rate and the resource idle rate. and The final value; when At that time, the final number of parallel cores The task pool is running at full capacity; when At that time, the final number of parallel cores The task pool operates in a reduced-load mode and triggers rate-limited submission control, extending the interval between adjacent task submissions to 0.3–0.5 seconds; when At that time, the final number of parallel cores The task pool operates in conservative mode, with a maximum number of concurrent tasks per session not exceeding [a certain threshold]. This triggers a memory warning log output; the three thresholds of 0.8 and 0.5 were obtained through stability experiments on the Sunway supercomputing platform under different node load conditions. The specific method is as follows: Experiments were conducted point-by-point within the range of 0.3 to 1.0, with a step size of 0.05. The operation was repeated 10 times for each value, and the number of task crashes and resource utilization were recorded. The value corresponding to the first time the crash rate dropped to 0 was recorded. The value is used as the lower threshold, and the value corresponding to the peak resource utilization rate is taken. The value is used as the upper threshold.

[0046] The core scheduling logic of the dynamic task pool management mechanism is as follows: the system iterates through the list of effective simulation time steps, counts the number of currently running background tasks in real time, and when the number of running background tasks reaches the final number of parallel cores... When the time step is complete, the scheduling process enters a waiting state until any background task finishes and releases a free slot. Then, it immediately submits the next time step reorganization task, repeating this process until all time step tasks have been submitted. Finally, it exits after all background tasks have finished. The pseudocode implementation of the scheduling logic is as follows: For each time step in the list of valid simulation time steps... Execute a loop to check if the current number of background tasks is less than If the condition is not met, call `wait -n` to wait for any task to finish. Once the condition is met, reset the time step using `SETID`. The corresponding reorganization task is submitted to the cluster scheduling system asynchronously in the background, and waits for 0.1 to 0.5 seconds after submission to avoid scheduling overload caused by high-frequency submissions; the dynamic task pool management mechanism strictly controls the number of instantaneously active tasks to not exceed This ensures that the total memory required by all active tasks is always lower than the node's physical memory limit, thus avoiding the risk of memory overflow in single-process mode from an architectural perspective.

[0047] The triple isolation mechanism consists of three layers: directory isolation, time-step fragmentation isolation, and spatial fragmentation isolation. Directory isolation means that each parallel reassembly process only reads and writes to the output directory corresponding to its assigned time step. The output directories of different processes are completely independent at the file system level and have no overlap. Time-step fragmentation isolation means that all effective simulation time steps are allocated to different processes according to the principle of single-step independent processing. Each process processes and only processes all field data of one time step, and processes do not share any time step read / write handles. Spatial fragmentation isolation means that when each process reads the processor's block subdirectory, it reads the local data of the time steps corresponding to processor0 to processorN in the order of OpenFOAM's native reassembly logic. All file handles are process-private and there is no cross-process file handle sharing. The triple isolation mechanism completely eliminates the risk of file lock contention and data overwriting during multi-process concurrent writing from both physical and logical levels, achieving zero-conflict parallel writing.

[0048] The upper bound of the acceleration performance of the parallel reassembly compared to the traditional serial reassembly is determined by Amdahl's law, assuming the serial portion accounts for a certain percentage. The number of parallel processes is Then the system's maximum speedup ratio is In the traditional reconstructPar serial mode, the data reassembly stage is a globally strong serial segment, with the serial portion accounting for a significant portion. Approaching 1, substituting into Amdahl's law, we get... This means that regardless of the number of parallel processes used in the numerical solution stage, the overall job speedup ratio approaches 1, and the supercomputing power is locked by the serial post-processing. This method reconstructs the data reconstruction stage from a globally strongly serial segment into a completely independent parallel task with a single time step granularity, reducing the proportion of the serial part to a much smaller percentage. Approaching 0, thus making the system speedup increase with the number of parallel processes. The increase in power linearly enhances the computing power of the supercomputing platform, fully unleashing its potential for high-concurrency computing.

[0049] The fully automated parallel recombining workflow is fully compatible with the native reconstructPar call interface. Users do not need to modify their operating habits or configure additional parallel parameters. Internally, it automatically maps serial logic to multi-core parallel tasks, and each process directly outputs the final global complete data without secondary merging loss. The workflow uses a background dynamic analysis mechanism. Once a new valid time step directory is generated during the calculation stage, the system automatically includes the time step in the task pool for staggered recombining, spreading the post-processing to the idle time period of the numerical solution cycle and eliminating the overall batch waiting overhead.

[0050] The specific implementation of step S01 is as follows: The system first recursively scans the simulation root directory, enumerating all subdirectories prefixed with "processor". It then uses a regular expression matching mechanism to filter out the "constant" directory and cached hidden files starting with a dot, constructing a list of valid processor block subdirectories. Subsequently, using "processor0" as the time step reference directory, it performs numerical validity checks on each of its internal subdirectories, removing non-time step directories named "constant" and "0". The remaining numerically named subdirectories are collected into a valid simulation time step set, and the total number of elements in the set is counted. The minimum time step is extracted through sorting operations. With the maximum time step The above three core parameters together constitute the input basis for subsequent parallel task scheduling, completing the core parameter adaptive calculation stage.

[0051] The specific implementation of step S02 is as follows: The system reads the node hardware information, including the maximum number of parallel cores. Available memory capacity of nodes Current free memory capacity The data volume of a single time step in the processor's block subdirectory is calculated by sampling. Query from historical assignment database Obtain the current bandwidth load rate from the file system monitoring interface. The aforementioned 7-dimensional feature vectors are fed into a pre-trained parallel kernel number adaptive evaluation model. This model sequentially passes through a batch normalization layer, a three-layer fully connected network, and a residual skip connection for nonlinear mapping, outputting the optimal parallel kernel number prediction. Then, the resource security factor is calculated. , and Determined by regression analysis of historical experiments. When At that time, the final number of parallel cores The task pool is running at full capacity; when hour, The task pool operates in a reduced-load mode, with the interval between adjacent task submissions extended to 0.3–0.5 seconds; when hour, The task pool operates in conservative mode, with a maximum number of concurrent tasks per session not exceeding [a certain threshold]. This triggers the output of a memory warning log. After the configuration information is output, the pre-run parameter verification and status output are completed.

[0052] The specific implementation of step S03 is as follows: The system creates a dedicated temporary log directory named with a timestamp in the simulation root directory. The directory naming rule uses a string concatenated with year, month, day, hour, minute, and second to ensure that a unique log storage path is generated for each run. Subsequently, the standard output stream and standard error stream of each parallel reassembly process are independently written to a log file named with a time step value in this directory through an output redirection mechanism, realizing centralized archiving of running status and abnormal information, providing a foundation for subsequent fault location and full-process traceability.

[0053] The specific implementation of step S04 is as follows: The system takes the set of effective simulation time steps as input and uses a dynamic task pool management mechanism for task scheduling. For each time step... The system constructs a reorganization command with the given time step as a parameter and submits the command to the operating system asynchronously in the background via the `SETID` mechanism. This allows each reorganization process to run in an independent session, completely decoupled from the main scheduling process. A triple isolation mechanism takes effect simultaneously at this stage: directory isolation ensures that each process only operates on the output directory corresponding to its own time step; time step fragmentation isolation ensures that each process processes only one time step's total data; and spatial fragmentation isolation ensures that all file handles are process-private when reading from processor0 to processorN, with no cross-process sharing. This triple isolation mechanism completely eliminates the risk of file lock contention and data overwriting during concurrent multi-process writes from both physical and logical perspectives, achieving zero-conflict parallel writing. The output of each process is redirected to the log directory created in step S03.

[0054] The specific implementation of step S05 is as follows: After submitting each time step task, the system queries the background process list to count the number of currently active tasks in real time. When the number of active tasks reaches... When the task completion event is triggered, the main scheduling process calls the `wait -n` primitive to enter a blocked waiting state until any background task finishes and releases a free slot. After the task completion event is triggered, the main scheduling process immediately retrieves the next time step from the uncommitted time step queue and repeats the task submission process in step S04, promptly filling any empty computing resource slots to keep the task pool always fully loaded. This polling-replenishment loop continues until all reassembly tasks corresponding to all time steps have been submitted.

[0055] The specific implementation of step S06 is as follows: After all time-step tasks have been submitted, the scheduling master process calls the wait command without parameters to block and wait until all background parallel reassembly processes have finished executing, and the operating system reclaims the corresponding process resources. The system then scans the temporary log directory, counts the number of error message entries in each time-step log file, summarizes the list of successfully reassembled time steps and the list of failed time steps, writes the statistical results to the overall status log file, outputs the reassembly completion flag, and ends the entire data reassembly operation.

[0056] It should be noted that the key technologies of this invention include: a fully independent parallel task construction technology with a single time step granularity, which completely decouples data read and write operations at different time steps at the logical and physical levels, ensuring that there are no data dependencies between parallel processes. This mechanism brings the Amdahl serial proportion of the serial reassembly segment close to 0, achieving a linear increase in speedup with the number of cores; a deep learning parallel core number adaptive evaluation technology based on 7-dimensional heterogeneous resource characteristics, which captures the nonlinear mapping relationship between heterogeneous constraints such as memory, bandwidth, and historical latency through batch normalization and residual networks, combined with a three-level dynamic correction mechanism of the resource safety coefficient function, maximizing computational resource utilization while ensuring memory safety; and a dynamic task pool management technology, which eliminates the core idling problem under static task allocation through an event-driven real-time replenishment mechanism. The synergistic effect of these three technologies ensures that the three objectives of adaptive parallelism determination, conflict-free concurrent writes, and maximized resource utilization are simultaneously guaranteed, raising the overall computational efficiency of the data reassembly stage to the theoretical upper limit of multi-core parallelism on supercomputing platforms.

[0057] It should be noted that in the post-processing stage of large-scale CFD simulations on a supercomputing platform, when the node's memory capacity is limited while the number of effective simulation time steps is extremely large, multiple parallel reassembly processes simultaneously loading large volumes of field data into memory may exhaust the node's physical memory, triggering the operating system to forcibly terminate the processes and causing the entire reassembly task to crash. The reason for this technical problem is that the fixed parallel core count strategy cannot perceive the dynamic changes in the node's real-time memory status and file system bandwidth load. When multiple processes simultaneously load time step data from processor 0 to processor N, the combined memory usage of each process exceeds the node's physical memory limit, triggering the operating system's memory overflow protection mechanism, leading to the forced termination of the processes. The usual solution to this technical problem is to manually pre-set a fixed number of parallel cores and rely on engineering experience to estimate the memory limit. However, fixed empirical values ​​carry a significant risk of overestimation or underestimation under various scenarios, such as changes in the scale of the simulation, dynamic fluctuations in node load, or file system bandwidth congestion. Overestimation leads to memory overflow crashes, while underestimation results in a large number of cores idling. This invention effectively solves this technical problem by using the node's available memory capacity... Current free memory Parallel file system bandwidth load rate Input features, output by deep learning model Then, through the resource security factor function Perform three-level dynamic correction, in In the conservative mode Compress to And limit the maximum number of concurrent tasks in a single session to no more than At the architectural level, the total memory usage of all active tasks is constrained within the safe boundary of the node's physical memory, completely eliminating the risk of memory overflow and crash. At the same time, the parallelism is automatically increased when memory is sufficient, achieving a dynamic balance between safety and efficiency.

[0058] A second aspect of the present invention provides a computer-readable storage medium storing program instructions that, when executed in a computer, perform the aforementioned OpenFoam data reconstruction method.

[0059] A third aspect of the present invention provides an OpenFoam data reconstruction system comprising the aforementioned computer-readable storage medium, wherein the system is any one of a computer, a server, or a microcontroller, the computer-readable storage medium is disposed within the system, and the system is provided with a microprocessor for executing program instructions stored in the computer-readable storage medium.

[0060] Specifically, the principle of this invention is:

[0061] The fundamental reason why this invention can solve the above-mentioned technical problems is that there is a natural logical data independence between effective simulation time steps. That is, there is no dependency between the field quantity data of different time steps at the reading and writing levels. This characteristic provides a theoretical basis for constructing completely independent parallel tasks at the granularity of a single time step. The traditional reconstructPar tool fails to utilize this independence and instead incorporates all time steps into the same serial execution flow, resulting in a serious waste of computing resources.

[0062] This invention first obtains all valid simulation time steps by scanning the processor0 subdirectory, establishing a precise task metadata database to provide complete input for subsequent parallel scheduling. Based on this, a deep learning-based adaptive evaluation model for the number of parallel cores is introduced, taking 7-dimensional heterogeneous resource characteristics as input. This model eliminates dimensional differences through a batch normalization layer and captures the nonlinear relationships between heterogeneous resource constraints such as node memory capacity, file system bandwidth load, and historical latency through a multi-layer fully connected network and residual skip connections. It outputs the optimal predicted value for the number of parallel cores, overcoming the inherent defect of large estimation deviations in fixed empirical formulas under different example sizes. The predicted value is further corrected by a resource safety coefficient function, dynamically switching between full load, reduced load, and conservative operating modes based on the current bandwidth load rate and idle memory ratio, ensuring that parallel tasks run within memory safety boundaries.

[0063] The dynamic task pool management mechanism monitors the number of active background tasks in real time and immediately replenishes new tasks after any task is completed, ensuring that computing resources are always running at full capacity. This avoids the core idle problem caused by uneven task execution time under the static task allocation strategy. The triple isolation mechanism, through directory isolation, time-step fragmentation isolation, and spatial fragmentation isolation, completely eliminates file lock contention and data overwrite risks during concurrent multi-process writes at both the file system and process dimensions, ensuring the correctness and stability of parallel writes. Since each process directly outputs the final global complete data without secondary merging, additional serial bottlenecks are further eliminated. The synergistic effect of these modules makes the serial portion of the data reassembly stage approach zero, and the overall system speedup increases linearly with the number of parallel processes, achieving full utilization of supercomputing power from a mechanism perspective.

[0064] The following provides a specific embodiment 1 of the present invention, and the specific implementation of each step in this embodiment 1 is described in detail below.

[0065] The specific implementation of step S01 is as follows: the system recursively scans all subdirectories prefixed with "processor" under the simulation root directory, removes the constant directory and cache files, then enters the processor0 subdirectory to perform directory traversal, and includes all numerically named subdirectories into the effective time step set. ,in To determine the total number of effective simulation time steps, For the minimum time step, The maximum time step, together with these three factors, constitute the core parameters for subsequent parallel scheduling.

[0066] The specific implementation of step S02 is that the system collects 7-dimensional feature vectors. As input to the parallel core number adaptive evaluation model, where This represents the average data volume per time step, in MB. The total number of subdirectories in the processor block. This refers to the available memory capacity of the node, in GB. This represents the maximum number of parallel cores per node. This represents the current bandwidth load rate of the parallel file system, with a value ranging from 0 to 1. The average single-core reassembly time for historical jobs is given in seconds. The batch normalization layer standardizes the input features. The normalization formula for the 3D feature is as follows:

[0067] ;

[0068] In the formula, For the first The eigenvalues ​​after normalization are dimensional. For the first dimensional original eigenvalues, For the first The mean of the dimensional feature within the training batch. For the first The variance of the dimensional feature within the training batch, To prevent extremely small constants from being divided by zero, the default value is 0. , This is the feature dimension index, with values ​​ranging from 1 to 7, used during the inference stage. and Instead of using the moving average statistic from training, the above normalization operation eliminates the dimensional differences between features, transforming each feature into a dimensionless standardized value before entering subsequent network layers. The normalized features are mapped through the first fully connected layer, and the calculation formula is as follows:

[0069] ;

[0070] In the formula, Here is the weight matrix of the first fully connected layer, with dimension 1. , Let be the bias vector of the first fully connected layer, with dimension . , This is the normalized 7-dimensional feature column vector. To modify the activation function of the linear unit, it is defined as follows: ,in The scalar input to the activation function. This is the output vector of the first fully connected layer, with dimension . The weight matrix of this layer With bias vector All parameters are dimensionless, output It is also a dimensionless intermediate feature vector. The first Dropout layer pairs... By discard rate Perform a random mask operation. The value range is 0.2 to 0.3, and the formula is as follows:

[0071] ;

[0072] In the formula, To and A Bernoulli random mask vector of the same dimension, each element independently with probability Choose 1, with probability Take 0, For element-wise multiplication, This is the output vector of the first Dropout layer, with dimension . , divided by The goal is to maintain a consistent expected output amplitude during both the training and inference phases. The calculation formula for the second fully connected layer is as follows:

[0073] ;

[0074] In the formula, This is the weight matrix of the second fully connected layer, with dimension . , Let be the bias vector of the second fully connected layer, with dimension . , This is the output vector of the second fully connected layer, with dimension . To alleviate the vanishing gradient problem, a residual skip connection is introduced. After linear projection and The formula for addition is as follows:

[0075] ;

[0076] In the formula, The linear projection weight matrix has dimensions of . , Let be a linear projection bias vector with dimension . , The vector is the result of residual fusion, with dimension . The purpose of linear projection is to... The dimensions were aligned from 256 to 128 to make the dimensions of the vectors on both sides of the addition operation consistent. The second Dropout layer is configured according to the dropout rate. right Execution mask, The value range is 0.2 to 0.3. The masking operation is the same as that of the first Dropout layer, resulting in the output vector. , dimension The calculation formula for the third fully connected layer is as follows:

[0077] ;

[0078] In the formula, This is the weight matrix of the third fully connected layer, with dimension . , This is the bias vector of the third fully connected layer, with dimension . , This is the output vector of the third fully connected layer, with dimension . The output layer uses linear activation, as shown in the following formula:

[0079] ;

[0080] In the formula, The output layer weight matrix has dimensions of . , Scalar bias for the output layer This is the optimal parallel core count prediction value, which is a dimensionless core count value, and is related to the label. The units are consistent. The mean squared error loss function is used during the training phase, as shown in the following formula:

[0081] ;

[0082] In the formula, This represents the number of samples per batch, ranging from 64 to 128. For the first The model prediction value for each sample. For the first The measured optimal number of parallel cores for each sample. Let be the mean squared error loss value. Both sides of the equation above are units of the square of the kernel number, ensuring consistent dimensions. Training uses the Adam optimizer with an initial learning rate of . The training rounds are 200–500. When the validation set loss does not decrease for 20 consecutive rounds, the learning rate is updated with a decay factor of 0.5. After prediction is complete, the resource safety factor is... The calculation formula is as follows:

[0083] ;

[0084] In the formula, and These are dimensionless weighting coefficients. This represents the current free memory capacity of the node, in GB. The dimensionless resource safety coefficient and Since both are dimensionless ratios, the dimensions on both sides of the equal sign are consistent. and By testing 100 sets of experimental conditions on the Sunway supercomputing platform ( Ten points were sampled evenly from 0 to 1. The regression analysis was performed using a sample of 10 points uniformly between 0.1 and 1, with each group repeated 3 times and the mean value taken. The objective was to minimize the sum of the actual reorganization crash rate and the resource idle rate. At that time, the final number of parallel cores ;when hour, The interval between adjacent task submissions is extended to 0.3–0.5 seconds; when hour, The maximum number of concurrent tasks in a single session shall not exceed And trigger the output of the memory warning log, in which The floor function truncates consecutive real numbers to the largest integer not exceeding this value, ensuring that the actual number of allocated cores is a positive integer. The three thresholds of 0.8 and 0.5 were obtained through stability experiments on the Sunway supercomputing platform under different node load conditions. Experiments were conducted point-by-point within the range of 0.3 to 1.0, with a step size of 0.05. The operation was repeated 10 times for each value, and the number of task crashes and resource utilization were recorded. The value corresponding to the first time the crash rate dropped to 0 was recorded. The value is used as the lower threshold, and the value corresponding to the peak resource utilization rate is taken. The value is used as the upper threshold.

[0085] The specific implementation of step S03 is that the system creates a temporary log directory in the simulation root directory, and the standard output and error output of each parallel reassembly task are redirected to an independent log file named after the time step in this directory to ensure that the status of the entire process is traceable.

[0086] The specific implementation of step S04 is that the system iterates through the set of effective time steps. By relying on a triple isolation mechanism, each time step is transformed into an independent background task and submitted in the form of a setsid. Each process only reads and writes the directory corresponding to its own assigned time step, and the file handles are all private to the process, completely eliminating the risk of file lock contention and data overwriting.

[0087] The specific implementation of step S05 is that the scheduling process counts the number of background tasks in real time. ,in The number of active background reorganization tasks at the moment. The system enters a waiting state, and immediately submits the next time step task after any task is completed. This cycle continues until all tasks have been submitted. The upper bound of the parallel speedup ratio is determined by Amdahl's Law, as shown in the following formula:

[0088] ;

[0089] In the formula, This represents the percentage of time spent on the serial portion of the overall task. The number of parallel processes. The maximum speedup ratio of the system is given in this method. Approaching 0, Follow The increase is linear, fully unleashing the high-concurrency computing power potential of the supercomputing platform.

[0090] The specific implementation of step S06 is that the system waits for all background parallel reorganization tasks to be completed, releases computing resources, outputs a reorganization completion status log, and ends the entire data reorganization process.

[0091] To better understand and implement this invention, the following is a specific application scenario of this invention, Example 2:

[0092] This embodiment uses a typical large-scale CFD example on the Sunway supercomputing platform to verify the complete execution process of the OpenFOAM data reconstruction method proposed in this invention. The example uses an unstructured mesh with a total of 1.2 × 10⁻⁶ cells. Total number of processor block subdirectories The total number of effective simulation time steps is 512. The minimum time step is 1200. The maximum time step is 0.001. The average data volume per time step is 1.200. It is 3.8 Each time step stores physical field numerical data such as velocity, pressure, and temperature.

[0093] In step S01, the system performs a full scan of the simulation root directory, enumerating 512 processor block subdirectories from processor0 to processor511. Using processor0 as the baseline, time steps are extracted. After removing the constant and 0 directories, a set of 1200 valid time steps is obtained. , , The core parameters were adaptively calculated, as shown in Table 1.

[0094] Table 1 Summary of Core Parameters

[0095]

[0096] In step S02, the system reads the node hardware information, including the maximum number of parallel cores on the node. The available memory capacity of the node is 64. 512 Current free memory 420 Historical average single-core recombination time The current parallel file system bandwidth load rate is 18 seconds. The value is 0.31. The 7-dimensional feature vector is fed into the pre-trained parallel kernel number adaptive evaluation model, and the model output is... Calculate the resource security factor. Determined through historical regression , Calculated This falls within the load reduction mode range, and the final number of parallel cores... The interval between adjacent task submissions is set to 0.4s, and the runtime configuration information output is shown in Table 2.

[0097] Table 2. Operation Configuration Parameters

[0098]

[0099] In step S03, the system creates a temporary log directory in the simulation root directory, named recon_log_20240918_143022. Subsequently, the standard output stream and standard error stream of all parallel reassembly processes are redirected to independent log files named with time step values ​​in this directory, realizing centralized archiving of the entire process running status.

[0100] In steps S04 and S05, such as Figure 4 As shown, the dynamic task pool management mechanism operates according to the following logic: The system sequentially retrieves time steps from a queue of 1200 valid time steps, submits a reorganization command asynchronously in the background using the `SETID` method, with a 0.4-second interval after each task submission. When the number of active background tasks reaches 39, the scheduling process calls `wait -n` to block and wait. New tasks are immediately added after any task finishes, keeping the task pool always fully loaded. A triple isolation mechanism is activated simultaneously; each process only operates on the output directory corresponding to its own time step, all file handles are process-private, there is no cross-process sharing, and conflict-free parallel writing is achieved. Figure 3 As shown, the distributed stitching architecture based on time-sliced ​​pre-parsing distributes all time steps in parallel to each computing process. Each process sequentially reads local data from processor0 to processor511 and directly outputs the complete global data without secondary merging loss. Figure 5 As shown, the overall architecture of the parallel data reassembly system based on the Sunway supercomputer of this invention consists of four layers: a parameter scanning layer, an intelligent scheduling layer, a triple-isolation execution layer, and a log archiving layer. Information is exchanged between these layers through standardized interfaces. In the task supplementation stage of step S05, the system uses a background-resident dynamic analysis mechanism to detect the newly generated valid time step catalog during the computation stage. The newly added time steps are automatically included in the task pool for staggered reassembly, distributing the post-processing overhead evenly across the idle periods of the numerical solution cycle.

[0101] In step S06, after all 1200 time-step reassembly tasks have been submitted, the scheduling master process calls the wait command to block and wait until all 39 active processes have finished executing, at which point the operating system reclaims process resources. The system scans the recon_log_20240918_143022 directory, counts error entries in the logs of each time step, confirms that all 1200 time steps have been successfully reassembled without error entries, outputs the overall status log, and ends the entire data reassembly operation. The reassembly results are summarized in Table 3.

[0102] Table 3 Summary of Reorganization Completion Status

[0103]

[0104] like Figure 2 As shown, the traditional reconstructPar tool uses a global serial process, processing 1200 time steps sequentially. The reconstructing time is linearly proportional to the number of time steps, leaving many computing cores on the supercomputing node idle during this stage. Figure 6 As shown, under the same computational scale, this invention reconstructs serial reassembly segments into completely independent parallel tasks with single-time-step granularity. The proportion of the serial portion in Amdahl's law approaches zero, and the system speedup increases linearly with the number of parallel processes, fully releasing the high-concurrency computing power potential of the supercomputing platform. Figure 6 In this designation, "sw parallel" indicates that the Sunway architecture uses multiple main cores in a parallel manner, "sw main core" indicates that it only uses a single main core on the Sunway supercomputing node, and "x86" represents the traditional x86 computing platform. Traditional serial methods, when processing large-scale time steps, suffer from a global serial segment that locks the overall speedup at a level close to 1. This invention, however, fundamentally eliminates the risk of conflicts from concurrent writes by multiple processes through time-step-level task decoupling and triple isolation mechanisms. This allows each parallel process to complete the reassembly operation truly independently and synchronously. The overall system throughput increases synchronously with the increase in the number of cores, eliminating the need to rely on serial post-processing waiting and completely changing the unbalanced pattern of parallel computation and serial post-processing in traditional methods.

[0105] It should be noted that the variables involved in this invention are explained in detail in Table 4.

[0106] Table 4. Variable Explanation Table

[0107]

[0108] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. An OpenFOAM data reconstruction method, characterized in that, Includes the following steps: Scan all processor block subdirectories under the simulation root directory, remove constant directories and cache files, count the total number of effective simulation time steps, extract the minimum and maximum time steps, and complete the adaptive calculation of core parameters. Based on the total number of effective simulation time steps, the optimal number of parallel cores is calculated through the parallel core number adaptive evaluation model, and the running configuration information is output to complete the parameter verification and status output before running. Create a temporary log directory to centrally store the running logs and error messages of each parallel reassembly task; The system iterates through all valid simulation time steps, transforms each time step into an independent scheduling and computing unit based on a dynamic task pool management mechanism, and achieves conflict-free parallel writing by relying on a triple isolation mechanism. Each process independently completes the reorganization of the grid and flow field data for the corresponding time step. The system monitors the number of concurrent tasks in the dynamic task pool in real time. When the number of concurrent tasks is lower than the optimal number of parallel cores threshold, it automatically pulls subsequent pending time step tasks to supplement the task pool. Wait for all background parallel reassembly tasks to complete, release computing resources, and output a reassembly completion status log. The core scheduling logic of the dynamic task pool management mechanism is as follows: the system iterates through the list of effective simulation time steps, counts the number of currently running background tasks in real time, and when the number of running background tasks reaches the final number of parallel cores, the scheduling process enters a waiting state. After any background task is completed, the next time step reorganization task is submitted immediately, and the reorganization task is submitted asynchronously in the background using the setsid method. The triple isolation mechanism consists of three layers: directory isolation, time-step fragmentation isolation, and spatial fragmentation isolation, which eliminates the risk of file lock contention and data overwriting during concurrent writing by multiple processes from both physical and logical perspectives.

2. The OpenFOAM data reconstruction method according to claim 1, characterized in that, The directory isolation refers to the fact that each parallel reassembly process only performs read and write operations on the output directory corresponding to its own allocated time step, and the output directories of different processes are completely independent at the file system level and have no overlap.

3. The OpenFOAM data reconstruction method according to claim 2, characterized in that, The adaptive evaluation model for the number of parallel cores is based on deep learning. The input layer receives a 7-dimensional feature vector, which consists of the total number of effective simulation time steps, the average data volume per time step, the total number of processor block subdirectories, the available memory capacity of the node, the maximum number of parallel cores of the node, the current bandwidth load rate of the parallel file system, and the average single-core reassembly time of historical jobs.

4. The OpenFOAM data reconstruction method according to claim 3, characterized in that, The structure of the parallel core number adaptive evaluation model is as follows: after the input layer, a batch normalization layer is connected, and after the batch normalization layer, a first fully connected layer, a first dropout layer, a second fully connected layer, a second dropout layer, a third fully connected layer, and an output layer are connected in sequence to output the optimal parallel core number prediction value, and a residual jump connection is introduced between the first fully connected layer and the second fully connected layer.

5. The OpenFOAM data reconstruction method according to claim 4, characterized in that, The training of the parallel core number adaptive evaluation model uses the Adam optimizer and the loss function is the mean squared error loss. When the validation set loss continuously reaches the non-decreasing epoch threshold, the learning rate decay is triggered, and the decay coefficient is the decay coefficient value. The model weight corresponding to the minimum validation set loss is saved.

6. The OpenFOAM data reconstruction method according to claim 5, characterized in that, The time step segmentation and isolation specifically refers to the allocation of all effective simulation time steps to different processes according to the principle of independent processing of each step. Each process processes and only processes all field data of one time step, and the processes do not share any time step read / write handles.

7. The OpenFOAM data reconstruction method according to claim 6, characterized in that, The optimal parallel core count prediction is corrected by a resource safety factor function, which divides the task pool into three operating modes based on an upper and lower threshold: full load mode, load reduction mode, and conservative mode. The three operating modes correspond to different final parallel core count calculation rules and task submission interval control strategies.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program instructions that, when executed in a computer, perform the method described in any one of claims 1-7.

9. An OpenFoam data reconstruction system, characterized in that, The system comprises the computer-readable storage medium of claim 8, wherein the system is a computer, the computer-readable storage medium is disposed within the system, and the system is provided with a microprocessor that executes program instructions stored in the computer-readable storage medium.

Citation Information

Patent Citations

  • GPU (Graphics Processing Unit) calculation method for computational fluid mechanics simulation

    CN119271366A

  • Multi-thread parallel simulation and task scheduling optimization method and system for thermal power plant controller

    CN121411890A