Simulation task resource allocation method and device, equipment and medium
Through global mesh division and load prediction model, the electromagnetic full-wave simulation tasks are statically allocated to CPU/GPU, which solves the problems of insufficient fine-grained task division and large dynamic scheduling overhead, improves computing efficiency and resource utilization, and reduces system complexity and synchronization delay.
Patent Information
- Application Number
- CN202510445566.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, the fine-grained task division of electromagnetic full-wave simulation tasks is insufficient and the dynamic scheduling overhead is large, resulting in low resource utilization and reduced system efficiency.
Through the global meshing and load prediction model, the key information of the mesh cells is extracted, load indicators are generated, divided into multiple sub-regions, and the sub-tasks are statically allocated to the most suitable computing unit to avoid the additional overhead caused by dynamic scheduling.
It improves computing efficiency and resource utilization, reduces system complexity and synchronization delay, and ensures simulation accuracy.
Smart Images

Figure CN120276862A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of electromagnetic wave simulation technology, and particularly to a method, device, equipment and medium for allocating simulation task resources. Background Art
[0002] With the wide application of electromagnetic full-wave simulation technology in fields such as optics, radio frequency, and microwave, the traditional computing mode based on a single CPU has been difficult to meet the stringent requirements of large-scale and complex device simulation for computing speed and resource utilization rate. In recent years, the GPU has been introduced into simulation computing due to its powerful large-scale parallel computing ability, significantly accelerating the computing efficiency of high-density grid regions. However, the CPU is still an ideal choice for processing complex data interaction and task scheduling due to its inherent advantages in control logic and data management. In this context, heterogeneous computing architectures (such as CPU+GPU hybrid systems) have emerged, aiming to give full play to the collaborative advantages of the two processors. However, how to efficiently divide large-scale simulation tasks and reasonably allocate them to the most suitable processing units has become a key technical problem to be solved urgently.
[0003] In related technologies, the current mainstream solutions mainly focus on multi-node or distributed scheduling systems, improving the overall efficiency through macro task allocation at the cluster level. However, such methods lack fine-grained partitioning within a single simulation task and cannot perform precise resource allocation based on pre-known parameters such as local grid density and material properties, resulting in insufficient or unbalanced utilization of computing resources. In addition, although dynamic load balancing strategies alleviate the problem of uneven load through real-time monitoring and task migration, in electromagnetic full-wave simulation, since the task load can be predicted in advance by parameters such as grid type, size, and material properties, real-time scheduling reduces the overall system efficiency due to excessive communication and synchronization overhead, which urgently needs to be solved. Summary of the Invention
[0004] This application provides a method, device, equipment and medium for allocating simulation task resources to solve the problems of insufficient fine-grained task partitioning and large dynamic scheduling overhead in the prior art, improve the overall computing efficiency and resource utilization rate, and reduce the system complexity and synchronization delay on the premise of ensuring simulation accuracy.
[0005] An embodiment of the first aspect of the present application provides a simulation task resource allocation method, including the following steps: obtaining a current simulation task; performing global grid division on the current simulation task to obtain a plurality of grid cells, and extracting key information of each grid cell; inputting the key information of each grid cell into a target load prediction model, the target load prediction model outputs corresponding load metrics, and generates global load distribution data according to the load metrics of all grid cells; dividing the global simulation area into a plurality of sub-regions according to the global load distribution data, and obtaining load prediction data of each sub-region; dividing the simulation task into a plurality of sub-tasks according to the load prediction data of each sub-region, and allocating each sub-task to a target computing unit.
[0006] According to an embodiment of the present application, the training method of the target load prediction model includes: obtaining historical simulation data; training an initial load prediction model based on the historical simulation data, and after the training is completed, validating the initial load prediction model based on a preset validation method to obtain a validation result. If the validation result does not meet the preset validation conditions, adjusting the parameters of the initial load prediction model based on a preset parameter adjustment strategy until the validation result meets the preset validation conditions, obtaining a validated load prediction model, and using the validated load prediction model as the target load prediction model.
[0007] According to an embodiment of the present application, the step of dividing the simulation task into a plurality of sub-tasks according to the load prediction data of each sub-region and allocating each sub-task to a target computing unit includes: allocating the first type of sub-task to a graphics processing unit, and allocating the second type of sub-task, the third type of sub-task, the fourth type of sub-task, the fifth type of sub-task, and the sixth type of sub-task to a central processing unit according to the load prediction data of each sub-region; where the first type of sub-task is a preset computationally intensive task, the second type of sub-task is a preset control logic task, the third type of sub-task is a preset boundary condition processing task, the fourth type of sub-task is a preset excitation source injection task, the fifth type of sub-task is a preset data acquisition and storage task, and the sixth type of sub-task is a preset time stepping loop control task.
[0008] According to an embodiment of the present application, after allocating the sub-tasks of each sub-region to the target computing unit, it further includes: asynchronously executing the sub-tasks of each sub-region based on the target computing unit to obtain local calculation results of each sub-region; generating global simulation data according to the local calculation results of each sub-region.
[0009] According to an embodiment of the present application, after dividing the global simulation area into multiple sub-areas according to the global load distribution data, it further includes: generating a grid file for each sub-area, where the grid file for each sub-area includes at least one of grid size, position information, and material properties; generating a load prediction report according to the grid file for each sub-area, where the load prediction report includes the load prediction data for each sub-area.
[0010] According to an embodiment of the present application, the load metrics of each classified grid cell include at least one of computational load, memory occupancy, and communication overhead, and the key information of each grid cell includes: the size, position, and corresponding material classification of each grid cell.
[0011] An embodiment of the second aspect of the present application provides a simulation task resource allocation device, including: an acquisition module for acquiring a current simulation task; a division module for performing global grid division on the current simulation task to obtain multiple grid cells and extracting the key information of each grid cell; a generation module for inputting the key information of each grid cell into a target load prediction model, the target load prediction model outputting corresponding load metrics and generating global load distribution data according to the load metrics of all grid cells; a processing module for dividing the global simulation area into multiple sub-areas according to the global load distribution data and obtaining the load prediction data of each sub-area; and an allocation module for dividing the simulation task into multiple sub-tasks according to the load prediction data of each sub-area and allocating each sub-task to a target computing unit.
[0012] According to an embodiment of the present application, the generation module is further configured to: acquire historical simulation data; train an initial load prediction model based on the historical simulation data, and after the training is completed, verify the initial load prediction model based on a preset verification method to obtain a verification result. If the verification result does not meet the preset verification conditions, adjust the parameters of the initial load prediction model based on a preset parameter adjustment strategy until the verification result meets the preset verification conditions, obtain a verified load prediction model, and use the verified load prediction model as the target load prediction model.
[0013] According to an embodiment of the present application, the allocation module is configured to: allocate the first type of subtasks to the graphics processing unit according to the load prediction data of each sub-region, and allocate the second type of subtasks, the third type of subtasks, the fourth type of subtasks, the fifth type of subtasks, and the sixth type of subtasks to the central processing unit; wherein, the first type of subtasks is a preset compute-intensive task, the second type of subtasks is a preset control logic task, the third type of subtasks is a preset boundary condition processing task, the fourth type of subtasks is a preset excitation source injection task, the fifth type of subtasks is a preset data acquisition and storage task, and the sixth type of subtasks is a preset time stepping loop control task.
[0014] According to an embodiment of the present application, after allocating the subtasks of each sub-region to the target computing unit, the allocation module is further configured to: asynchronously execute the subtasks of each sub-region based on the target computing unit to obtain the local calculation results of each sub-region; generate global simulation data according to the local calculation results of each sub-region.
[0015] According to an embodiment of the present application, after dividing the global simulation region into multiple sub-regions according to the global load distribution data, the processing module is further configured to: generate a grid file for each sub-region, where the grid file for each sub-region includes at least one of grid size, position information, and material properties; generate a load prediction report according to the grid file for each sub-region, where the load prediction report includes the load prediction data for each sub-region.
[0016] According to an embodiment of the present application, the load metrics of each classified grid cell include at least one of compute load, memory occupancy, and communication overhead, and the key information of each grid cell includes the size, position, and corresponding material classification of each grid cell.
[0017] An embodiment of the third aspect of the present application provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the program to implement the simulation task resource allocation method as described in the above embodiment.
[0018] An embodiment of the fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and the program is executed by a processor to be used to implement the simulation task resource allocation method as described in the above embodiment.
[0019] An embodiment of the fifth aspect of the present application provides a computer program product, including a computer program or instruction, and when the computer program or instruction is executed, it is used to implement the simulation task resource allocation method as described in the above embodiment.
[0020] Therefore, the present application has at least the following beneficial effects:
[0021] The current simulation task performs global mesh division to obtain multiple mesh cells, extracts the key information of each mesh cell, outputs the load index corresponding to each mesh cell based on the target load prediction model, generates global load distribution data according to the load indexes of all mesh cells, divides the global simulation area into multiple sub-regions, divides the simulation task into multiple sub-tasks according to the load prediction data of each sub-region, and assigns each sub-task to the target computing unit. Thus, the problems of insufficient fine-grained task division and large dynamic scheduling overhead in the prior art are solved, the overall computing efficiency and resource utilization rate are improved, and the system complexity and synchronization delay are reduced on the premise of ensuring the simulation accuracy.
[0022] The additional aspects and advantages of the present application will be partly given in the following description, partly will become obvious from the following description, or will be understood through the practice of the present application. Description of the Drawings
[0023] The above and / or additional aspects and advantages of the present application will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0024] Figure 1 is a flowchart of a simulation task resource allocation method according to an embodiment of the present application;
[0025] Figure 2 is a schematic diagram of the completion of mesh division and classification within a rectangular simulation area according to an embodiment of the present application;
[0026] Figure 3 is a flowchart of the training of a load prediction model according to an embodiment of the present application;
[0027] Figure 4 is a schematic diagram of the overall FDTD simulation process according to an embodiment of the present application;
[0028] Figure 5 is a schematic diagram of heterogeneous cooperative execution according to an embodiment of the present application;
[0029] Figure 6 is a block schematic diagram of a simulation task resource allocation device according to an embodiment of the present application;
[0030] Figure 7 is a schematic diagram of the structure of an electronic device according to an embodiment of the present application. Detailed Embodiments
[0031] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present application, and should not be construed as limiting the present application.
[0032] A simulation task resource allocation method, apparatus, device, and medium according to an embodiment of the present application will be described below with reference to the accompanying drawings.
[0033] Before introducing the simulation task resource allocation method proposed in the embodiments of the present application, the task allocation method in the related art will be introduced first.
[0034] When performing complex electromagnetic full-wave simulations (such as radio frequency devices, optoelectronic devices, etc.), large-scale grid discretization and a large number of numerical calculations are usually required. Traditional methods either only use the CPU, resulting in a slow calculation speed; or only use the GPU, resulting in an increased burden on control management and I / O operations. Most existing parallel methods focus on task scheduling at the cluster level, but the solution for internal partitioning of a single simulation task to fully utilize the hardware advantages of the CPU / GPU is still not perfect.
[0035] In the related art, some current high-performance computing (HPC) platforms have proposed scheduling systems at the multi-node or distributed level. For example, a cluster composed of multiple CPU (Central Processing Unit) / GPU (Graphics Processing Unit) nodes can handle different simulation tasks simultaneously, but it mainly focuses on the distribution of tasks among different nodes and resource scheduling. Since these solutions mainly focus on the macro scheduling of the entire cluster and lack research on fine-grained partitioning within a single simulation task, it is difficult to perform accurate task partitioning based on pre-known parameters such as local grid density and material properties in a complex electromagnetic full-wave simulation task, resulting in possible problems of insufficient or unbalanced utilization of local computing resources.
[0036] In addition, the dynamic load balancing scheduling strategy focuses on real-time monitoring of the load status of each processing unit during the simulation process and reallocating tasks through the dynamic load balancing algorithm. Such solutions generally include real-time monitoring, scheduling algorithms, task migration, and data synchronization, and can reallocate tasks and synchronize the boundaries of relevant data without interrupting the simulation process. Although the dynamic scheduling strategy can better alleviate the problem of uneven load in some scenarios, in the full-wave simulation scenario with strong task predictability, it is often difficult to achieve the optimal utilization effect of resources due to excessive scheduling overhead and data synchronization problems. Specifically, in electromagnetic full-wave simulation, since the task load can usually be predicted through pre-set parameters (such as mesh type, size, material properties, etc.), real-time dynamic scheduling not only increases the system complexity but also introduces additional communication and synchronization overhead, resulting in increased calculation latency and decreased overall efficiency.
[0037] It can be seen that the defects of the existing technology are mainly reflected in insufficient fine-grained task division, failure to accurately divide the inside of a single simulation task, resulting in low resource utilization; large dynamic scheduling overhead, and the communication and synchronization delays introduced by real-time load monitoring and task migration, reducing the overall operating efficiency of the system.
[0038] Therefore, the simulation task resource allocation method of the present application aims to improve the overall efficiency and scalability of electromagnetic full-wave simulation, while taking into account computational accuracy and resource utilization. Use a load prediction model to evaluate the resource requirements of simulation tasks in advance, so as to achieve precise and fine-grained division inside a single simulation task; adopt a static resource allocation strategy, and based on the prediction results, pre-allocate computationally intensive tasks to the GPU, while allocating control logic and data management tasks to the CPU, thus avoiding the additional overhead brought by real-time scheduling in dynamic load balancing; improve the overall computational efficiency and resource utilization, reduce system complexity and synchronization delays on the premise of ensuring simulation accuracy, and provide an efficient and stable solution for large-scale electromagnetic full-wave simulation.
[0039] Specifically, Figure 1 is a schematic flow chart of a simulation task resource allocation method provided by an embodiment of the present application.
[0040] As Figure 1 shown, the simulation task resource allocation method includes the following steps:
[0041] In step S101, obtain the current simulation task.
[0042] Among them, the current simulation task data includes device geometry information, material properties, and other initial parameters.
[0043] Specifically, the embodiment of the present application can obtain the current simulation task data through a graphical user interface or a command line interface, which is not specifically limited here.
[0044] In step S102, a global mesh division is performed on the current simulation task to obtain a plurality of mesh units, and key information of each mesh unit is extracted.
[0045] Among them, the preset classification strategy can classify the meshes according to a predefined enumeration type.
[0046] Specifically, the embodiment of the present application can adopt the FDTD (Finite-Difference Time-Domain) mesh division algorithm to divide the entire simulation area into a plurality of mesh units. The basis for mesh division is the simulation frequency band, device geometry information, and material properties in the current simulation task. Specifically, the FDTD mesh division algorithm discretizes the simulation area into regular mesh units according to the frequency band of the simulation task, the geometric shape of the device (such as size, structure, etc.), and the material distribution (such as the type of different materials, dielectric constant, magnetic permeability, etc.).
[0047] Furthermore, the mesh unit information is extracted. Specifically, for each mesh unit, key parameters (such as mesh size, position information, material properties, etc.) are extracted, and the meshes are classified according to a predefined enumeration type (MaterialType). Among them, the MaterialType enumeration type includes: Normal: ordinary material, Dispersion: dispersion model, Conformal: conformal, Dis_conformal: dispersion conformal, PML: boundary material, Anisotropic: anisotropic material.
[0048] Exemplarily, as Figure 2 shown, Figure 2 is an example of mesh division within a rectangular simulation area and the completion of classification of mesh units according to a predefined enumeration type.
[0049] In step S103, the key information of each mesh unit is input into the target load prediction model, and the target load prediction model outputs the corresponding load metrics, and global load distribution data is generated according to the load metrics of all mesh units.
[0050] Among them, in some embodiments, the load metrics corresponding to each mesh unit include at least one of computational load, memory occupancy, and communication overhead. The key information of each mesh unit includes: the size, position, and corresponding material classification of each mesh unit.
[0051] Specifically, for each grid cell, key load metrics are calculated based on its material type and grid characteristics, mainly including: Computational load: Evaluate the computational complexity required for this grid cell during the FDTD iteration, considering that different materials (such as dispersive materials or anisotropic materials) may require more complex numerical calculations; Memory occupancy: Estimate the memory resources required for this grid during the simulation according to the grid data structure. Communication overhead: For grids involving boundary data exchange, calculate the communication resources that may be consumed during cross-grid data synchronization.
[0052] Further, in some embodiments, the training method of the target load prediction model includes: obtaining historical simulation data; training the initial load prediction model based on the historical simulation data, and after the training is completed, validating the initial load prediction model based on a preset validation method to obtain a validation result. If the validation result does not meet the preset validation conditions, adjust the parameters of the initial load prediction model based on a preset parameter adjustment strategy until the validation result meets the preset validation conditions, obtaining a validated load prediction model, and using the validated load prediction model as the target load prediction model.
[0053] Specifically, the embodiments of the present application can integrate the load metric parameters (computational load, memory occupancy, communication overhead) of all grid cells to form a feature set that comprehensively describes the entire simulation task, and perform preprocessing such as normalization and denoising on the original data to ensure the stability and accuracy of the data when constructing a mathematical model later.
[0054] Further, based on the preprocessed feature data, construct a mathematical model (such as a LinearWeighted Model, a linear weighted model), that is, an initial load prediction model, to map the grid cell features to the estimated load value.
[0055] For each material type, adjust the weight coefficients of the computational load COMPUTE_WEIGHT, memory occupancy MEMORY_WEIGHT, and communication overhead COMM_WEIGHT so that the initial load prediction model can accurately reflect the resource consumption of various grids during the FDTD iteration.
[0056] Further, use the historical simulation data (the actual load conditions and computing power utilization rates under various simulation scenarios and resource configurations) to train the initial load prediction model, evaluate the prediction accuracy through methods such as cross-validation, and adjust the weight parameters to ensure that the load prediction value output by the initial load prediction model matches the actual simulation load, so as to ensure that the target load prediction model can accurately evaluate the computational load of each grid cell during the FDTD iteration and provide a quantitative basis for static resource allocation.
[0057] In addition, once the initial load prediction model is trained and verified to be accurate, the trained model is directly called during each subsequent simulation process without the need for retraining.
[0058] To facilitate a clearer and more intuitive understanding of the construction process of the target load prediction model in the embodiments of the present application by those skilled in the art, the following is a detailed description in conjunction with Figure 3 for detailed illustration.
[0059] Specifically, as Figure 3 shown, the construction process of the target load prediction model includes the following steps:
[0060] Start stage: Input simulation task data, including device geometry, material properties, and initial parameters.
[0061] Data preprocessing: Normalize, denoise, and convert the format of the original data to ensure data consistency and quality. This step is completed when initially constructing the feature dataset. For similar input conditions, the preprocessing step can be reused or only slightly adjusted, rather than starting from scratch each time.
[0062] Feature extraction: Extract key features such as grid size, position, and material information from the preprocessed data.
[0063] Construct load feature set: Calculate the load metrics for each grid based on the extracted key features, such as computing load, memory occupancy, and communication overhead.
[0064] Model selection / construction: Select a suitable load prediction model (such as a linear weighted model or a machine learning model).
[0065] Model training: Use historical data and cross-validation methods to train the model and continuously adjust the parameters to improve the prediction accuracy.
[0066] Model verification: Compare the model prediction results with the actual load to evaluate the model performance.
[0067] Model optimization: Adjust the model parameters or structure according to the verification results and retrain if necessary.
[0068] Model deployment: Save the trained model for subsequent static resource allocation and task scheduling.
[0069] End: The process ends.
[0070] Optionally, the load prediction model constructed in the embodiments of the present application can be a linear weighted model or a prediction model based on machine learning.
[0071] Among them, the linear weighted model is based on the preprocessed feature data and constructs a linear weighted model by setting fixed weights (computing load COMPUTE_WEIGHT, memory occupancy MEMORY_WEIGHT, communication overhead COMM_WEIGHT), directly mapping the grid features to the load prediction values. The prediction model based on machine learning uses more complex machine learning algorithms (such as regression trees, neural networks, etc.) to fit the grid data, and the prediction accuracy may be higher, but more training data and computing resources are required at the same time.
[0072] Both of the above two schemes can achieve load prediction in terms of basic functions. However, the linear model scheme is simple to implement and has low computing overhead, while the machine learning scheme may perform better in complex scenarios, but it introduces additional model training and tuning complexities.
[0073] Therefore, according to parameters such as the type, computing intensity, grid type and scale of the simulation task, combined with historical data, a mathematical model is established to accurately predict the resource requirements of each future subtask. This model can evaluate the computing complexity of each subdomain in advance and provide a quantitative basis for subsequent resource allocation.
[0074] In step S104, the global simulation area is divided into multiple sub-areas according to the global load distribution data, and the load prediction data of each sub-area is obtained.
[0075] Specifically, based on the calculation results of the target load prediction model, the global simulation area is divided into multiple sub-areas. The target load prediction model predicts the computing load (COMPUTE_WEIGHT), memory occupancy (MEMORY_WEIGHT) and communication overhead (COMM_WEIGHT) of each grid cell to generate the load distribution data of the global simulation area. According to these load distribution data, combined with the number of processes specified by the user, a preset load balancing strategy is used to divide the global simulation area into multiple load-balanced sub-areas. The goal of the division is to make the total load in each sub-area as balanced as possible. This fine division is the premise for realizing efficient static task allocation, ensuring the balance of the computing load within each sub-area, and thus ensuring that in subsequent parallel computing, each computing unit (such as GPU or CPU) can efficiently process the tasks assigned to them.
[0076] Optionally, the embodiments of the present application can adopt adaptive division based on the load balancing goal, or adopt uniform division under a fixed number of processes.
[0077] Among them, the adaptive division based on the load balancing goal uses the region growing or divide-and-conquer strategy according to the load prediction report to divide the global simulation area into multiple load-balanced sub-areas. The uniform division under a fixed number of processes directly performs uniform division according to the preset number of processes and then makes local adjustments to the division results later.
[0078] It can be seen that adaptive partitioning can more accurately reflect the load distribution and ensure load balance within each sub-domain; while uniform partitioning is simple and easy to implement, but may lead to a decrease in resource utilization in the case of uneven load. Those skilled in the art can adopt different partitioning methods according to the actual situation to divide the global simulation area into multiple load-balanced sub-areas, which will not be specifically limited here.
[0079] Further, in some embodiments, after dividing the global simulation area into multiple sub-areas according to the global load distribution data, it further includes: generating a grid file for each sub-area, where the grid file for each sub-area includes at least one of grid size, position information, and material properties; generating a load prediction report according to the grid file of each sub-area, where the load prediction report includes load prediction data for each sub-area.
[0080] Specifically, for each sub-area, a respective grid file is generated, and these grid files contain detailed information about all the grids within the sub-area, including the size of the grids, position information, and material properties (based on the MaterialType enumeration type, such as Normal, Dispersion, Conformal, etc.), and finally a load prediction report is generated, containing load prediction data for each sub-area (such as computational load, memory occupancy, and communication overhead).
[0081] In step S105, the simulation task is divided into multiple subtasks according to the load prediction data of each sub-area, and each subtask is assigned to a target computing unit.
[0082] Further, in some embodiments, dividing the simulation task into multiple subtasks according to the load prediction data of each sub-area and assigning each subtask to a target computing unit includes: according to the load prediction data of each sub-area, assigning the first type of subtask to a graphics processing unit, and assigning the second type of subtask, the third type of subtask, the fourth type of subtask, the fifth type of subtask, and the sixth type of subtask to a central processing unit; where the first type of subtask is a preset computationally intensive task, the second type of subtask is a preset control logic task, the third type of subtask is a preset boundary condition processing task, the fourth type of subtask is a preset excitation source injection task, the fifth type of subtask is a preset data acquisition and storage task, and the sixth type of subtask is a preset time stepping loop control task.
[0083] Optionally, the resource allocation strategy of the embodiments of the present application can be a pure static allocation scheme, or a semi-static / semi-dynamic scheme.
[0084] Among them, in the pure static allocation scheme, before the simulation starts, each subtask is directly and statically allocated to the GPU or CPU according to the load prediction result without runtime adjustment. The semi-static / semi-dynamic scheme mainly uses static allocation, but allows local and limited dynamic adjustment under certain boundary or abnormal conditions to cope with prediction errors or load fluctuations.
[0085] It can be seen that the pure static scheme avoids the communication and synchronization overheads brought by dynamic scheduling, and the system structure is simple and efficient; while the semi-static scheme has certain advantages in flexibility, but at the same time introduces additional scheduling complexity and real-time adjustment costs. Those skilled in the art can adopt different resource allocation strategies according to the actual situation, which are not specifically limited here.
[0086] Specifically, according to the load prediction result, a static allocation algorithm is developed to reasonably pre-allocate subtasks to the GPU (such as computationally intensive subtasks) and the CPU (such as control, boundary handling, and data management subtasks). For example, the detailed allocation strategy is as follows:
[0087] (1) Initialization stage
[0088] Subtasks: Initialize the calculation area, model, material parameters, time step, boundary condition parameters, excitation source, recorder, etc.; perform sub-region division and mesh generation.
[0089] Since the initialization stage involves complex logical judgments, data parsing, and configuration management, it needs to be executed serially and depends on the general processing ability of the CPU. Therefore, the execution unit is the CPU.
[0090] (2) Field component update (computationally intensive)
[0091] Subtasks: Electric field (E-field) update: Calculate the electric field components at each grid point according to the difference formula of Maxwell's equations. Magnetic field (H-field) update: Calculate the magnetic field components at each grid point according to the difference formula of Maxwell's equations. Write a kernel function using CUDA, and each thread processes one grid point.
[0092] Since the field component update involves parallel calculations of a large number of grid points (each grid point is calculated independently), the parallel architecture of the GPU (such as the CUDA kernel function) can efficiently handle such tasks. Therefore, the execution unit is the GPU.
[0093] (3) Boundary condition handling
[0094] Boundary (Bloch, periodic, symmetric / antisymmetric): At the boundaries of the computational domain, the respective boundary conditions are processed. Except for the PML absorbing boundary, since the internal region of the PML layer is placed as a special material for iteration within the GPU, Bloch, periodic, and symmetric / antisymmetric boundary processing usually involves complex logical judgments and conditional branches, which are more suitable for the serial processing capabilities of the CPU. Therefore, the execution unit is the CPU.
[0095] (4) Excitation source injection
[0096] Add excitation at the specified location and time step. The position and time logic of the excitation source need to be precisely controlled, which is suitable for the task scheduling of the CPU. Therefore, the execution unit is the CPU.
[0097] (5) Data acquisition and storage
[0098] Field value recording: Save the field values at specific locations or time steps for post-processing. Data storage involves file I / O and memory management, which is suitable for the general processing capabilities of the CPU. Therefore, the execution unit is the CPU.
[0099] (6) Time stepping loop control
[0100] Iteration management: Control the total number of time steps of the simulation and coordinate the execution order of each step. Loop control requires global state management, which is suitable for the logical control of the CPU. Therefore, the execution unit is the CPU.
[0101] Thus, using the load prediction report of the sub-regions, adopting a static allocation strategy, the computationally intensive sub-regions are allocated to the GPU, while tasks such as control, boundary processing, and excitation injection are allocated to the CPU. The static allocation strategy avoids the additional overhead caused by real-time monitoring and data synchronization in dynamic scheduling, thereby achieving the optimization of the overall computing efficiency.
[0102] Furthermore, based on the above task division, a hybrid GPU-CPU computing architecture is designed. Among them, the data flow design includes:
[0103] Data resident in GPU memory: Main field component arrays (Ex, Ey, Ez, Hx, Hy, Hz), material parameters.
[0104] Data resident in CPU memory: Boundary condition parameters, excitation source configuration, data storage buffer.
[0105] Furthermore, the task scheduling process includes:
[0106] Initialization (CPU): Initialize, partition, and allocate tasks, and upload the field data and material parameters to the GPU.
[0107] Time loop (controlled by CPU):
[0108] Step 1: The CPU triggers the GPU kernel function to update the electric field (E-field).
[0109] Step 2: The CPU triggers the GPU kernel function to update the magnetic field (H-field).
[0110] Step 3: The CPU processes the boundary conditions.
[0111] Step 4: The CPU injects the excitation source.
[0112] Step 5: The CPU copies the field data from the GPU to the memory for storage as needed.
[0113] End (CPU): Release the GPU video memory and save the result data.
[0114] Further, in some embodiments, after allocating the subtasks of each sub-region to the target computing unit, it further includes: asynchronously executing the subtasks of each sub-region based on the target computing unit to obtain the local calculation results of each sub-region; generating global simulation data according to the local calculation results of each sub-region.
[0115] Specifically, the preset asynchronous execution and data synchronization mechanism in the embodiments of the present application uses CUDA streams to overlap GPU computing and CPU tasks. For example, when the GPU is calculating the n+1 step, the CPU processes the boundary conditions of the n step. In addition, data transfer from the GPU to the CPU is only performed when data is saved or necessary synchronization is required, minimizing data transfer. At the same time, buffers are reused in the GPU video memory to avoid frequent memory allocation and release.
[0116] Further, summarize the local simulation results of all sub-regions and each computing unit (GPU / CPU), perform necessary post-processing and data correction, generate global electromagnetic field distribution maps, time-domain evolution curves, and other simulation data reports, and display them through the user interface or save them to the database.
[0117] It should be noted that after the simulation calculation in the embodiments of the present application, the results of each computing unit are summarized and spliced. These post-processing steps are used for data analysis and result display, and have strong reusability and pre-setting, and can be executed batch after the simulation is completed.
[0118] In addition, the embodiments of the present application can also test the overall system performance under different simulation scenarios and different hardware architectures, debug the load prediction and allocation parameters, and ensure efficient and stable operation.
[0119] Thus, through the "prediction - allocation" integrated method, the present invention not only fully utilizes the capabilities of each computing unit, but also ensures data synchronization and boundary interaction between tasks, thereby greatly reducing the computing bottleneck and improving the stability and execution efficiency of the simulation task while ensuring the simulation accuracy.
[0120] To facilitate a clearer and more intuitive understanding of the overall FDTD simulation process of the embodiments of the present application by those skilled in the art, the following will be combined with Figure 4 for detailed description.
[0121] Specifically, as Figure 4 shown, the overall FDTD simulation process includes the following steps:
[0122] Input simulation task data: The system receives simulation task data including device geometry, material properties, and other initial parameters.
[0123] Global mesh generation: Using the FDTD mesh generation algorithm, the entire simulation area is divided into multiple mesh cells.
[0124] Mesh load assessment: Calculate key load metrics (computational load, memory occupancy, communication overhead) for each mesh to provide a basis for subsequent partitioning.
[0125] Sub-region partitioning: According to the mesh load assessment results, the global region is divided into multiple sub-regions using a load balancing strategy.
[0126] Static resource allocation: According to the load conditions of the sub-regions, tasks are statically allocated: computationally intensive tasks are allocated to the GPU, and control logic, boundary condition handling, and data management tasks are allocated to the CPU.
[0127] Parallel computing: The GPU and CPU execute their respective allocated tasks simultaneously to achieve efficient computing during the FDTD simulation process.
[0128] Data synchronization and communication: During the calculation process of each sub-region, the overall simulation data consistency is ensured through boundary data exchange and synchronization control.
[0129] Result integration and output: Integrate the calculation results of each sub-region, generate global simulation data, and output the final result.
[0130] End: Complete the entire FDTD simulation process.
[0131] Thus, by utilizing the prediction results, tasks are pre-partitioned to the most suitable processing units in a static allocation manner: computationally intensive tasks with high requirements for parallelism (such as numerical iteration in high-density mesh regions in the FDTD algorithm) are allocated to the GPU to utilize its high-throughput advantage of large-scale parallel computing; while tasks such as initialization of complex structure type parameters, control logic, and data management are allocated to the CPU. This strategy avoids the additional overhead and uncertainty brought by the real-time adjustment of dynamic load balancing, ensuring an optimal match in resource utilization and execution efficiency throughout the simulation process.
[0132] Further, to enable those skilled in the art to more clearly and intuitively understand the heterogeneous collaborative execution process of the embodiments of the present application, the following will be described in detail in conjunction with Figure 5 for elaboration.
[0133] Specifically, as Figure 5 shown, the heterogeneous collaborative execution process includes the following parts:
[0134] (1) CPU part:
[0135] Initialization of simulation tasks: This includes the initialization of simulation information and material data, which is usually completed before the start of the simulation task and is the work of the preparation stage. The CPU is responsible for initializing the simulation tasks.
[0136] Data preprocessing and uploading: The CPU preprocesses the input data and uploads it to the GPU asynchronously.
[0137] Boundary condition processing and excitation injection: The CPU processes the boundary conditions and the injection tasks of the excitation sources.
[0138] Result integration and output: The CPU summarizes the calculation results of each part and outputs the final simulation result.
[0139] (2) GPU part (using CUDA Streams):
[0140] Receiving data: The GPU receives the data uploaded from the CPU through CUDA Streams.
[0141] E-field update calculation: Parallel calculation of the electric field (E-field) is performed on CUDA Stream 1.
[0142] H-field update calculation: Parallel calculation of the magnetic field (H-field) is performed on CUDA Stream 2.
[0143] Asynchronous communication: Dashed arrows are used to represent the asynchronous data transfer and boundary data synchronization between the CPU and the GPU, ensuring that the boundary condition information is transferred consistently between different computing units, thereby realizing collaborative computing.
[0144] Next, the simulation task resource allocation system related to the simulation task resource allocation method of the present application will be introduced.
[0145] Specifically, the simulation task resource allocation system of the embodiments of the present application adopts a modular design, mainly including a mesh division and load calculation module, a sub-region division and mesh generation module, a static resource allocation and parallel calculation module, and a result integration and output module. In addition, it also includes an offline load prediction model training module for constructing the load prediction model. Each module is closely connected through data streams and control signals, and together they complete the whole process from input data to the output of the final simulation result.
[0146] Among them, the offline load prediction model training module is used to train the load prediction model by using historical simulation data (including the real load conditions and computing power utilization rates under various simulation scenarios and different resource configurations). The prediction accuracy is evaluated by methods such as cross-validation, and the weight parameters such as the computing load COMPUTE_WEIGHT, memory occupancy MEMORY_WEIGHT, and communication overhead COMM_WEIGHT of each material are adjusted to ensure that the load prediction value output by the model is consistent with the actual simulation load. In addition, the offline load prediction model training module outputs the trained load prediction model, which is directly called in the subsequent simulation process to evaluate the load conditions of each grid cell.
[0147] The offline load prediction model training module completes offline training before the system goes online, and the training results are input as parameter configurations into the grid division and load calculation module.
[0148] Furthermore, the grid division and load calculation module includes the following functions and capabilities:
[0149] Input data processing: Receive simulation task data, including device geometry information, material properties, and other initial parameters.
[0150] Global grid division: Use the FDTD grid division algorithm to divide the entire simulation area into multiple grid cells, and extract the basic information (such as size and position) of each grid and the corresponding material classification (based on the MaterialType enumeration: Normal, Dispersion, Conformal, Dis_conformal, PML, Anisotropic).
[0151] Load evaluation: Call the offline trained load prediction model to calculate the load metrics (including computing load COMPUTE_WEIGHT, memory occupancy MEMORY_WEIGHT, and communication overhead COMM_WEIGHT) for each grid, and form global load distribution data.
[0152] Module connection: The output global grid data and the load metrics of each grid are used as the basic input for subsequent sub-region division.
[0153] Further, the sub-region division and mesh generation module is used to divide the global simulation region into multiple sub-domains based on the calculation results of the load prediction model. The load prediction model predicts the compute load (COMPUTE_WEIGHT), memory occupancy (MEMORY_WEIGHT), and communication overhead (COMM_WEIGHT) of each grid to generate the load distribution data of the global simulation region. According to these load distribution data and combined with the number of processes specified by the user, a load balancing strategy is adopted to divide the global region into multiple sub-domains, with the goal of making the total load in each sub-domain as balanced as possible, thereby ensuring the efficiency of subsequent parallel computing. For each sub-region, a mesh file containing detailed mesh information is generated, including the size, position information, and material properties of the mesh (based on the MaterialType enumeration type, such as Normal, Dispersion, Conformal, etc.). Finally, a load prediction report is generated, recording the load prediction data of each sub-region (such as compute load, memory occupancy, and communication overhead).
[0154] Module connection: The generated sub-region mesh data and load report are passed to the static resource allocation and parallel computing module as the basis for task scheduling.
[0155] Further, the static resource allocation and parallel computing module includes the following functions and capabilities:
[0156] Static resource allocation: According to the sub-region load prediction report, a static allocation strategy is adopted to pre-allocate the tasks of each sub-region to appropriate computing units. Allocation principle: Compute-intensive tasks (such as high-load grid regions) are allocated to GPUs (suitable for large-scale parallel computing), while tasks such as control logic, boundary condition handling, excitation source injection, data acquisition, and time stepping control are allocated to CPUs.
[0157] Parallel computing: GPU part: Use CUDA kernel functions to implement the parallel update of the electric and magnetic fields, and handle a large number of numerical calculations in the FDTD algorithm. CPU part: Execute tasks such as logical judgment of boundary conditions, injection of excitation sources, data acquisition, and time loop control. Synchronization and communication: Through asynchronous execution (such as using CUDA Streams) and data synchronization mechanisms, achieve efficient cooperation between the CPU and GPU to ensure the consistency of boundary data between sub-regions.
[0158] Module connection: During the parallel computing process, the local calculation results of each sub-region are transferred to the result integration module through data synchronization.
[0159] Further, the result integration and output module includes functions and capabilities:
[0160] Data integration: Summarize the local simulation results of all sub-regions and each computing unit (GPU / CPU), and perform necessary post-processing and data correction.
[0161] Result output: Generate a global electromagnetic field distribution map, a time-domain evolution curve, and other simulation data reports, and display or save them to the database through the user interface.
[0162] Module connection: Receive and integrate the result data from the static resource allocation and parallel computing module, and finally display the comprehensive simulation results to the user.
[0163] The overall data flow and the relationship between modules are introduced in detail below.
[0164] In the offline training stage, the load prediction model is trained with historical data, and the trained model parameters are configured into the system.
[0165] At the beginning of the simulation process, the mesh generation and load calculation module generates a global mesh from the input data, and uses the load prediction model to evaluate the load of each mesh, and outputs the global mesh and load data.
[0166] In the sub-region division stage, the sub-region division and mesh generation module evenly divides the simulation region based on the global load data, and generates a detailed sub-region mesh file and a load prediction report.
[0167] In the task allocation and parallel computing stage, the static resource allocation and parallel computing module: according to the sub-region load report, statically allocate tasks to the GPU or CPU, start parallel computing, and ensure data synchronization and boundary interaction through an asynchronous mechanism.
[0168] In the result integration stage, the result integration and output module summarizes the output results of each computing unit, performs post-processing, and then displays or stores the final simulation data.
[0169] According to the simulation task resource allocation method of the embodiment of the present application, the current simulation task performs global mesh division to obtain a plurality of mesh units, extracts the key information of each mesh unit, outputs the load index corresponding to each mesh unit based on the target load prediction model, generates global load distribution data according to the load indexes of all mesh units, divides the global simulation region into a plurality of sub-regions, divides the simulation task into a plurality of sub-tasks according to the load prediction data of each sub-region, and allocates each sub-task to the target computing unit. Thereby, the problems of insufficient fine-grained task division and large dynamic scheduling overhead in the prior art are solved, the overall computing efficiency and resource utilization rate are improved, and the system complexity and synchronization delay are reduced on the premise of ensuring the simulation accuracy.
[0170] Next, a simulation task resource allocation device according to an embodiment of the present application is described with reference to the accompanying drawings.
[0171] Figure 6 It is a block diagram of the simulation task resource allocation device according to an embodiment of the present application.
[0172] As Figure 6 shown, the simulation task resource allocation device 10 includes: an acquisition module 100, a division module 200, a generation module 300, a processing module 400, and an allocation module 500.
[0173] Among them, the acquisition module 100 is used to acquire the current simulation task; the division module 200 is used to perform global grid division on the current simulation task to obtain a plurality of grid cells, and extract the key information of each grid cell; the generation module 300 is used to input the key information of each grid cell into the target load prediction model, the target load prediction model outputs the corresponding load metrics, and generates global load distribution data based on the load metrics of all grid cells; the processing module 400 is used to divide the global simulation area into a plurality of sub-areas according to the global load distribution data, and obtain the load prediction data of each sub-area; the allocation module 500 is used to divide the simulation task into a plurality of subtasks according to the load prediction data of each sub-area, and allocate each subtask to the target computing unit.
[0174] Further, in some embodiments, the generation module 300 is further used to: acquire historical simulation data; train the initial load prediction model based on the historical simulation data, and after the training is completed, verify the initial load prediction model based on a preset verification method to obtain a verification result. If the verification result does not meet the preset verification conditions, adjust the parameters of the initial load prediction model based on a preset parameter adjustment strategy until the verification result meets the preset verification conditions, obtain the verified load prediction model, and use the verified load prediction model as the target load prediction model.
[0175] Further, in some embodiments, the allocation module 500 is used to: allocate the first type of subtask to the graphics processor according to the load prediction data of each sub-area, and allocate the second type of subtask, the third type of subtask, the fourth type of subtask, the fifth type of subtask, and the sixth type of subtask to the central processing unit; where the first type of subtask is a preset computationally intensive task, the second type of subtask is a preset control logic task, the third type of subtask is a preset boundary condition processing task, the fourth type of subtask is a preset excitation source injection task, the fifth type of subtask is a preset data acquisition and storage task, and the sixth type of subtask is a preset time stepping loop control task.
[0176] Further, in some embodiments, after allocating the subtasks of each sub-region to the target computing unit, the allocation module 500 is further configured to: asynchronously execute the subtasks of each sub-region based on the target computing unit to obtain the local computing results of each sub-region; generate global simulation data according to the local computing results of each sub-region.
[0177] Further, in some embodiments, after dividing the global simulation region into multiple sub-regions according to the global load distribution data, the processing module 400 is further configured to: generate a grid file for each sub-region, where the grid file of each sub-region includes at least one of grid size, position information, and material properties; generate a load prediction report according to the grid file of each sub-region, where the load prediction report includes the load prediction data of each sub-region.
[0178] Further, in some embodiments, the load metrics of each classified grid cell include at least one of computing load, memory occupancy, and communication overhead, and the key information of each grid cell includes: the size, position, and corresponding material classification of each grid cell.
[0179] It should be noted that the foregoing explanation of the embodiments of the simulation task resource allocation method is also applicable to the simulation task resource allocation device of this embodiment, and will not be repeated here.
[0180] According to the simulation task resource allocation device of the embodiments of the present application, the current simulation task performs global grid division to obtain multiple grid cells, extracts the key information of each grid cell, outputs the load metrics corresponding to each grid cell based on the target load prediction model, generates global load distribution data according to the load metrics of all grid cells, divides the global simulation region into multiple sub-regions, divides the simulation task into multiple subtasks according to the load prediction data of each sub-region, and allocates each subtask to the target computing unit. Thereby, the problems of insufficient fine-grained task division and large dynamic scheduling overhead in the prior art are solved, the overall computing efficiency and resource utilization rate are improved, and the system complexity and synchronization delay are reduced on the premise of ensuring the simulation accuracy.
[0181] Figure 7 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device may include:
[0182] A memory 701, a processor 702, and a computer program stored on the memory 701 and executable on the processor 702.
[0183] When the processor 702 executes the program, it implements the simulation task resource allocation method provided in the foregoing embodiments.
[0184] Further, the electronic device further includes:
[0185] A communication interface 703 for communication between the memory 701 and the processor 702.
[0186] A memory 701 for storing a computer program that can run on the processor 702.
[0187] The memory 701 may include high-speed RAM memory and may also include non-volatile memory, such as at least one disk memory.
[0188] If the memory 701, the processor 702, and the communication interface 703 are implemented independently, the communication interface 703, the memory 701, and the processor 702 can be interconnected via a bus to complete communication with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0189] Optionally, in a specific implementation, if the memory 701, the processor 702, and the communication interface 703 are integrated on a single chip, the memory 701, the processor 702, and the communication interface 703 can complete communication with each other through an internal interface.
[0190] The processor 702 may be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0191] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-described simulation task resource allocation method is implemented.
[0192] The embodiments of the present application also provide a computer program product, including a computer program or instructions, and when the computer program or instructions are executed, the above-described simulation task resource allocation method is implemented.
[0193] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or N embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0194] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of this application, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically and clearly defined.
[0195] Although the embodiments of this application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting this application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.
Claims
1. A method for allocating simulation task resources, characterized in that, It includes the following steps: Obtain the current simulation task; Perform global mesh division on the current simulation task to obtain multiple mesh cells, and extract the key information of each mesh cell; Input the key information of each mesh cell into the target load prediction model, the target load prediction model outputs the corresponding load metrics, and generate global load distribution data based on the load metrics of all mesh cells; Divide the global simulation area into multiple sub-areas according to the global load distribution data, and obtain the load prediction data of each sub-area; Divide the simulation task into multiple sub-tasks according to the load prediction data of each sub-area, and allocate each sub-task to the target computing unit.
2. The simulation task resource allocation method according to claim 1, wherein The training method of the target load prediction model includes: Obtain historical simulation data; Train the initial load prediction model based on the historical simulation data, and after the training is completed, verify the initial load prediction model based on a preset verification method to obtain a verification result. If the verification result does not meet the preset verification conditions, adjust the parameters of the initial load prediction model based on a preset parameter adjustment strategy until the verification result meets the preset verification conditions, obtain the verified load prediction model, and use the verified load prediction model as the target load prediction model.
3. The simulation task resource allocation method according to claim 1, wherein The step of dividing the simulation task into multiple sub-tasks according to the load prediction data of each sub-area and allocating each sub-task to the target computing unit includes: According to the load prediction data of each sub-area, allocate the first type of sub-task to the graphics processor, and allocate the second type of sub-task, the third type of sub-task, the fourth type of sub-task, the fifth type of sub-task, and the sixth type of sub-task to the central processing unit; Among them, the first type of sub-task is a preset compute-intensive task, the second type of sub-task is a preset control logic task, the third type of sub-task is a preset boundary condition processing task, the fourth type of sub-task is a preset excitation source injection task, the fifth type of sub-task is a preset data acquisition and storage task, and the sixth type of sub-task is a preset time stepping loop control task.
4. The simulation task resource allocation method according to claim 1, wherein After allocating the sub-tasks of each sub-area to the target computing unit, it further includes: Asynchronously execute the sub-tasks of each sub-area based on the target computing unit to obtain the local calculation results of each sub-area; Generate global simulation data based on the local calculation results of each sub-area.
5. The simulation task resource allocation method according to claim 1, wherein After dividing the global simulation area into multiple sub-areas according to the global load distribution data, it further includes: Generate a mesh file for each sub-area, where the mesh file for each sub-area includes at least one of mesh size, position information, and material properties; Generate a load prediction report according to the mesh file of each sub-area, where the load prediction report includes the load prediction data of each sub-area.
6. The simulation task resource allocation method according to claim 1, wherein The load metrics of each classified mesh cell include at least one of compute load, memory occupancy, and communication overhead, and the key information of each mesh cell includes the size, position, and corresponding material classification of each mesh cell.
7. A simulation task resource allocation device, characterized in that, It includes: An acquisition module, configured to acquire a current simulation task; A division module, configured to perform global grid division on the current simulation task to obtain a plurality of grid cells, and extract key information of each grid cell; A generation module, configured to input the key information of each grid cell into a target load prediction model, the target load prediction model outputs corresponding load metrics, and generate global load distribution data according to the load metrics of all grid cells; A processing module, configured to divide the global simulation area into a plurality of sub-areas according to the global load distribution data, and obtain load prediction data of each sub-area; An allocation module, configured to divide the simulation task into a plurality of subtasks according to the load prediction data of each sub-area, and allocate each subtask to a target computing unit.
8. An electronic device, characterized in that, Comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, the processor executes the program to implement the simulation task resource allocation method according to any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to be used for implementing the simulation task resource allocation method according to any one of claims 1-6.
10. A computer program product, characterized in that, Including a computer program, when the computer program is executed by the processor, it is used for implementing the simulation task resource allocation method according to any one of claims 1-6.
Citation Information
Cited By
Unstructured grid division method based on structured grid and electronic equipment
CN120726262A
Accelerator card simulation method and device, equipment, storage medium and program product
CN121210280A
Accelerator card emulation methods, devices, equipment, storage media and software products
CN121210280B