Parallel Performance Optimization Method for Electromagnetic Transient Simulation Based on RoCE Network

By building a submodule list, partition structure and labeling processing, combined with bandwidth priority queue and hierarchical packaging delay, the simulation performance degradation caused by the difference in communication frequency between modules in the RoCE network is solved, and efficient parallel computing and network communication are realized in electromagnetic transient simulation.

CN119788609BActive Publication Date: 2025-07-25FANGXIN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510057309.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-07-25
Estimated Expiration
2045-01-14

AI Technical Summary

Technical Problem

In the prior art, electromagnetic transient simulation based on RoCE networks is significantly different in communication frequency between modules due to the multi-level time step parallel algorithm, which cannot achieve efficient allocation and utilization of network bandwidth resources, resulting in a degradation of simulation performance.

Method used

By analyzing the configuration parameters of multi-level time step length, building a submodule list and clarifying the time step length settings of each submodule, combining the simulation partition structure to allocate the simulation subdomain range in the RoCE network, labeling the interface variables, generating a time step priority queue and setting a bandwidth allocation strategy, and implementing a hierarchical encapsulation delay and parallel dependency verification algorithm to ensure data transmission consistency.

Benefits of technology

It significantly improves the modular management efficiency of simulation tasks, realizes efficient dynamic scheduling of network bandwidth resources, alleviates the congestion problem caused by uneven allocation of communication resources, and improves the parallel computing performance and network communication utilization efficiency of electromagnetic transient simulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119788609B_ABST
    Figure CN119788609B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for optimizing the parallel performance of electromagnetic transient simulation based on the RoCE network, specifically relating to the technical field of simulation performance optimization. By analyzing the configuration parameters of multi-level time steps, constructing a sub-module list and clarifying the time step settings, a simulation partition structure is established based on the sub-module list and the range of simulation sub-domains is allocated; through the tagging process of interface variables and the configuration of memory areas by the Remote Direct Memory Access protocol, efficient management of data transmission is achieved; a time step priority queue is generated according to the time step differences of sub-modules, and a bandwidth allocation strategy is set, combined with the hierarchical encapsulation delay of high- and low-frequency interface variables, to optimize the utilization efficiency of network bandwidth; before global convergence, a parallel dependency verification algorithm is applied to retrieve the interface variable tag data, and the data that has not been transmitted is re-sent to ensure the consistency of simulation data, improving the utilization rate of communication resources and solving the performance bottleneck problem caused by communication frequency differences in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of simulation performance optimization, and more specifically, to a method for optimizing the parallel performance of electromagnetic transient simulation based on the RoCE network. Background Art

[0002] Electromagnetic transient simulation is an important tool for evaluating the dynamic performance of a system and verifying control strategies. In order to accelerate the simulation speed and meet the real-time computing requirements of complex scenarios, a distributed parallel computing method based on the RoCE (Remote Direct Memory Access over Converged Ethernet) network is gradually adopted. In this method, different computing modules are scheduled through a multi-level time-step parallel algorithm, and the efficient communication capabilities of the RoCE network are utilized to achieve cross-node data interaction and collaborative computing. However, due to the high computational complexity and large-scale data transmission requirements of electromagnetic transient simulation, many challenges are faced in the allocation and utilization of network resources.

[0003] In the prior art, in the electromagnetic transient simulation based on the RoCE network, due to the significant difference in communication frequencies between modules caused by the multi-level time-step parallel algorithm, it is impossible to achieve efficient allocation and utilization of network bandwidth resources, which easily results in waste of communication resources or local congestion, and further leads to a decline in simulation performance. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a method for optimizing the parallel performance of electromagnetic transient simulation based on the RoCE network to solve the problems raised in the above background art.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] A method for optimizing the parallel performance of electromagnetic transient simulation based on the RoCE network, comprising the following steps:

[0007] S1: Analyze the configuration parameters of the multi-level time step, construct a sub-module list corresponding to the electromagnetic transient simulation model, and clarify the time step settings of each sub-module;

[0008] S2: Establish a cross-node simulation partition structure based on the sub-module list, and allocate the simulation sub-domain range of each node in the RoCE network;

[0009] S3: Perform tagging processing on the interface variables of each simulation sub-domain, assign a unique tag identifier to each interface variable, and configure its corresponding memory area through the Remote Direct Memory Access protocol;

[0010] S4: Generate a time step priority queue according to the time step differences of the sub-modules, and set a bandwidth allocation strategy corresponding to the time step priority in the lossless flow control mechanism of the RoCE network;

[0011] S5: When the sub-module data of the high-frequency time step is updated, perform hierarchical encapsulation delay on the interface variable sending window of the low-frequency time step.

[0012] S6: Before the global convergence of the overall simulation is completed, apply the parallel dependency verification algorithm to retrieve the interface variable tag data of all computing nodes one by one, and perform a secondary transmission on the interface variable data that has not completed transmission.

[0013] In a preferred embodiment, parse the configuration parameters of the multi-level time step, construct a sub-module list corresponding to the electromagnetic transient simulation model, and clarify the time step settings of each sub-module, specifically including:

[0014] Obtain the physical attributes and dynamic behavior characteristics of each module in the electromagnetic transient simulation model, and analyze the computing requirements of each module in combination with the network topology structure of the electromagnetic transient simulation model.

[0015] Determine the time step range applicable to different modules according to the dynamic behavior characteristics and computing requirements of each module.

[0016] Divide the electromagnetic transient simulation model into multiple sub-modules, generate a list corresponding to each sub-module, and mark the time step settings and associated input / output data types of each sub-module in the list.

[0017] In a preferred embodiment, establish a cross-node simulation partition structure based on the sub-module list, and allocate the simulation sub-domain range of each node in the RoCE network, specifically including:

[0018] According to the time step settings and input / output data types in the sub-module list, divide the simulation model into multiple computing regions, and determine the boundary interfaces of each computing region.

[0019] Based on the network topology structure of the simulation model and the computing requirements of each computing region, allocate each computing region to a computing node.

[0020] When allocating computing nodes, adjust the boundaries of the computing regions according to the computing complexity and data transmission requirements of each computing region to achieve load balancing.

[0021] Number the allocated computing regions, and generate a partition structure file recording the computing region node allocation information and the positions of interface variables.

[0022] Set the corresponding simulation sub-domain range for each computing node in the RoCE network, and the simulation sub-domain range includes the boundary interface variables of the computing region and their associated memory access permissions.

[0023] In a preferred embodiment, the interface variables of each simulation sub-domain are tagged, a unique tag identifier is assigned to each interface variable, and its corresponding memory area is configured through the Remote Direct Memory Access (RDMA) protocol. Specifically, it includes:

[0024] Obtain the boundary interface variables of each simulation sub-domain, and identify the unique characteristics of the interface variables based on the interface variable data types and boundary positions recorded in the simulation partition structure file;

[0025] Assign a unique tag identifier to each interface variable. The tag identifier includes the simulation sub-domain number to which the interface variable belongs, the data type of the interface variable, and the memory offset address of the interface variable;

[0026] Bind the tag identifier of each interface variable to its corresponding memory area, and allocate read and write permissions to each memory area through the Remote Direct Memory Access (RDMA) protocol;

[0027] When allocating read and write permissions, establish a memory address mapping table for the boundary interface variables of each simulation sub-domain. The memory address mapping table records the tag identifier of the interface variable, the starting address of the memory area, and the read and write permission levels;

[0028] Upload the established memory address mapping table to the configuration file of the Remote Direct Memory Access (RDMA) protocol to complete the configuration of the memory area of the interface variable.

[0029] In a preferred embodiment, a time step priority queue is generated based on the time step differences of the sub-modules, and a bandwidth allocation policy corresponding to the time step priority is set in the lossless flow control mechanism of the RoCE network. Specifically, it includes:

[0030] Obtain the time step settings of the sub-modules, classify the sub-modules according to the length of the time steps, assign higher priorities to the sub-modules with shorter time steps, and assign lower priorities to the sub-modules with longer time steps;

[0031] Generate a time step priority queue based on the priority classification results. The time step priority queue contains the priority numbers of each sub-module and their corresponding time steps;

[0032] In the lossless flow control mechanism of the RoCE network, allocate bandwidth quotas to each priority queue. The bandwidth quotas are calculated based on the priority numbers and data transmission frequencies;

[0033] By adjusting the parameters of the bandwidth allocation policy in the priority queue, ensure that the bandwidth ratio of the high-priority queue is higher than that of the low-priority queue and avoid resource competition;

[0034] Record the bandwidth allocation information of the time step priority queue in the lossless flow control mechanism of the RoCE network. The bandwidth allocation information includes the priority number, the corresponding bandwidth ratio, and the time step range of the sub-module.

[0035] In a preferred embodiment, when the sub-module data of the high-frequency time step is updated, a hierarchical encapsulation delay is implemented for the interface variable sending window of the low-frequency time step, specifically including:

[0036] Detect the request status of the sub-module for updating data of the high-frequency time step, and obtain the list of boundary interface variables of the high-frequency time step sub-module;

[0037] Classify the interface variables of the low-frequency time step, and divide the low-frequency interface variables into multiple priority queues according to the length of the low-frequency time step and the data transmission priority;

[0038] Set the encapsulation delay time for each priority queue. The encapsulation delay time is determined according to the level of the priority queue where the interface variable is located. The lower the priority level, the longer the encapsulation delay time;

[0039] During the update cycle of the high-frequency time step sub-module, adjust the sending window of the low-frequency interface variables, and batch send the encapsulated low-frequency interface variables to the corresponding receiving nodes;

[0040] Record the encapsulation delay parameters for each adjustment of the sending window, including the priority number, delay time, and transmission timestamp of the low-frequency interface variable.

[0041] In a preferred embodiment, before the global convergence of the overall simulation is completed, apply the parallel dependency verification algorithm to retrieve the interface variable label data of all computing nodes one by one, and perform a secondary send for the interface variable data that has not been transmitted, specifically including:

[0042] Obtain the list of interface variable label data of each computing node. The interface variable label data includes the label identifier of the interface variable, the computing node number to which it belongs, and its transmission status;

[0043] During the global convergence calculation process, retrieve the interface variable label data one by one according to the parallel dependency verification algorithm, and verify whether the transmission status of each interface variable is the completed status;

[0044] For the interface variable data whose transmission status is found to be incomplete during the retrieval process, record the computing node number to which it belongs and the amount of data that has not been transmitted;

[0045] Based on the recorded amount of data that has not been transmitted and the computing node number, initiate a secondary send request to the corresponding computing node;

[0046] After receiving the second transmission data, update the tag data status of the corresponding interface variable to completed, and record the final transmission status of all interface variables in the convergence log file.

[0047] Technical effects and advantages of the electromagnetic transient simulation parallel performance optimization method based on RoCE network of the present invention:

[0048] 1. By parsing the configuration parameters of multi-level time steps, constructing a sub-module list corresponding to the electromagnetic transient simulation model, and clarifying the time step settings of each sub-module, and allocating the simulation sub-domain range in the RoCE network in combination with the simulation partition structure, the modular management efficiency of the simulation task is significantly improved. At the same time, through the tagging process of interface variables and the configuration of memory areas by the Remote Direct Memory Access protocol, the data transmission consistency of interface variables and the accuracy of resource management are ensured, laying a foundation for the efficient cooperation of distributed simulation computing.

[0049] 2. Aiming at the significant differences in the communication frequencies between modules, a method for generating a time step priority queue and a corresponding bandwidth allocation strategy are proposed, and combined with the hierarchical encapsulation delay mechanism of high- and low-frequency interface variables, an efficient dynamic scheduling of network bandwidth resources is achieved. Combined with the parallel dependency check algorithm, the present invention can timely retrieve and process the data that has not been transmitted, and improve the data consistency and the reliability of global simulation convergence through the second transmission mechanism, thereby effectively alleviating the congestion problem caused by uneven allocation of communication resources in the prior art, and comprehensively improving the parallel computing performance of electromagnetic transient simulation and the utilization efficiency of network communication. Description of the Drawings

[0050] Figure 1 Schematic diagram of the electromagnetic transient simulation parallel performance optimization method based on RoCE network of the present invention. Detailed Embodiments

[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0052] Embodiment: Figure 1 The electromagnetic transient simulation parallel performance optimization method based on RoCE network of the present invention is given, which includes the following steps:

[0053] S1: Parse the configuration parameters of multi-level time steps, construct a sub-module list corresponding to the electromagnetic transient simulation model, and clarify the time step settings of each sub-module.

[0054] S2: Establish a cross - node simulation partition structure based on the sub - module list and allocate the simulation sub - domain range for each node in the RoCE network.

[0055] S3: Perform tagging processing on the interface variables of each simulation sub - domain, assign a unique tag identifier to each interface variable, and configure its corresponding memory area through the Remote Direct Memory Access (RDMA) protocol.

[0056] S4: Generate a time - step priority queue based on the time - step differences of sub - modules and set the bandwidth allocation policy corresponding to the time - step priority in the lossless flow control mechanism of the RoCE network.

[0057] S5: When the data of the sub - module with a high - frequency time - step is updated, implement hierarchical encapsulation delay for the sending window of the interface variables with a low - frequency time - step.

[0058] S6: Before the global convergence of the overall simulation is completed, apply the parallel dependency verification algorithm to retrieve the interface variable tag data of all computing nodes one by one, and perform a secondary transmission on the interface variable data that has not been completed for transmission.

[0059] Analyze the configuration parameters of the multi - level time - step, construct a sub - module list corresponding to the electromagnetic transient simulation model, and clarify the time - step settings of each sub - module, specifically including:

[0060] Obtain the physical attributes and dynamic behavior characteristics of each module in the electromagnetic transient simulation model, and analyze the computing requirements of each module in combination with the network topology structure of the electromagnetic transient simulation model:

[0061] Analyze each module in the model (such as generators, transformers, transmission lines, etc.) one by one, extract its key physical parameters, including parameters such as impedance, admittance, inductance, capacitance, etc., and record these parameters as the basic physical attributes of the module. Analyze the dynamic behavior characteristics of the module. For example, analyze its dynamic response characteristics based on the electromagnetic rotor model of the generator, or analyze its high - frequency dynamic characteristics based on the switching frequency of the power electronics module. Associate these physical attributes and dynamic behavior characteristics with the computing requirements of the electromagnetic transient simulation as the basic data input for further module calculations.

[0062] Obtain the overall network topology diagram of the simulation model, including the connection relationships between nodes, the interaction boundaries between modules, and the interconnected topological structure, and record the positional relationships of each module in the network. Through topological structure analysis, identify the possible strong coupling relationships between modules, such as electrical interconnections or magnetic couplings between adjacent modules. According to the physical properties and dynamic behavior characteristics of each module, combined with the role of the module in the topology, calculate the magnitude of the impact of its dynamic behavior on the simulation calculation. For example, for a high-frequency switching module, calculate its contribution to the rate of change of voltage in the network and evaluate its requirement for the time step size. Based on the complexity of the dynamic behavior and the calculation accuracy requirements of the module, evaluate the calculation requirements of each module one by one to provide a basis for time step size division.

[0063] Based on the dynamic behavior characteristics and calculation requirements of each module, determine the time step size ranges applicable to different modules:

[0064] According to the complexity of the dynamic behavior of the module, divide the time step size range suitable for its calculation. For example, set the high-frequency switching module to a microsecond-level time step size, while set the slow mechanical module to a millisecond-level time step size.

[0065] Group the modules with similar time step size ranges into a sub-module group. For example, group the high-frequency electronic modules into one sub-module group, while group the power equipment modules into another sub-module group.

[0066] Create a separate description file for each sub-module, recording the units, time step size range, and input / output data types included in the sub-module. For example, the input of the high-frequency module may be the instantaneous voltage, and the output is the switching action timing sequence; while the input of the slow module is the motor torque, and the output is the steady-state current.

[0067] Divide the electromagnetic transient simulation model into multiple sub-modules, generate a list corresponding to each sub-module, and mark the time step size settings and associated input / output data types of each sub-module in the list:

[0068] Summarize the description files of each sub-module into a sub-module list, sort the list in descending order according to the time step size settings of the modules, and ensure that the time step size settings of each sub-module are clearly marked in the list. Record the associated input / output data types of each sub-module in the list for subsequent steps of simulation partitioning and management of interface variables. Save this sub-module list as a data file as the input for subsequent simulation partitioning structure and interface variable configuration.

[0069] Based on the sub-module list, establish a cross-node simulation partitioning structure and allocate the simulation sub-domain ranges of each node in the RoCE network, specifically including:

[0070] Divide the simulation model into multiple computational regions according to the time step settings and input / output data types in the sub-module list, and determine the boundary interfaces of each computational region:

[0071] Parse the sub-module list, extract the time step settings, input data types, and output data types of each sub-module, and clarify the simulation attributes of the sub-modules.

[0072] Based on the time step settings, group the sub-modules with similar time steps into one computational region. For example, for modules with microsecond-level time steps, such as high-frequency power electronic devices, their computational requirements are significantly higher than those of mechanical modules with millisecond-level time steps, so they form separate computational regions.

[0073] Determine the boundary interface variables of the computational regions. Boundary interface variables refer to the data variables that need to be exchanged between computational regions and other regions, such as voltage, current, etc. The boundary positions of these variables are determined by the input / output data types of the sub-modules.

[0074] Record the complete information of the sub-module scope, time step settings, and boundary interface variables of each computational region as the basis for the partitioning structure.

[0075] Based on the network topology of the simulation model and the computational requirements of each computational region, allocate each computational region to a computing node:

[0076] Parse the network topology of the simulation model, obtain the node connection relationships, device interaction methods, and communication paths, and identify the positions of each computational region in the overall topology.

[0077] Combine the dynamic computational requirements of each computational region and analyze its computational complexity. For example, calculate the computational complexity of each computational region through the formula: ; where, represents the computational complexity of the th computational region, represents the total time step of all sub-modules in the th computational region, represents the data interaction frequency (i.e., the update frequency of the boundary interface variables) of the th computational region, is the number of the computational region.

[0078] According to the computational complexity and communication requirements, allocate the computational regions to each computing node. For example, for regions with high computational complexity and large communication frequencies, they are preferentially allocated to nodes with high computing capabilities and low communication latencies.

[0079] When allocating computing nodes, adjust the boundaries of the computational regions according to the computational complexity and data transmission requirements of each computational region to achieve load balancing:

[0080] Perform a load assessment on the computing area of each computing node to ensure that the computing load is evenly distributed among all nodes.

[0081] Achieve load balancing by adjusting the boundaries of the computing area. For example, if the total computing complexity of a certain computing node exceeds that of other nodes, the load is balanced by adjusting the formula: ; where represents the computing resources allocated to the th computing node, represents the total number of computing nodes in the simulation system, represents the total computing resources available in the system, the total computing complexity of the th specific computing node, represents the summary item of the computing complexity of all nodes, represents a traversal index in the set of computing nodes, refers to the specific target node.

[0082] It should be noted that is the calculation for the specific node , is the summary of the total computing complexity of all nodes on a global scale; the difference is that is a specific single node, is an iteration variable among all nodes, used for summation or global calculation.

[0083] Adjust the boundary interface variables of the computing area to ensure that the data transmission after the computing area is divided still meets the simulation accuracy requirements.

[0084] Number the allocated computing areas and generate a partition structure file that records the node allocation information of the computing area and the positions of the interface variables:

[0085] Allocate a unique number to each computing area and record the correspondence between the number and the computing node.

[0086] Record the following information for each computing area in the partition structure file: the composition of the sub-module and its time step setting; the boundary interface variables and their data types; the computing node number to which it belongs.

[0087] The partition structure file serves as the basic data for the subsequent tagging of interface variables and RDMA configuration.

[0088] Set the corresponding simulation sub-domain range for each computing node in the RoCE network. The simulation sub-domain range includes the boundary interface variables of the computing area and their associated memory access permissions:

[0089] In the RoCE network, a simulation sub-domain range is allocated to each computing node. The simulation sub-domain range includes the computing area allocated to the computing node and its boundary interface variables. Remote direct memory access permissions are configured for the simulation sub-domain range of each computing node to ensure that the boundary interface variables of the computing area can be efficiently transmitted in the RoCE network. After configuration, the sub-domain range and memory access permissions of each computing node are recorded to support subsequent interface variable tagging and flow control policies.

[0090] Tag the interface variables of each simulation sub-domain, assign a unique tag identifier to each interface variable, and configure its corresponding memory area through the remote direct memory access protocol, specifically including:

[0091] Obtain the boundary interface variables of each simulation sub-domain, and identify the unique characteristics of the interface variables based on the interface variable data type and boundary position recorded in the simulation partition structure file:

[0092] Read the computing area information of each simulation sub-domain from the simulation partition structure file, and extract the boundary interface data between the computing areas. The boundary interface variable refers to the variable that needs to perform data interaction between different simulation sub-domains, such as dynamic variables such as voltage and current in adjacent computing areas.

[0093] For the interface variables of each computing area, analyze the unique characteristics of the interface variables one by one according to the simulation sub-domain number, data type (such as floating-point type or integer type), and boundary position to which they belong. The unique characteristics are used to identify the unique attributes of the interface variables in the entire simulation system to avoid data conflicts or duplicate processing.

[0094] Record the boundary attributes of each interface variable, including the simulation sub-domain number to which the variable belongs, the boundary position index, and the data type information, to provide a basis for subsequent tagging processing.

[0095] Assign a unique tag identifier to each interface variable. The tag identifier includes the simulation sub-domain number to which the interface variable belongs, the data type of the interface variable, and the memory offset address of the interface variable:

[0096] Generate a unique tag identifier for each interface variable according to the boundary attributes of the interface variable.

[0097] The composition of the tag identifier includes the following:

[0098] The simulation sub-domain number to which the interface variable belongs: indicates the sub-domain where the interface variable is located, used to distinguish different computing areas;

[0099] The data type of the interface variable: identifies the data format of the interface variable (such as 32-bit floating-point type or 64-bit integer type) to ensure consistency in transmission and processing;

[0100] Memory offset address of interface variable: It is used to indicate the specific position of the interface variable in the allocated memory area, facilitating data storage and retrieval.

[0101] Generate a label identifier through an algorithm, for example: ; where represents the label identifier of the interface variable, which is a unique integer used to identify the interface variable; represents the simulation subdomain number to which the interface variable belongs; represents the data type number of the interface variable (1 for floating-point type, 2 for integer type, and so on); represents the memory offset address of the interface variable, used to uniquely identify the position in memory.

[0102] Record the generated label identifier and associate it with the corresponding interface variable to ensure the uniqueness and traceability of the label identifier.

[0103] Bind the label identifier of each interface variable to its corresponding memory area, and allocate read and write permissions to each memory area through the Remote Direct Memory Access protocol:

[0104] Allocate a memory area for each interface variable. According to the memory offset address of the interface variable, map the variable to the shared memory of the simulation node.

[0105] Through the Remote Direct Memory Access protocol, set the read and write permissions for each memory area to ensure that the interface variable can be efficiently accessed by the corresponding simulation subdomain. For example, interface variables in subdomains with high-frequency time steps are preferentially allocated a higher level of read and write permissions to meet their real-time requirements.

[0106] Ensure that the label identifier of each interface variable corresponds one-to-one with the allocated memory area, and record this binding relationship in the simulation configuration file.

[0107] When allocating read and write permissions, establish a memory address mapping table for the boundary interface variables of each simulation subdomain. The memory address mapping table records the label identifier of the interface variable, the starting address of the memory area, and the read and write permission level:

[0108] Establish a memory address mapping table to record the label identifier of each interface variable and the starting address of its allocated memory area.

[0109] The fields of the memory address mapping table include: the label identifier of the interface variable; the starting address of the memory area; the read and write permission level of the memory area.

[0110] Determine the starting address of the memory area through the following formula: ; where represents the starting address of the memory area, represents the memory base address allocated to the simulation subdomain, Represents the memory offset address of the interface variable.

[0111] Generate the memory address mapping table based on the above fields, and uniformly manage the boundary interface variables of all simulation subdomains.

[0112] Upload the established memory address mapping table to the configuration file of the Remote Direct Memory Access (RDMA) protocol to complete the memory area configuration of the interface variables:

[0113] Upload the memory address mapping table to the configuration file of the Remote Direct Memory Access (RDMA) protocol to ensure that the simulation subdomains of each computing node can access their corresponding memory areas.

[0114] Record the following in the configuration file: the memory address range of each simulation subdomain; the label identification of each interface variable and its corresponding memory area information; the read / write permission level of the interface variable.

[0115] After the configuration is completed, verify whether the memory area of each interface variable is accessible to ensure the consistency of the binding between the label identification and the memory area.

[0116] Generate a time step priority queue based on the time step differences of the submodules, and set the bandwidth allocation policy corresponding to the time step priority in the lossless flow control mechanism of the RoCE network, specifically including:

[0117] Obtain the time step settings of the submodules, classify the submodules according to the length of the time step, and assign higher priority to the submodules with shorter time steps and lower priority to the submodules with longer time steps:

[0118] Extract the time step settings of each submodule from the submodule list. The time step represents the update period of the submodule in the simulation calculation, in seconds. For example, the time step of a high-frequency electronic module may be in the microsecond range, while the time step of a slow mechanical module may be in the millisecond range.

[0119] Classify the submodules according to the length of the time step: assign higher priority to the submodules with shorter time steps because their calculation frequency is higher and their real-time requirement for bandwidth is higher; assign lower priority to the submodules with longer time steps because their calculation frequency is lower and their bandwidth requirement is relatively loose.

[0120] Assign a priority number to each submodule. The priority number is an integer, and the smaller the value, the higher the priority. For example, the high-frequency electronic module may be assigned a priority number 1, while the slow mechanical module is assigned a priority number 2 or greater.

[0121] Generate a time-step priority queue based on the priority classification results. The time-step priority queue contains the priority numbers of each sub-module and their corresponding time steps:

[0122] Sort the sub-modules according to the priority numbers, arrange the sub-modules with increasing priority numbers in sequence, and generate a time-step priority queue.

[0123] Each entry in the time-step priority queue includes the following information: the priority number of the sub-module; the time step of the sub-module; the data transfer frequency of the sub-module (i.e., the update frequency of the boundary interface variables).

[0124] Record the generated time-step priority queue in a queue file for subsequent call by the bandwidth allocation strategy.

[0125] In the lossless flow control mechanism of the RoCE network, allocate bandwidth quotas for each priority queue. The bandwidth quotas are calculated based on the priority numbers and data transfer frequencies:

[0126] First, according to the levels of each priority number in the time-step priority queue, allocate a larger bandwidth proportion to the queues with smaller priority numbers (i.e., higher priorities), and at the same time allocate a smaller bandwidth proportion to the queues with larger priority numbers. Second, combined with the total data transfer frequencies of the sub-modules in each priority queue, further increase the bandwidth allocation ratio for the queues with high-frequency data transfer requirements to ensure that high-priority queues can meet the requirements of real-time transmission. Finally, the allocated bandwidth quotas are recorded as bandwidth parameters to guide the lossless flow control mechanism of the RoCE network to dynamically adjust transmission resources and ensure the rationality of bandwidth allocation and data transfer efficiency among multiple priority queues.

[0127] By adjusting the parameters of the bandwidth allocation strategy in the priority queue, ensure that the bandwidth proportion of high-priority queues is higher than that of low-priority queues and avoid resource contention:

[0128] Optimize the bandwidth proportion of high-priority queues. By adjusting the parameters of the bandwidth allocation strategy, such as the weighting factor, further increase the bandwidth proportion of high-priority queues. When adjusting the bandwidth allocation strategy, ensure that the minimum bandwidth quota of low-priority queues is not lower than the basic transmission requirements to avoid simulation task delays caused by resource contention. Use a hierarchical flow control model to record the adjusted bandwidth allocation parameters, including priority numbers, bandwidth proportions, and weighting factors.

[0129] Record the bandwidth allocation information of the time-step priority queue in the lossless flow control mechanism of the RoCE network. The bandwidth allocation information includes the priority number, the corresponding bandwidth proportion, and the time-step range of the sub-module:

[0130] Record the bandwidth allocation information of the time step priority queue in the configuration file of the lossless flow control mechanism of the RoCE network. The bandwidth allocation information includes: the priority number of each priority queue; the bandwidth ratio of each priority queue; the time step range of the sub-module corresponding to each priority queue.

[0131] Configure the bandwidth allocation parameters in the lossless flow control mechanism to ensure that the bandwidth ratio of each priority queue is allocated according to the adjusted policy.

[0132] After completing the configuration of the bandwidth allocation information, verify whether the bandwidth of each priority queue meets the requirements of the simulation system to ensure the real-time performance of the high-priority queue and the integrity of the low-priority queue.

[0133] When the data of the sub-module with a high-frequency time step is updated, implement hierarchical encapsulation delay for the send window of the interface variables with a low-frequency time step, specifically including:

[0134] Detect the request status of the sub-module that updates the data with a high-frequency time step, and obtain the list of boundary interface variables of the high-frequency time step sub-module:

[0135] Perform an update detection on the sub-module with a high-frequency time step to determine whether there is a data update request. The update request is usually identified by the status flag bit of the boundary interface variable. When the flag bit changes, it indicates that the high-frequency time step sub-module needs to update the data.

[0136] Obtain the list of sub-modules with update requests, extract the boundary interface variables of the high-frequency time step sub-modules involved, and record the number, the simulation sub-domain to which each interface variable belongs, and the data type.

[0137] Use the extracted boundary interface variables as the basis for the subsequent encapsulation delay processing of the low-frequency time step interface variables.

[0138] Classify the interface variables with a low-frequency time step, and divide the low-frequency interface variables into multiple priority queues according to the length of the low-frequency time step and the data transmission priority:

[0139] Obtain the list of low-frequency time step interface variables recorded in the simulation partition structure file, and classify them in combination with the time step settings and data transmission priorities of each interface variable.

[0140] Divide the low-frequency interface variables into different priority queues according to the length of the time step. For example: Interface variables with a shorter time step (close to the threshold of the high-frequency time step) are assigned to the high-priority queue; Interface variables with a longer time step are assigned to the low-priority queue.

[0141] Further refine the priority queue in combination with the data transmission frequency, and preferentially allocate the interface variables with high data transmission frequency to the higher-priority queue to optimize resource allocation.

[0142] Set an encapsulation delay time for each priority queue. The encapsulation delay time is determined according to the level of the priority queue where the interface variable is located. The lower the priority level, the longer the encapsulation delay time:

[0143] Set an encapsulation delay time for each priority queue. The length of the encapsulation delay time is associated with the priority level: The encapsulation delay time of the high-priority queue is set shorter to meet the real-time transmission requirements; the encapsulation delay time of the low-priority queue is set longer to save bandwidth resources.

[0144] The specific value of the encapsulation delay time is determined by a preset rule. For example: Make the encapsulation delay time inversely proportional to the time step to ensure reasonable hierarchical regulation of the priority queue.

[0145] While setting the encapsulation delay time, record the delay parameters of each priority queue in the encapsulation delay configuration file for subsequent operation calls.

[0146] During the update period of the high-frequency time step sub-module, adjust the sending window of the low-frequency interface variables, and batch send the low-frequency interface variables after encapsulation delay to the corresponding receiving nodes:

[0147] During the update period of the high-frequency time step sub-module, dynamically adjust the sending window of the low-frequency interface variables according to the delay parameters in the encapsulation delay configuration file. For the interface variables in the high-priority queue, immediately open the sending window and preferentially transmit the delayed data to the corresponding receiving nodes; for the interface variables in the low-priority queue, open the sending window after the set encapsulation delay time. When adjusting the sending window, optimize the order of data transmission to ensure that high-priority data is transmitted within the update period of the high-frequency sub-module, and low-priority data is sent in batches to avoid transmission conflicts.

[0148] Record the encapsulation delay parameters for each adjustment of the sending window, including the priority number, delay time, and transmission timestamp of the low-frequency interface variable:

[0149] After the adjustment of the sending window is completed, record the encapsulation delay parameters, including the following: The priority number of each low-frequency interface variable; the encapsulation delay time of the interface variable; the actual transmission timestamp of the interface variable. Store the recorded encapsulation delay parameters in the delay log file of the simulation system. The delay log file is used for subsequent simulation result verification and performance analysis. Verify the recorded encapsulation delay parameters to ensure that the delay time for each adjustment is consistent with the priority setting, and avoid data loss or conflicts.

[0150] Before the global convergence of the overall simulation is completed, the parallel dependency verification algorithm is applied to retrieve the interface variable label data of all computing nodes one by one, and the interface variable data that has not been transmitted is retransmitted. Specifically, it includes:

[0151] Obtain the interface variable label data list of each computing node. The interface variable label data includes the label identifier of the interface variable, the computing node number it belongs to, and its transmission status:

[0152] The label identifier of the interface variable is used to uniquely identify the variable; the computing node number to which the interface variable belongs is used to locate the storage and transmission nodes of the variable; the transmission status of the interface variable is used to indicate whether the variable has been transmitted, and the status values include "incomplete" and "complete".

[0153] Store the interface variable label data list in the verification data structure used in the global convergence process to provide data support for subsequent retrieval and status verification one by one.

[0154] During the global convergence calculation process, retrieve the interface variable label data one by one according to the parallel dependency verification algorithm, and verify whether the transmission status of each interface variable is the completed status:

[0155] During the global convergence calculation process, retrieve all interface variable label data lists one by one, and read information such as the label identifier, computing node number, and transmission status item by item.

[0156] Apply the parallel dependency verification algorithm to verify the dependency relationship of the interface variables and determine whether there are variables that have not been transmitted. When verifying the dependency relationship, consider the following rules:

[0157] If the dependent variable of the interface variable has not been transmitted yet, the current variable status is marked as "incomplete"; if the current interface variable has no incomplete dependencies, the status is marked as "complete".

[0158] Update the verification result to the label data list in real time for subsequent retransmission operations.

[0159] For the interface variable data whose transmission status is found to be incomplete during the retrieval process, record the computing node number it belongs to and the amount of data that has not been transmitted:

[0160] For the interface variables whose transmission status is found to be "incomplete" during the verification process, record the following information: the computing node number to which the interface variable belongs; the amount of data that has not been transmitted of the interface variable, and the amount of data is recorded in bytes; the label identifier of the interface variable for subsequent positioning of retransmission.

[0161] Store the recorded information in the global convergence incomplete data list as the basis for initiating a retransmission request.

[0162] Initiate a secondary transmission request to the corresponding computing node based on the recorded amount of incomplete transmitted data and the computing node number:

[0163] According to the records in the global convergence incomplete data list, send a secondary transmission request to each computing node. The request includes: the label identifier of the interface variable that needs to be resent; the amount of data with incomplete transmission; the number and memory address information of the receiving node.

[0164] After the computing node receives the secondary transmission request, extract the incomplete data and resend it to the target receiving node according to the specified memory address.

[0165] After receiving the secondary transmitted data, update the label data status of the corresponding interface variable to completed, and record the final transmission status of all interface variables in the convergence log file:

[0166] After receiving the secondary transmitted data, update the label data status of the corresponding interface variable, and modify the transmission status from "incomplete" to "completed".

[0167] Uniformly record the final transmission status of all interface variables to generate a global convergence log file. The log file contains the following content: the label identifier of each interface variable; the final transmission status; the timestamp of transmission completion.

[0168] Store the generated global convergence log file in the simulation system log management module to provide a basis for subsequent simulation performance analysis and result verification.

[0169] The above formulas are all dimensionless and take their numerical calculations. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the real situation. The preset parameters and threshold selection in the formulas are set by those skilled in the art according to the actual situation.

[0170] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0171] Those of ordinary skill in the art will realize that the modules and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled artisans can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0172] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and modules described above can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.

[0173] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces, and the indirect couplings or communication connections of the devices or modules can be in electrical, mechanical, or other forms.

[0174] The module described as a separation component may or may not be physically separated. The component shown as a module may or may not be a physical module. It may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0175] In addition, in each embodiment of this application, the functional modules can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.

[0176] If the above-mentioned function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0177] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

[0178] Finally: The above is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for optimizing the parallel performance of electromagnetic transient simulation based on the RoCE network, characterized in that It includes the following steps: S1: Analyze the configuration parameters of multi-level time steps, construct a list of sub-modules corresponding to the electromagnetic transient simulation model, and clarify the time step settings for each sub-module; S2: Based on the list of sub-modules, establish a cross-node simulation partition structure, and allocate the simulation sub-domain range of each node in the converged Ethernet Remote Direct Memory Access (RoCE) network; S3: Tag the interface variables of each simulation sub-domain, assign a unique tag identifier to each interface variable, and configure its corresponding memory area through the Remote Direct Memory Access protocol; S4: Generate a time step priority queue according to the time step differences of sub-modules, and set the bandwidth allocation policy corresponding to the time step priority in the lossless flow control mechanism of the RoCE network; S5: When the data of the sub-module with a high-frequency time step is updated, implement a hierarchical encapsulation delay for the window of the interface variable with a low-frequency time step; S6: Before the global convergence of the overall simulation is completed, apply a parallel dependency check algorithm to retrieve the tag data of the interface variables of all computing nodes one by one, and perform a secondary transmission on the interface variable data that has not been completed for transmission.

2. The electromagnetic transient simulation parallel performance optimization method based on the RoCE network according to claim 1, wherein Analyze the configuration parameters of multi-level time steps, construct a list of sub-modules corresponding to the electromagnetic transient simulation model, and clarify the time step settings for each sub-module, specifically including: Obtain the physical attributes and dynamic behavior characteristics of each module in the electromagnetic transient simulation model, and analyze the computing requirements of each module in combination with the network topology structure of the electromagnetic transient simulation model; Determine the time step range applicable to different modules according to the dynamic behavior characteristics and computing requirements of each module; Divide the electromagnetic transient simulation model into multiple sub-modules, generate a list corresponding to each sub-module, and mark the time step settings and associated input / output data types of each sub-module in the list.

3. The electromagnetic transient simulation parallel performance optimization method based on the RoCE network according to claim 2, wherein Based on the list of sub-modules, establish a cross-node simulation partition structure, and allocate the simulation sub-domain range of each node in the RoCE network, specifically including: According to the time step settings and input / output data types in the list of sub-modules, divide the simulation model into multiple computing regions, and determine the boundary interfaces of each computing region; Based on the network topology structure of the simulation model and the computing requirements of each computing region, allocate each computing region to a computing node; When allocating computing nodes, adjust the boundaries of the computing regions according to the computing complexity and data transmission requirements of each computing region to achieve load balancing; Number the allocated computing regions, and generate a partition structure file recording the node allocation information of the computing regions and the positions of the interface variables; Set the corresponding simulation sub-domain range for each computing node in the RoCE network, and the simulation sub-domain range includes the boundary interface variables of the computing region and their associated memory access permissions.

4. The electromagnetic transient simulation parallel performance optimization method based on the RoCE network according to claim 3, wherein Tag the interface variables of each simulation sub-domain, assign a unique tag identifier to each interface variable, and configure its corresponding memory area through the Remote Direct Memory Access protocol, specifically including: Obtain the boundary interface variables of each simulation sub-domain, and identify the unique characteristics of the interface variables based on the interface variable data types and boundary positions recorded in the simulation partition structure file; Assign a unique label identifier to each interface variable. The label identifier includes the simulation sub - domain number to which the interface variable belongs, the data type of the interface variable, and the memory offset address of the interface variable; Bind the label identifier of each interface variable to its corresponding memory area, and allocate read - write permissions to each memory area through the Remote Direct Memory Access (RDMA) protocol; When allocating read - write permissions, establish a memory address mapping table for the boundary interface variables of each simulation sub - domain. The memory address mapping table records the label identifier of the interface variable, the starting address of the memory area, and the read - write permission level; Upload the established memory address mapping table to the configuration file of the Remote Direct Memory Access (RDMA) protocol to complete the memory area configuration of the interface variable.

5. The electromagnetic transient simulation parallel performance optimization method based on the RoCE network according to claim 4, characterized in that Generate a time - step priority queue based on the time - step differences of sub - modules, and set the bandwidth allocation policy corresponding to the time - step priority in the lossless flow control mechanism of the RoCE network. Specifically, it includes: Obtain the time - step settings of sub - modules, classify the sub - modules according to the length of the time - step, assign higher priorities to sub - modules with shorter time - steps, and lower priorities to sub - modules with longer time - steps; Generate a time - step priority queue based on the priority classification results. The time - step priority queue contains the priority number of each sub - module and its corresponding time - step; In the lossless flow control mechanism of the RoCE network, allocate bandwidth quotas to each priority queue. The bandwidth quotas are calculated based on the priority number and the data transmission frequency; Adjust the parameters of the bandwidth allocation policy in the priority queue to ensure that the bandwidth ratio of the high - priority queue is higher than that of the low - priority queue and avoid resource competition; Record the bandwidth allocation information of the time - step priority queue in the lossless flow control mechanism of the RoCE network. The bandwidth allocation information includes the priority number, the corresponding bandwidth ratio, and the time - step range of the sub - module.

6. The electromagnetic transient simulation parallel performance optimization method based on the RoCE network according to claim 5, wherein When the data of the sub - module with a high - frequency time - step is updated, implement hierarchical encapsulation delay for the send window of the interface variables with a low - frequency time - step. Specifically, it includes: Detect the request status of the sub - module with a high - frequency time - step to update data, and obtain the list of boundary interface variables of the high - frequency time - step sub - module; Classify the interface variables with a low - frequency time - step, and divide the low - frequency interface variables into multiple priority queues according to the length of the low - frequency time - step and the data transmission priority; Set the encapsulation delay time for each priority queue. The encapsulation delay time is determined according to the level of the priority queue where the interface variable is located. The lower the priority level, the longer the encapsulation delay time; During the update period of the high - frequency time - step sub - module, adjust the send window of the low - frequency interface variables, and send the encapsulated low - frequency interface variables to the corresponding receiving nodes in batches; Record the encapsulation delay parameters of each send window adjustment, including the priority number of the low - frequency interface variable, the delay time, and the transmission timestamp.

7. The electromagnetic transient simulation parallel performance optimization method based on RoCE network according to claim 6, characterized in that Before the global convergence of the overall simulation is completed, apply the parallel dependency check algorithm to retrieve the label data of the interface variables of all computing nodes one by one, and perform secondary transmission on the interface variable data that has not been completed for transmission. Specifically, it includes: Obtain the list of interface variable label data for each computing node, where the interface variable label data includes the label identifier of the interface variable, the computing node number it belongs to, and its transmission status; During the global convergence calculation process, retrieve the interface variable label data one by one according to the parallel dependency verification algorithm, and verify whether the transmission status of each interface variable is the completed status; For the interface variable data whose transmission status is not completed found during the retrieval process, record the computing node number it belongs to and the amount of data not yet transmitted; Based on the recorded amount of data not yet transmitted and the computing node number, initiate a secondary sending request to the corresponding computing node; After receiving the secondary sent data, update the label data status of the corresponding interface variable to completed, and record the final transmission status of all interface variables in the convergence log file.

Citation Information

Patent Citations

  • Simulation method and simulation device for real-time electromagnetic transient state of power system

    CN116861768A

  • Method and device for joint simulation of electric power simulation software and user program

    CN118194735A