Simulation device for rail transit power supply system
By adopting ARM and FPGA parallel computing units in the urban rail transit power supply system simulation device and combining it with distributed simulation technology, efficient simulation calculations of the power supply system are achieved, solving the problem of insufficient simulation calculation speed and supporting the optimization and verification of rail transit network optimization solutions.
Patent Information
- Application Number
- CN202410318874.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-20
- Publication Date
- 2025-09-23
Smart Images

Figure CN120688420A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of simulation technology, and in particular to a simulation device for a rail transit power supply system. Background Art
[0002] With the rapid development of urban rail transit, the subway has become a transportation backbone due to its large capacity, energy saving, environmental protection, and high punctuality. According to statistics, China's urban rail transit passenger volume in 2022 was 19.4 billion, the number of passengers entering the station was 11.69 billion, and the passenger turnover volume was 156 billion passenger-kilometers. However, urban rail transit is still facing the significant problem of huge energy consumption. According to the data in the "Urban Rail Transit 2021 Annual Statistics and Analysis Report" released by the China Urban Rail Transit Association, the total electricity consumption of urban rail transit in 2021 was 21.31 billion kWh, a year-on-year increase of 23.6%, accounting for 2.56‰ of China's total electricity consumption in the same year. With the continuous increase in newly commissioned lines, the overall energy consumption index continues to grow. Therefore, for complex engineering systems like urban rail transit systems, comprehensive performance optimization, centered around improving energy efficiency, holds significant practical significance and technical research value. High-performance simulation models and environments are essential support tools for research into comprehensive performance optimization technologies for urban rail transit systems. This is primarily reflected in the following two aspects. 1) Supporting the search for optimization solutions: Due to the complexity and large scale of subway power supply systems, mathematical analytical models are difficult to establish. Therefore, iterative optimization using mathematical models is unsuitable for optimizing subway power supply system operation plans. Simulation models are required, and iterative optimization methods based on simulation are employed. 2) Supporting the verification of optimization solutions: Urban rail transit faces drawbacks such as high cost, low safety factors, and limited test coverage when conducting various field experiments. Simulation, as a non-physical experimental verification method, can overcome these challenges faced by field experimental verification of complex engineering systems and safely complete various experimental verifications at low cost. Since iterative optimization requires numerous simulation iterations, and simulation verification also requires simulating complex system-level optimization solutions within a limited timeframe, simulation system performance becomes a key factor. Therefore, establishing high-performance simulation models and environments for urban rail transit systems holds significant application and research value. The currently established urban rail transit system simulation environment mainly includes infrastructure models, train models, power supply system models, signal system models and dispatching system models. The power supply system model has the highest simulation accuracy requirements and is also the most time-consuming part. Therefore, how to achieve efficient simulation of the urban rail transit power supply system model has become an important issue that needs to be solved.
[0003] Currently, power system simulations are mostly performed using CPUs. However, the serial instructions and sequential operations used by conventional computers restrict the speed of simulation execution. When the simulation scale is large, the CPU's serial computing mode leads to slow calculations. Summary of the Invention
[0004] According to one aspect of the present disclosure, a rail transit power supply system simulation device is provided, the device comprising:
[0005] Train model simulation calculation module, used to simulate the train running status and send the train location information;
[0006] a power supply system proxy model simulation calculation module, connected to the train model simulation calculation module, for determining and sending a node conductance matrix and a node current matrix based on the location information, pre-configured circuit netlist information, and operating line information, wherein the node conductance matrix includes the conductances of multiple nodes, and the node current matrix includes the currents of multiple nodes;
[0007] a power supply system simulation calculation module, connected to the power supply system proxy model simulation calculation module, configured to perform adaptive calculations based on the node conductance matrix and the node current matrix to obtain calculation results, and send the calculation results to the power supply system proxy model simulation calculation module, wherein the adaptive calculations include parallel calculations, and the calculation results include a node voltage matrix, which includes voltages of multiple nodes;
[0008] The power supply system proxy model simulation calculation module is also used to send the calculation result to the train model simulation calculation module, so that the train model simulation calculation module updates the train operation status.
[0009] In a possible implementation, the power supply system simulation calculation module includes: an ARM-based network interface calculation unit and an FPGA-based parallel simulation calculation unit.
[0010] The network interface calculation unit is provided between the power supply system proxy model simulation calculation module and the parallel simulation calculation unit, and the network interface calculation unit is used to realize information exchange between the power supply system proxy model simulation calculation module and the parallel simulation calculation unit;
[0011] The parallel simulation calculation unit is used to perform adaptive calculations based on the node conductance matrix and the node current matrix to obtain calculation results, and the parallel calculations include parallel LU decomposition calculations.
[0012] In a possible implementation, the power supply system agent model simulation calculation module is connected to the network interface calculation unit via an Ethernet interface, wherein:
[0013] The node conductance matrix, node current matrix, and node voltage matrix stored in the power supply system proxy model simulation calculation module and the parallel simulation calculation unit are all floating-point data, wherein the power supply system proxy model simulation calculation module and the parallel simulation calculation unit are both used to: convert the data to be sent into character data and send it, and convert the received data into floating-point data and store it;
[0014] The network interface calculation unit is used to send at least one row of data of the node conductance matrix or the node current matrix to the parallel simulation calculation unit each time.
[0015] In one possible implementation, the network interface computing unit is configured to:
[0016] When detecting that the first sending flag is a first preset value, sending a line of data, setting the first sending flag to a second preset value, wherein the first sending flag indicates whether the network interface computing unit has completed sending a line of data;
[0017] If it is detected that the first receiving flag is a second preset value, the count value is increased by 1, wherein the first sending flag, the first receiving flag, and the initial value of the count value are all first preset values, and the first receiving flag indicates whether the parallel simulation calculation unit has completed receiving a row of data;
[0018] In a possible implementation, the parallel simulation computing unit is used to:
[0019] receiving the current row data sent by the network interface computing unit when detecting that the first sending flag is a second preset value;
[0020] If the current row of data is received, the first reception flag is set to a second preset value;
[0021] If the count value is less than the preset value, the first receiving flag is set to the first preset value.
[0022] In a possible implementation, the parallel simulation computing unit is further configured to:
[0023] When the count value reaches a preset value, stopping data reception;
[0024] An adaptive operation is performed according to the received node conductance matrix and node current matrix to obtain an operation result.
[0025] In a possible implementation, the parallel simulation computing unit is used to:
[0026] Sending the operation result, and setting a second sending flag bit to a second preset value, where the initial value of the second sending flag bit is the first preset value, and the second sending flag bit indicates whether the parallel simulation calculation unit has completed data sending;
[0027] In one possible implementation, the network interface computing unit is configured to:
[0028] receiving the operation result when detecting that the second sending flag is a second preset value;
[0029] When receiving the operation result, setting a second reception flag bit to a second preset value, where the initial value of the second reception flag bit is the first preset value, and the second reception flag bit indicates whether the network interface computing unit has completed receiving the data;
[0030] The calculation result is sent to the power supply system agent model simulation calculation module.
[0031] In a possible implementation, the parallel simulation computing unit is further configured to:
[0032] When it is detected that the second receiving flag is changed from the initial value to the second preset value, the second receiving flag and the second sending flag are both set to the first preset value, and the sending process of the operation result is ended.
[0033] In a possible implementation, the power supply system proxy model simulation calculation module includes:
[0034] Circuit simulation class, including simulation circuit structures composed of multiple passive component simulation models and multiple active component simulation models;
[0035] The rail transit line class includes a line information model and an update model. The line information model stores rail transit line information. The update model is used to update the simulation circuit structure and update the node conductance matrix and node current matrix.
[0036] In a possible implementation, the power supply system simulation calculation module includes at least one of a DC simulation type, an AC simulation type, and an AC / DC simulation type, wherein:
[0037] The DC simulation class includes a simulation initialization class and an equivalent capacitance class, which are used to assist simulation calculations, wherein the equivalent capacitance class is an equivalent model of a capacitor under a set simulation step size;
[0038] The AC simulation class is used to calculate the power flow equation on the AC side and obtain the voltage of the traction substation;
[0039] The AC / DC simulation class is used for iterative joint simulation of AC and DC.
[0040] In a possible implementation, the device further includes:
[0041] Signal system model simulation calculation module, used to simulate and model the signal transmission system of train operation and signal transmission in rail transit;
[0042] Infrastructure simulation calculation module, used to simulate and model the infrastructure in the rail transit system;
[0043] The dispatching system model simulation calculation module is used for simulation modeling of the dispatching system of the rail transit system;
[0044] Passenger flow simulation calculation module is used to simulate and model the passenger flow of rail transit systems.
[0045] The train model simulation calculation module is further configured to send the calculation result to the dispatching system model simulation calculation module so that the dispatching system model simulation calculation module updates the dispatching information of the train.
[0046] The embodiment of the present disclosure simulates the train running status through a train model simulation calculation module, sends the train's location information, determines and sends the node conductance matrix and the node current matrix through the power supply system simulation calculation module, performs parallel adaptive calculations based on the node conductance matrix and the node current matrix through the power supply system simulation calculation module to obtain calculation results, and uses the calculation results to update the train running status. The embodiment of the present disclosure realizes fast and efficient simulation calculations through parallel calculations, solves the problems of insufficient simulation calculation speed and insufficient CPU calculation parallelism caused by the increase in circuit scale and calculation amount in related technologies, and solves the system performance bottleneck problem caused by insufficient CPU calculation parallelism, thereby effectively supporting the optimization and verification of rail transit network optimization solutions.
[0047] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, rather than limiting the present disclosure. Other features and aspects of the present disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The accompanying drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.
[0049] Figure 1 A schematic diagram of a rail transit power supply system simulation device according to an embodiment of the present disclosure is shown.
[0050] Figure 2A schematic diagram of a rail transit power supply system simulation device according to an embodiment of the present disclosure is shown.
[0051] Figure 3-1 The diagram shows the process of LU decomposition by FPGA and data interaction with PC and ARM.
[0052] Figure 3-2 The simulation results of FPGA solving a 24-dimensional equation system.
[0053] Figure 4 A schematic diagram of a power supply system agent model simulation calculation module and a power supply system simulation calculation module according to an embodiment of the present disclosure is shown.
[0054] Figure 5-1 Shows the PC-side network communication interface program flow chart.
[0055] Figure 5-2 A schematic diagram showing how to set the IP address and port number of the network debugging tool.
[0056] Figure 5-3 A schematic diagram showing the test results on the PC side is shown. Figure 5-4 A schematic diagram showing the test results of the network debugging tool.
[0057] Figure 5-5 Shows the ARM-side network communication interface program flow chart.
[0058] Figure 5-6 A schematic diagram showing the connection to the COM5 serial port is shown.
[0059] Figure 5-7 A schematic diagram showing the test results on the ARM side is shown.
[0060] Figure 5-8 A schematic diagram of debugging the network debugging tool test end is shown.
[0061] Figure 5-9 A schematic diagram showing the data storage form and transmission form.
[0062] Figure 5-10 A schematic diagram of two-station-one-interval matrix A and matrix B is shown.
[0063] Figure 5-11 It shows a schematic diagram of the ARM end receiving the results of A and B.
[0064] Figure 5-12 A schematic diagram of matrix X with two stations and one interval is shown.
[0065] Figure 5-13 It shows a schematic diagram of receiving X on the PC side.
[0066] Figure 5-14Shows the architectural diagram of FPGA and ARM.
[0067] Figure 5-15 A schematic diagram of using Vivado to build a system hardware platform is shown.
[0068] Figure 5-16 A schematic diagram of parameter configuration is shown.
[0069] Figure 5-17 A schematic diagram of custom IP parameter settings is shown.
[0070] Figure 5-18 A schematic diagram showing the number of master and slave interfaces is shown.
[0071] Figure 5-19 Shows a schematic diagram of module parameter settings.
[0072] Figure 5-20 Shows a schematic diagram of module parameter settings.
[0073] Figure 5-21 A schematic diagram of the address mapping between Master and Slave is shown.
[0074] Figure 5-22 Shows a schematic diagram of the ARM data transmission program flow.
[0075] Figure 5-23 Shows the ARM data receiving program flow chart.
[0076] Figure 5-24 Shows a schematic diagram of ARM and FPGA data reception and transmission verification.
[0077] Figure 6-1 A schematic diagram of the least squares fitting method for solving small-scale equations is shown.
[0078] Figure 6-2 A schematic diagram of the least squares fitting method for large-scale equation solution time is shown.
[0079] Figure 6-3 A schematic diagram of quadratic fitting of large-scale equation solution time is shown. DETAILED DESCRIPTION
[0080] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.
[0081] In the description of the present disclosure, it should be understood that the terms "length", "width", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, and are only for the convenience of describing the present disclosure and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present disclosure.
[0082] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. Throughout the present disclosure, "plurality" means two or more, unless otherwise specifically defined.
[0083] In this disclosure, unless otherwise expressly specified or limited, terms such as "mounted," "connected," "connect," and "fixed" should be understood broadly. For example, they may refer to fixed connections, detachable connections, or integration; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal connections between two components or interactions between two components. Those skilled in the art will understand the specific meanings of these terms in this disclosure based on specific circumstances.
[0084] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.
[0085] The term "and / or" herein simply describes an association relationship between associated objects, indicating that three relationships can exist. For example, "A and / or B" can represent the existence of three situations: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" herein refers to any combination of at least two of any one or more of a plurality of items. For example, "at least one of A, B, and C" can represent any one or more elements selected from the set consisting of A, B, and C.
[0086] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.
[0087] See also Figure 1, Figure 1 A schematic diagram of a rail transit power supply system simulation device according to an embodiment of the present disclosure is shown.
[0088] like Figure 1 As shown, the device includes:
[0089] The train model simulation calculation module 10 is used to simulate the train running state and send the train position information;
[0090] a power supply system proxy model simulation calculation module 20, connected to the train model simulation calculation module 10, for determining and sending a node conductance matrix and a node current matrix based on the location information, pre-configured circuit netlist information, and operating line information, wherein the node conductance matrix includes the conductances of multiple nodes, and the node current matrix includes the currents of multiple nodes;
[0091] a power supply system simulation calculation module 30, connected to the power supply system proxy model simulation calculation module 20, configured to perform adaptive calculations based on the node conductance matrix and the node current matrix to obtain calculation results, and send the calculation results to the power supply system proxy model simulation calculation module 20, wherein the adaptive calculations include parallel calculations, and the calculation results include a node voltage matrix, which includes voltages of multiple nodes;
[0092] The power supply system proxy model simulation calculation module 20 is further used to send the calculation result to the train model simulation calculation module 10 so that the train model simulation calculation module 10 updates the train operation status.
[0093] The embodiment of the present disclosure simulates the train running status through a train model simulation calculation module, sends the train's location information, determines and sends the node conductance matrix and the node current matrix through the power supply system simulation calculation module, performs parallel adaptive calculations based on the node conductance matrix and the node current matrix through the power supply system simulation calculation module to obtain calculation results, and uses the calculation results to update the train running status. The embodiment of the present disclosure realizes fast and efficient simulation calculations through parallel calculations, solves the problems of insufficient simulation calculation speed and insufficient CPU calculation parallelism caused by the increase in circuit scale and calculation amount in related technologies, and solves the system performance bottleneck problem caused by insufficient CPU calculation parallelism, thereby effectively supporting the optimization and verification of rail transit network optimization solutions.
[0094] The embodiments of the present disclosure do not limit the types of specific nodes in the node conductance matrix, the node current matrix, and the node voltage matrix. A node can be a specific device in a circuit (such as a resistor, a capacitor, etc.), or a section of wire, or a module of a circuit. Those skilled in the art can determine the meaning of a node based on actual conditions and needs.
[0095] The embodiments of the present disclosure do not limit the specific types of rail transit. Rail transit may include subways, light rails, motor trains, high-speed rails, etc. Those skilled in the art may apply the technical solutions of the embodiments of the present disclosure to any rail transit system according to actual conditions and needs.
[0096] The embodiments of the present disclosure do not limit the specific implementation methods of the train model simulation calculation module 10, the power supply system proxy model simulation calculation module 20, and the power supply system simulation calculation module 30. Those skilled in the art can adopt appropriate technical means to implement them according to actual conditions and needs.
[0097] The embodiments of the present disclosure do not limit the specific types of train running status. Those skilled in the art can set it according to actual conditions and needs. For example, the train running status may include acceleration state, deceleration state, uniform speed running state, temporary parking state, stopped running state, etc.
[0098] The embodiments of the present disclosure do not limit the specific types of the location information, pre-configured circuit netlist information, and operation line information. Those skilled in the art can set them according to actual conditions and needs. For example, the operation line information may include the stations passed by the train during operation, the location information may be the position of the train on the operation line, which can be obtained through a positioning component (such as GPS, etc.), and the circuit netlist information is the netlist information of the rail transit power supply system that supplies power to the train, which may include the parameters of all components such as voltage source (traction substation), capacitance, conductivity, and connection relationship, such as R12, which represents the resistance value between node 1 and node 2. The embodiment of the present disclosure does not limit the setting method of the parameters in the circuit netlist information. For example, it can be set according to the parameters provided by the National High-speed Train Technology Innovation Center.
[0099] The present embodiment of the present disclosure does not limit the specific implementation method of the power supply system proxy model simulation calculation module 20 determining and sending the node conductance matrix and node current matrix based on the position information, pre-configured circuit netlist information, and operating line information. Those skilled in the art can refer to relevant technologies for implementation according to actual conditions and needs. The node conductance matrix and node current matrix determined in the present embodiment of the present disclosure are correlated with the position information and operating line information of the train. The node conductance matrix and node current matrix determined based on the position information, pre-configured circuit netlist information, and operating line information can represent the conductance and current of each node at a specific location on a specific operating line. In this way, the power supply system simulation calculation module 30 performs adaptive calculations based on the node conductance matrix and node current matrix to obtain a node voltage matrix including the voltages of each node at a specific location on the specific operating line as the calculation result.
[0100] The following is an exemplary introduction to the preferred implementation of the power supply system simulation calculation module 30.
[0101] See also Figure 2 , Figure 2 A schematic diagram of a rail transit power supply system simulation device according to an embodiment of the present disclosure is shown.
[0102] In one possible implementation, Figure 2 As shown, the power supply system simulation calculation module 30 includes: an ARM-based network interface calculation unit 310 and an FPGA-based parallel simulation calculation unit 320,
[0103] The network interface calculation unit 310 is provided between the power supply system proxy model simulation calculation module 20 and the parallel simulation calculation unit 320, and the network interface calculation unit 310 is used to realize information exchange between the power supply system proxy model simulation calculation module 20 and the parallel simulation calculation unit 320;
[0104] The parallel simulation calculation unit 320 is used to perform adaptive calculations based on the node conductance matrix and the node current matrix to obtain calculation results. The parallel calculations include parallel LU decomposition calculations, which will be introduced below.
[0105] The disclosed embodiments utilize the parallel computing and high integration characteristics of FPGAs to achieve simulation acceleration and high efficiency for power supply systems, specifically: 1) Parallel computing: By parallelizing the LU decomposition calculation tasks, the parallel computing capabilities of FPGAs are utilized to achieve rapid solution of the equation set, thereby improving simulation efficiency; 2) High integration: FPGAs integrate a large number of programmable logic blocks and programmable interconnection networks, while combining them with ARM processor cores to form a highly integrated hardware platform. This high integration allows FPGAs to implement multiple functions on a single chip, including processing complex simulation calculation tasks and efficient communication and data exchange with other modules. These features enable FPGAs to achieve rapid solution of equation sets and power supply system simulation, which can, to a certain extent, solve the problems of insufficient simulation calculation speed and insufficient CPU calculation parallelism that exist in the previously constructed subway power supply system simulation platform as the circuit scale and calculation amount increase, and the system performance bottleneck caused by insufficient CPU calculation parallelism, thereby effectively supporting the optimization and verification of rail transit network optimization solutions.
[0106] Of course, the hardware architecture for implementing parallel computing can also be other hardware architectures besides FPGA. The embodiments of the present disclosure do not limit this. Those skilled in the art can adopt a suitable hardware architecture that can implement parallel high-speed computing according to actual conditions and needs.
[0107] The disclosed embodiments adopt the characteristics of distributed computing to improve the efficiency of computing. A distributed system refers to a system composed of multiple computers that cooperate with each other, which are connected through a network and work together. Generally speaking, a distributed system consists of multiple components, each of which can perform specific tasks and communicate and coordinate with other components through the network. This distributed structure can improve the reliability, availability and performance of the system, and allow the system to be expanded to support more users or larger workloads. It is widely used in the fields of the Internet, cloud computing, big data, artificial intelligence, etc. Distributed simulation technology is a technology that speeds up simulation by allocating simulation tasks to multiple computer components on the basis of a distributed system. This technology can improve simulation efficiency and expand the scale of simulation.
[0108] In one possible implementation, Figure 2 As shown, the device may further include:
[0109] The signal system model simulation calculation module 40 is used to simulate and model the signal transmission system of train operation and signal transmission in rail transportation;
[0110] An infrastructure simulation calculation module 50 is used to simulate and model the infrastructure in the rail transit system;
[0111] The dispatching system model simulation calculation module 60 is used for simulation modeling of the dispatching system of the rail transit system;
[0112] The passenger flow simulation calculation module 70 is used to simulate and model the passenger flow of the rail transit system.
[0113] In a possible implementation, the train model simulation calculation module 10 is further configured to send the calculation result to the dispatching system model simulation calculation module 60 so that the dispatching system model simulation calculation module 60 updates the dispatching information of the train.
[0114] Table 1 is used below to provide an exemplary introduction to the signal system model simulation calculation module 40, the infrastructure simulation calculation module 50, the dispatching system model simulation calculation module 60, and the passenger flow simulation calculation module 70. Of course, the embodiments of the present disclosure do not limit the specific implementation methods of the signal system model simulation calculation module 40, the infrastructure simulation calculation module 50, the dispatching system model simulation calculation module 60, and the passenger flow simulation calculation module 70. Those skilled in the art can adopt relevant technologies to implement them according to actual conditions and needs.
[0115] Table 1
[0116]
[0117] Exemplarily, the power supply system proxy model simulation calculation module 20 can run on a terminal device or server. The terminal device can include, for example, a laptop, a desktop computer (hereinafter referred to as a PC), or other devices. The power supply system proxy model simulation calculation module 20 establishes a node voltage equation by reading in the circuit netlist (including all component parameters such as the voltage source (traction substation), capacitance, and conductance, and indicating the connection relationship, such as R12, representing the resistance between node 1 and node 2), subway line and train location information, thereby obtaining node matrix A (the node conductance matrix) and matrix B (node current matrix), which contain the current and conductance of each node at different locations. The location information can be used to calculate the equivalent resistance value between stations, calculated as: resistance per unit distance * distance. In order to hand over the computationally intensive solution work to the FPGA, the ARM-based distributed simulation system network interface is used to transmit matrix A and matrix B, and the calculated voltage matrix X (the node voltage matrix) is transmitted back to the PC. The TCP communication protocol is used for communication between the PC and the ARM. The FPGA-based parallel simulation computing unit 320 completes LU decomposition parallel computing and pipeline parallel technology, and exchanges data with ARM through the AXI4-Lite bus.
[0118] As an example, to implement data transmission and parallel computing processes, the Zynq hardware platform can be used for simulation acceleration. Zynq is a SoC series launched by Xilinx that integrates FPGAs and ARM processors, while providing many programmable interfaces and peripherals, allowing developers to implement a variety of different functions on a single chip. This series of chips has two parts: programmable logic and processor system. The programmable logic part (PL) contains a large number of FPGA logic blocks and programmable interconnect networks, while the processor system (PS) includes processor cores such as ARM Cortex-A9 or Cortex-A53. The two are connected via the AXI (Advanced eXtensible Interface) bus, enabling efficient data transmission and control.
[0119] The following is an exemplary introduction to the parallel LU decomposition operation performed by the parallel simulation calculation unit 320.
[0120] LU Decomposition is a matrix decomposition method in linear algebra with a wide range of applications, including solving linear equations, calculating determinants, and inverting matrices. It decomposes a matrix A into the product of a lower triangular matrix L and an upper triangular matrix U, that is, A=LU. Among them, the diagonal elements of the lower triangular matrix L are all 1, and the other elements are in lower triangular form. The diagonal elements of the upper triangular matrix U are the same as A, and the other elements are in upper triangular form. The advantage of this decomposition is that if a technician wants to solve the equation system Ax=b, he can solve x by first solving Ly=b and then solving Ux=y. If b changes, the technician only needs to re-solve y and x without re-decomposing A.
[0121] For example, the LU decomposition can be calculated using Gaussian-Jordan elimination. Specifically, the technician can record the upper triangular matrix obtained in each elimination step in the Gaussian-Jordan elimination method, and the resulting matrix is the U matrix; the lower triangular matrix L can be calculated from the product of each elimination step during the elimination process.
[0122] For example, the FPGA-based power supply system simulation platform mainly consists of a train model simulation calculation module 10 and a circuit calculation function module (a power supply system proxy model simulation calculation module 20 and a power supply system simulation calculation module 30). Its working process can be divided into the following steps:
[0123] First, the train model simulation calculation module 10 of the simulation platform is responsible for receiving signals from the time synchronization system, data from other simulation modules, and information such as circuit parameters stored in the database. For power supply system simulation, the information from the train module mainly includes the train's power, position, and direction of travel.
[0124] Next, the train model simulation calculation module 10 converts the received train information into corresponding events and stores these events in the event queue;
[0125] When the train model simulation calculation module 10 receives the time advance signal, it starts processing the events in the event queue. This step involves configuring the circuit calculation modules (power supply system proxy model simulation calculation module 20 and power supply system simulation calculation module 30) to perform related calculation tasks, such as solving equations and performing matrix operations.
[0126] After the circuit calculation module completes the calculation, it sends the result back to the train model simulation calculation module 10 through the internal channel; finally, the train model simulation calculation module 10 summarizes all internal data and outputs the simulation result.
[0127] This hardware platform system is designed based on a single-chip SoC (System on Chip) architecture that integrates FPGA and ARM. In this architecture, the FPGA performs logic processing and parallel data processing, primarily responsible for high-speed and real-time data processing. The FPGA's software consists of multiple modules, such as parallel computing modules, data sampling modules, and feature extraction modules, which together enable rapid data processing. The ARM component is responsible for data transmission, multi-task scheduling, and processing tasks with high real-time requirements, making it suitable for application scenarios requiring low latency and determinism. Within the chip, the FPGA and ARM synchronize data through internal channels to ensure coordinated and efficient data processing.
[0128] When using matrix inversion to solve linear equations, the computational complexity and resource requirements increase cubically with the order of the equations, resulting in increased computation time. Considering that the computational complexity and storage requirements of LU decomposition are far less than those of directly solving the inverse matrix, using the LU decomposition coefficient matrix algorithm for solving linear equations in FPGAs can achieve rapid solution.
[0129] Figure 3-1 The diagram shows the process of LU decomposition by FPGA and data interaction with PC and ARM.
[0130] For example, Figure 3-1 As shown in the figure, first, the PC generates matrices A and B based on information such as the circuit netlist and sends them to the ARM via the TCP protocol. The ARM, as the server, receives matrices A and B and writes the matrix information to the FPGA register via the AXI4-Lite bus, sending one row at a time. After the FPGA stores the received data in the memory of the solver module, it starts the LU operation and writes the result to the register via the AXI4-Lite bus. The ARM reads the register on the AXI4-Lite bus through the built-in library function to obtain the matrix X, which is then sent to the PC via the TCP protocol. After receiving the matrix X, the PC exchanges data with the train model computing node, completing one operation.
[0131] Figure 3-2 The simulation results of FPGA solving a 24-dimensional equation system.
[0132] Depend on Figure 3-2 It can be seen that the calculation starts at
[0133] t start =1.220470μs Formula 3-1
[0134] The calculation ends at
[0135] t end=36.200000μs Formula 3-2
[0136] Therefore, the time required to complete an LU decomposition operation is
[0137] t LU =t end -t start =36.200000μs-1.220470μs=34.97953μs Formula 3-3
[0138] For example, the PC operating system is Windows, the processor is an Intel(R) Core(TM) i7-8565U CPU @ 1.80GHz - 1.99GHz, and the computer memory is 8GB. Because the solution time is affected by the CPU's operating status, to obtain a more accurate solution time, multiple measurements are needed to calculate the average and reduce errors. The test was repeated 10 times, with the duration of each measurement shown in Table 3-2.
[0139] Table 3-2 Time for solving linear equations
[0140]
[0141] The average time to solve the linear equations is
[0142] t average =519.7μs Formula 3-4
[0143] Define the FPGA parallel computing acceleration as follows, T parallel_calculate T is the time spent on parallel computation of LU decomposition using FPGA. serial_calculate The time used for serial calculation using the CPU.
[0144]
[0145] From formula (3-3) and formula (3-4), it can be seen that the FPGA parallel computing acceleration ratio achieved by the embodiment of the present disclosure is
[0146]
[0147] It can be seen that the embodiment of the present disclosure uses FPGA to implement LU decomposition parallel computing to achieve an acceleration ratio of 14.86, that is, the time required to solve the linear equation system on the PC side is reduced by 14.86 times.
[0148] See also Figure 4 , Figure 4 A schematic diagram of a power supply system agent model simulation calculation module and a power supply system simulation calculation module according to an embodiment of the present disclosure is shown.
[0149] In one possible implementation, Figure 4 As shown, the power supply system agent model simulation calculation module 20 may include:
[0150] Circuit simulation class, including simulation circuit structures composed of multiple passive component simulation models and multiple active component simulation models;
[0151] The rail transit line class includes a line information model and an update model. The line information model stores rail transit line information. The update model is used to update the simulation circuit structure and update the node conductance matrix and node current matrix.
[0152] Exemplarily, the update model in the rail transit line class can use the calculation results (including the node voltage matrix) to update the simulation circuit structure (such as the information in the circuit class, including changes in resistance values, increases and decreases in resistance), and update the node conductance matrix and node current matrix. It can also use external input parameters such as power information and position information to update the simulation circuit structure and update the node conductance matrix and node current matrix.
[0153] The "class" described in the embodiments of the present disclosure may refer to a simulation model. The embodiments of the present disclosure do not limit the specific implementation methods of the circuit simulation class and the rail transit line class, and those skilled in the art may set them according to actual conditions and needs.
[0154] In one possible implementation, Figure 4 As shown, the power supply system simulation calculation module 30 may include at least one of the simulation types such as DC simulation, AC simulation, AC / DC simulation, etc., wherein:
[0155] The DC simulation class includes a simulation initialization class and an equivalent capacitance class, which are used to assist simulation calculations, wherein the equivalent capacitance class is an equivalent model of a capacitor under a set simulation step size;
[0156] The AC simulation class is used to calculate the power flow equation on the AC side and obtain the voltage of the traction substation;
[0157] The AC / DC simulation class is used for iterative joint simulation of AC and DC.
[0158] The embodiments of the present disclosure do not limit the specific implementation methods of the DC simulation class, the AC simulation class, and the AC / DC simulation class. Those skilled in the art can configure them according to actual conditions and needs.
[0159] Exemplarily, "performing adaptive calculations based on the node conductance matrix and the node current matrix to obtain calculation results" can be completed in the "DC simulation class, AC simulation class, and AC / DC simulation class". For example, the DC simulation class can establish a nonlinear circuit equation based on the circuit structure in the circuit simulation class, and then solve the equation according to the LU decomposition method and / or other methods (such as the Newton iteration method) to obtain the node voltage matrix X (including the voltage value of each node and the current value of the voltage source). For example, the AC simulation class calculates the power flow equation on the AC side to obtain the voltage of the traction substation on the AC side. This voltage is converted into the DC side voltage of the traction substation in the AC / DC simulation class, and is used as the voltage value of the traction substation voltage source in the DC simulation class to participate in the solution of the nonlinear circuit equation on the DC side. Among them, the traction substation is modeled as a voltage source and an equivalent resistor in the DC simulation class. The solved node voltage matrix X contains the current value of the voltage source, and the power is then calculated. For example, the AC / DC simulation class can be a combination of the DC simulation class and the AC simulation class, and some connections between the two types of alternating calculations are added, that is, the AC side voltage of the traction substation is converted into the DC side voltage for the DC simulation class, and the DC side power of the traction substation is converted into the AC side power for the AC simulation class.
[0160] For example, each simulation class can receive simulation parameters such as step size, simulation time, etc. from the outside and then perform simulation.
[0161] Exemplarily, the DC simulation class reads the circuit information in the circuit class, establishes the node voltage equation, and uses the solution results to update the current and voltage values in the circuit class. The DC simulation class includes a simulation initialization class and an equivalent capacitance class to assist in simulation calculations, where the equivalent capacitance class is an equivalent model of a capacitor under a set simulation step size. For example, by setting the DC simulation class in an FPGA and utilizing the FPGA's high-speed, parallel computing characteristics, the FPGA uses the DC simulation class to perform parallel LU settlement on the received node current matrix and node conductance matrix, and thus quickly obtain the node voltage matrix.
[0162] For example, the AC simulation class can be used to calculate the AC-side power flow equation and obtain the traction substation voltage. The AC / DC simulation class can implement iterative joint simulation of AC and DC. The present disclosure does not limit the specific implementation methods and functions of the AC and AC / DC simulation classes. Those skilled in the art can configure them based on actual circumstances and needs.
[0163] The disclosed embodiments use FPGA parallel LU computing for acceleration. For example, simulation computing tasks can be ported to a distributed system based on the Zynq platform. This platform integrates ARM processors and FPGAs, providing an efficient and flexible hardware acceleration environment for simulation. Compared to traditional simulation methods that rely solely on CPUs or GPUs for computing, the disclosed embodiments leverage the parallel processing capabilities of FPGAs and the software programmability of ARM to significantly improve simulation execution efficiency and data processing capabilities.
[0164] For example, the simulation class can update simulation results based on the input node voltage equations. The simulation results include the voltage vector X at each moment. Assuming that the train power and position are updated every 0.1s, the circuit needs to be updated every 0.1s. First, it can be determined whether the train is newly added or should be deleted. If so, the dimension of the node voltage equation has changed, and the node voltage equation needs to be re-established based on the modified circuit information and the circuit needs to be initialized. Otherwise, the existing node voltage equation and circuit information can be modified. Then, a nonlinear solver is used to simulate the current 0.1s duration.
[0165] Exemplarily, the subway line class and circuit class related programs in the embodiment of the present disclosure can be run on the PC side (that is, the power supply system agent model simulation calculation module 20 is set on the PC side), and the simulation class such as the AC / DC simulation class can be run on the FPGA (that is, the power supply system simulation calculation module 30 is set on the FPGA), and the data exchange of matrix A, matrix B, and matrix X is completed between the two through ARM. In the process of PC and ARM using the TCP protocol to complete data communication, PC acts as Client and ARM acts as Server. Use Socket to build a PC-side network communication interface, use LwIP to build an ARM-side network communication interface, and use network debugging tools to verify the correctness of the interface function respectively, and finally complete the data transmission and reception of matrix A, matrix B, and matrix X by PC and ARM. The embodiment of the present disclosure uses Zynq as the hardware platform, which integrates FPGA and ARM processors, and completes the data exchange between FPGA and ARM through the AXI4-Lite bus, with ARM as Master and FPGA as Slave.
[0166] Of course, the embodiment of the present disclosure may also adopt a method of direct communication between the FPGA and the PC, which is not limited by the embodiment of the present disclosure. Of course, in order to achieve more efficient and stable communication between the PPGA and the PC, the embodiment of the present disclosure sets up an ARM-based network interface computing unit 310. A notable feature of the Zynq platform is that it integrates an ARM processor and an FPGA. This architectural design provides developers with a flexible and efficient working environment. Compared to writing complex code directly on the FPGA for communicating with the PC, using the mature communication library already available on the ARM processor to realize data transmission between the FPGA and the PC can not only simplify the development process, but also greatly shorten the development cycle. At the same time, these library functions have been extensively tested and optimized, and can provide stable and efficient communication capabilities.
[0167] The following is an example of possible implementation methods for building a PC-side network communication interface based on Socket.
[0168] Figure 5-1 Shows the PC-side network communication interface program flow chart.
[0169] like Figure 5-1 As shown in the figure, the construction of the PC-side network communication interface based on Socket includes the following processes:
[0170] (1) Create a socket: First, you need to create a socket, specify the protocol type as TCP, and assign a unique identifier to the socket to identify the socket on the network.
[0171] (2) Request connection: For the client, it is necessary to connect to the remote server through Socket and specify the server's IP address and port number.
[0172] (3) Data Exchange: After the connection is established, the client can send data to the server through the socket. When sending data, the data must be packaged into packets and the recipient's IP address and port number must be specified. The receiver can receive data sent by the sender through the socket and must listen on the socket to receive data. After receiving the data, it must unpack the data and restore it to its original format.
[0173] (4) Disconnect: When the data transmission is completed, the socket connection needs to be closed. For both the client and the server, the connection can be closed by calling the close() function.
[0174] After the setup is complete, you can use the network debugging tool to verify whether the interface is set up correctly.
[0175] Figure 5-2A schematic diagram showing how to set the IP address and port number of the network debugging tool.
[0176] like Figure 5-2 As shown, set the network debugging tool to server (Sever), IP address to 192.168.1.106, port number to 10001, and click "Open".
[0177] Figure 5-3 A schematic diagram showing the test results on the PC side is shown. Figure 5-4 A schematic diagram showing the test results of the network debugging tool.
[0178] like Figure 5-3 As shown, run the PC program, the interface displays "Initialization successful, server connection successful, please enter the message to send", then enter "Hello, _I_am_Client." on the PC. Figure 5-4 You can see on the network debugging tool interface that the data sent by the PC has been received. Then the network debugging tool sends "Hello, _I_am_Server." to the PC. Figure 5-3 Similarly, you can see that the PC received information from the network debugging tool. Similarly, both devices sent arrays and successfully received array information from each other. The test results show that the socket-based PC network communication interface was successfully established.
[0179] The following is an example introduction to the possible implementation methods of building an ARM-side network communication interface based on the LwIP library.
[0180] Figure 5-5 Shows the ARM-side network communication interface program flow chart.
[0181] Figure 5-6 A schematic diagram showing the connection to the COM5 serial port is shown.
[0182] Figure 5-7 A schematic diagram showing the test results on the ARM side is shown.
[0183] Figure 5-8 A schematic diagram of debugging the network debugging tool test end is shown.
[0184] For example, using Figure 5-5 The process shown can easily implement the TCP protocol server on Zynq. After setting up the server on the ARM side, power on the development board and connect to the COM5 port, as shown in the following example: Figure 5-6 . Run the program, the ARM side results are as follows Figure 5-7 As can be seen from the figure, the ARM host address is 192.168.1.10 and the port number is 7.
[0185] Open the network debugging tool, connect to the remote host address 192.168.1.10, the remote host port number is 7, click "Open", and the connection is successful. Figure 5-8 Enter the received data description in the input box. After sending, the ARM server will send an array represented in floating-point hexadecimal format to the network debugging tool. It can be seen that the data is the same as the expected data, proving that the ARM server sent the data correctly.
[0186] The following is an example introduction to the network communication implementation between the PC and ARM terminals.
[0187] As previously described, the disclosed embodiments can implement a network communication interface between a PC and an ARM device, with the PC acting as the client and the ARM acting as the server. Functional testing can also be performed using network debugging tools. The following provides an exemplary introduction to implementing communication between a PC and an ARM device.
[0188] By implementing communication between PC and ARM, at least two functions can be achieved between PC and ARM:
[0189] (1) PC sends two matrices A and B to ARM;
[0190] (2) ARM sends to PC Send matrix X to ARM.
[0191] Figure 5-9 A schematic diagram showing the data storage form and transmission form.
[0192] For example, Figure 5-9 It can be seen that the storage format of matrix A and matrix B in PC and ARM are both floating point types, and they need to be converted into character types during transmission.
[0193] In a possible implementation, see also Figure 2 ,like Figure 2 As shown, the power supply system agent model simulation calculation module 20 is connected to the network interface calculation unit 310 through an Ethernet interface.
[0194] For example, Figure 5-9 As shown, the node conductance matrix A, node current matrix B, and node voltage matrix C stored in the power supply system proxy model simulation calculation module 20 and the parallel simulation calculation unit 320 are all floating-point data, wherein the power supply system proxy model simulation calculation module 20 and the parallel simulation calculation unit 320 are both used to: convert the data to be sent into character data and send it, and convert the received data into floating-point data and store it;
[0195] The network interface calculation unit 310 is configured to send at least one row of data of the node conductance matrix or the node current matrix to the parallel simulation calculation unit 320 each time.
[0196] Figure 5-10 A schematic diagram of two-station-one-interval matrix A and matrix B is shown.
[0197] For example, the PC side generates the node susceptance matrix A (24×24) and the current matrix B (24×1) generated by two stations and one interval by establishing the node voltage equation, as shown in FIG. Figure 5-10 When sending matrices from the PC to the ARM, type conversion is performed to convert Matrix A and Matrix B from floating point types to character types to facilitate data transmission.
[0198] Figure 5-11 It shows a schematic diagram of the ARM end receiving the results of A and B.
[0199] Figure 5-12 A schematic diagram of matrix X with two stations and one interval is shown.
[0200] Figure 5-13 It shows a schematic diagram of receiving X on the PC side.
[0201] For example, when receiving data, the ARM client first copies the data into a character buffer. This process uses the malloc function to dynamically allocate memory, preventing incomplete data reception errors caused by insufficient static memory allocation. The received buffer is then released, and the data in the buffer is finally parsed into floating-point data, completing the data reception. Figure 5-11 The figure shows the result of receiving data from ARM. Due to the limited space, only part of the result is captured. Figure 5-10 The sent data is consistent, indicating that the data reception on the ARM side is normal.
[0202] For example, the voltage matrix X (24×1) calculated from two stations and one interval is as follows: Figure 5-12 When the ARM side sends the matrix to the PC side, it uses the library function tcp_write(pcb,(constchar*)X,sizeof(X),1) to send the data, and uses the type coercion method to convert the matrix X into a floating point type and send it.
[0203] For example, when the PC receives the matrix X returned by the ARM, it first applies for a buffer with a floating point data type. Then it uses the reinterpret_cast type conversion operator to convert the received character data into a floating point type and store it in the buffer. The result of receiving the data on the PC is as follows: Figure 5-13 As shown by Figure 5-13 It can be seen that receiving data and Figure 5-12The sent data is consistent, indicating that the PC side receives the data normally.
[0204] The following is an example introduction to possible implementation methods of FPGA and ARM communication based on the AXI4-Lite bus protocol.
[0205] Figure 5-14 Shows the architectural diagram of FPGA and ARM.
[0206] For example, the IP of the FPGA part is interconnected to the AXI bus, and the ARM CPU is also interconnected to the AXI bus, so that the FPGA and ARM can exchange data.
[0207] Flexible use of AXI-4 bus technology can complete data exchange, allowing technicians to achieve efficient, high-speed and standardized advantages in building powerful FPGA internal bus data interconnection and communication.
[0208] like Figure 5-14 As shown, the ARM side (ARM-based network interface computing unit 310) communicates with the FPGA side (FPGA-based parallel simulation computing unit 320) via the AXI4-Lite bus, enabling debugging using an ILA (Integrated Logic Analyzer). Furthermore, the ARM side connects to the PC side via a UART-to-USB chip for bare-metal operation. Of course, the ARM side can also be connected to DDR memory.
[0209] Figure 5-15 A schematic diagram of using Vivado to build a system hardware platform is shown.
[0210] Figure 5-16 A schematic diagram of parameter configuration is shown.
[0211] Figure 5-17 A schematic diagram of custom IP parameter settings is shown.
[0212] Figure 5-18 A schematic diagram showing the number of master and slave interfaces is shown.
[0213] Figure 5-19 Shows a schematic diagram of module parameter settings.
[0214] Figure 5-20 Shows a schematic diagram of module parameter settings.
[0215] like Figure 5-16 As shown, the Zynq UltraScale+ MPSoC module is an ARM-based module, allowing you to write software programs to flexibly process data. Click "Presets", click "Apply Configuration...", select the configured .tcl file, and configure the user parameters for the chip.
[0216] Figure 5-15 The myip_v1.0 module in the project is the FPGA side, and you can set up your own IP to complete specific functions according to project requirements. The IP parameters set in this project are as follows Figure 5-17 As shown, select Lite for the interface type, Slave for the interface mode, and ARM as the Master. Since the transmitted data is floating point, select 32 bits for the data width.
[0217] The following test targets two stations and one section. Matrix A has a dimension of 24×24, Matrix B has a dimension of 24×1, and Matrix X has a dimension of 24×1. There are 60 registers, of which reg0-reg23 are used to send Matrix A and Matrix B, reg24-reg47 are used to receive the matrix, and the remaining registers can be used to set flags. For the specific data transmission and reception process, please refer to the previous description.
[0218] The two ends of the AXI bus protocol can be divided into the master and the slave, and they generally need to be connected through an AXI Interconnect (an IP core), which provides a switching mechanism for connecting one or more AXI master devices to one or more AXI slave devices. Specifically, when there are multiple hosts and slaves, AXIInterconnect is responsible for connecting and managing them. Because AXI supports out-of-order transmission, out-of-order transmission requires the support of the host's ID signal, and the IDs sent by different hosts may be the same, and AXI Interconnect solves this problem. It will process the ID signals of different hosts to make the ID unique. Figure 5-18 The module was parameterized and the master and slave interfaces were both set to 1.
[0219] The Processor System Reset module is a synchronous reset module that provides reset signals in the same clock domain. The parameter settings are as follows: Figure 5-19 .
[0220] The System ILA module is used for logic analysis to verify whether the IP function is correct and the parameter settings are as follows: Figure 5-20 .
[0221] Figure 5-21 A schematic diagram of the address mapping between Master and Slave is shown.
[0222] The ARM side needs to send matrix A and matrix B to the FPGA and receive the calculation result X from the FPGA.
[0223] ARM can send data to FPGA through the built-in API function of XILINX, namely Xil_Out32(addr, data), where addr is equal to the base address plus the offset address, data is the data to be sent, Out indicates the direction is ARM→FPGA, and 32 indicates that the data width read is 32 bits.
[0224] The function for ARM to read data from FPGA is Xil_In32(addr), where addr has the same meaning as above, In indicates the direction: FPGA→ARM, and 32 indicates that the data width to be read is 32 bits.
[0225] The base address of the embodiment of the present disclosure is 0x8000_0000, which is determined by the hardware settings. Figure 5-21 This is the address mapping diagram of Master (ie ARM) and Slave (ie FPGA).
[0226] In a possible implementation, the network interface computing unit 310 may be configured to:
[0227] When detecting that the first sending flag is a first preset value (e.g., 0), sending a line of data, setting the first sending flag to a second preset value (e.g., 1), the first sending flag indicating whether the network interface computing unit 310 has completed sending a line of data;
[0228] If it is detected that the first receiving flag is the second preset value, the count value is increased by 1, wherein the first sending flag, the first receiving flag, and the initial value of the count value are all the first preset values, and the first receiving flag indicates whether the parallel simulation calculation unit 320 has completed receiving a row of data;
[0229] In a possible implementation, the parallel simulation computing unit 320 may be used to:
[0230] When detecting that the first sending flag is a second preset value, receiving the current row data sent by the network interface computing unit 310;
[0231] If the current row of data is received, the first reception flag is set to a second preset value;
[0232] If the count value is less than the preset value, the first receiving flag is set to the first preset value.
[0233] In a possible implementation, the parallel simulation computing unit 320 is further configured to:
[0234] When the count value reaches a preset value, stopping data reception;
[0235] An adaptive operation is performed according to the received node conductance matrix and node current matrix to obtain an operation result.
[0236] Figure 5-22 Shows a schematic diagram of the ARM data transmission program flow.
[0237] For example, assume that the dimension of matrix A is n×n and the dimension of B is n×1. During the matrix sending process, matrix A and matrix B are stored in a (n+1)×n float array. One row of data, i.e., 24 data, is sent each time, and the number of rows sent is determined by the index value i.
[0238] The embodiment of the present disclosure sets a flag bit to ensure that the ARM and FPGA can complete the data transmission and reception work in an orderly manner, wherein the flag bit transfer_AB_finish_flag indicates whether the ARM has completed the transmission of a row of matrices, and the flag bit recev_AB_finish_flag indicates whether the FPGA has completed the reception of a row of matrices. The two flag bits are stored in registers in the FPGA, and the flag bit reading and writing work is completed through Xil_Out32(addr, data) and Xil_In32(addr).
[0239] For example, Figure 5-22 As shown, after the program runs, first the two flags are set to 0 and the counter i is set to 0;
[0240] For example, Figure 5-22 As shown, a row of matrix information is then sent, and the flag transfer_AB_finish_flag is set to 1 through the AXI4-Lite bus to inform the FPGA that it can start receiving the row of data;
[0241] For example, Figure 5-22 As shown, after FPGA finishes receiving, it sets the flag bit recev_AB_finish_flag to 1, telling ARM that it can send the next line of data;
[0242] For example, Figure 5-22 As shown, ARM continuously reads the value of the register storing recev_AB_finish_flag through polling. If it is 1, it increments the counter i by 1 and sets recev_AB_finish_flag to 0. When the counter reaches n, it indicates that matrix A and matrix B have been sent and the work can be ended. Otherwise, it sets transfer_AB_finish_flag to 0 and continues to send the next row of matrix information, repeating the above steps.
[0243] In a possible implementation, the parallel simulation computing unit 320 may be used to:
[0244] Send the operation result, and set the second sending flag to a second preset value, where the initial value of the second sending flag is the first preset value, and the second sending flag indicates whether the parallel simulation calculation unit 320 has completed data sending;
[0245] In a possible implementation, the network interface computing unit 310 may be configured to:
[0246] receiving the operation result when detecting that the second sending flag is a second preset value;
[0247] When receiving the operation result, setting a second reception flag bit to a second preset value, where the initial value of the second reception flag bit is the first preset value, and the second reception flag bit indicates whether the network interface computing unit 310 has completed receiving the data;
[0248] The calculation result is sent to the power supply system agent model simulation calculation module 20.
[0249] In a possible implementation, the parallel simulation computing unit 320 is further configured to:
[0250] When it is detected that the second receiving flag is changed from the initial value to the second preset value, the second receiving flag and the second sending flag are both set to the first preset value, and the sending process of the operation result is ended.
[0251] Figure 5-23 Shows the ARM data receiving program flow chart.
[0252] Assume that the dimension of matrix X is n×1. During the matrix receiving process, ARM reads the data placed by FPGA on the AXI4-Lite bus and stores the matrix X in an n×1 float array. The data receiving program flow is as follows: Figure 5-23 As shown in the figure, flags are also set here to ensure that the ARM and FPGA can complete data transmission and reception in an orderly manner. The flag transfer_X_finish_flag indicates whether the FPGA has placed the calculation result X on the AXI4-Lite bus, and the flag recev_X_finish_flag indicates whether the ARM has completed receiving the matrix X. These two flags are stored in registers in the FPGA and are read and written using Xil_Out32(addr, data) and Xil_In32(addr).
[0253] For example, Figure 5-23 As shown, after the program runs, the two flags are first set to 0.
[0254] For example, Figure 5-23 As shown, the FPGA then sends the calculation result X to the AXI4-Lite bus and sets the flag transfer_X_finish_flag to 1, indicating that the ARM can receive data;
[0255] For example, Figure 5-23 As shown in the figure, after ARM receives the data, it sets the flag bit recev_X_finish_flag to 1, indicating that the reception is complete. Otherwise, it polls the register storing recev_X_finish_flag until its value is 1, then exits the loop. After completing the data reception, it sets the two flags to 0, ending the task.
[0256] Figure 5-24 Shows a schematic diagram of ARM and FPGA data reception and transmission verification.
[0257] The matrix scale here is comparable to the matrix scale of a circuit with two stations and one interval. The disclosed embodiment can simulate the improvement in simulation speed compared to the CPU solution when using FPGA to simulate a circuit with two stations and one interval.
[0258] To verify whether the ARM and FPGA can complete data exchange, this section simulates data exchange between two stations and one interval. The dimension of the matrix A corresponding to the two stations and one interval is 24×24, the dimension of the matrix B is 24×1, and the dimension of the matrix X is 24×1.
[0259] The disclosed embodiment performs the following test: The ARM writes a row of matrix information, namely 24 float values, to registers reg0-reg23 on the FPGA and stores them in other FPGA registers. The FPGA then places the values of these registers into registers reg24-reg47 on the AXI4-Lite bus for the ARM to read. If the sent and received data are the same, the data exchange function is functioning properly.
[0260] Depend on Figure 5-24 It can be seen that the data written by ARM to the FPGA register is the same as the data read, which is in line with the expected effect. Therefore, the data transmission and reception functions between ARM and FPGA are normal.
[0261] The simulation experiment results are analyzed below.
[0262] Figure 6-1 A schematic diagram of the least squares fitting method for solving small-scale equations is shown.
[0263] Figure 6-2 A schematic diagram of the least squares fitting method for large-scale equation solution time is shown.
[0264] Figure 6-3 A schematic diagram of quadratic fitting of large-scale equation solution time is shown.
[0265] FPGA is used to measure even-dimensional time data between 20 and 32 dimensions. The results are shown in Table 5-1:
[0266] Table 5-1 FPGA solution small-scale equation solution time
[0267]
[0268] Table 5-2 shows the time data for solving a system of linear equations using the Windows Lapack library. Each data point is the average of 100 measurements. The computer has 32.0 GB of memory and an Intel Xeon E5-2678 v3 CPU with 12 cores and 24 logical processors.
[0269] Table 5-2 CPU solution small-scale equation solution time
[0270]
[0271] Use MATLAB to perform least squares fitting on the FPGA and CPU calculation times, and the fitting lines are as follows: Figure 6-1 :
[0272] The linear equation solved by FPGA is:
[0273] y=2.0357x-24.5000 formula (5-1)
[0274] The linear equation solved by the CPU is:
[0275] y=2.5747x-26.1230 Formula (5-2)
[0276] Table 5-3 shows the CPU timing for solving equations with larger dimensions. Due to hardware limitations, FPGAs are currently unable to solve equations with dimensions greater than 34.
[0277] Table 5-3 CPU solution large-scale equation solution time
[0278]
[0279] Add large-scale equation solving data and use MATLAB to perform least squares fitting on the FPGA and CPU solution time. The fitting line is as follows: Figure 6-2 .
[0280] Perform quadratic fitting on the FPGA and CPU solution time, and the fitting curve is as follows Figure 6-3 .
[0281] The disclosed embodiments utilize parallel structures and computational modules to increase computational speed, and can design parallel execution instructions and appropriate processing arrays to improve execution efficiency. In the distributed system of the disclosed embodiments, the hardware acceleration platform utilizes a Zynq system that integrates ARM and FPGA. This system-on-chip combines the software programmability of the processor with the hardware programmability of the FPGA. ARM can be used to construct a network communication interface to facilitate communication and data exchange between the FPGA and other modules, and high-speed data transmission between the ARM and FPGA can be achieved via the AXI4 bus.
[0282] This demonstrates that the use of an FPGA-based parallel simulation method to achieve efficient parallel simulation of the subway power supply system, and the utilization of an ARM-based distributed system network communication interface to enable efficient communication and data exchange between the FPGA and other modules, has important practical engineering applications and technical research value. Furthermore, this method has the potential for widespread application in other circuit-related simulation and analysis fields.
[0283] While various embodiments of the present disclosure have been described above, the above descriptions are illustrative, non-exhaustive, and not intended to be limiting of the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A rail transit power supply system simulation device, characterized in that: The device comprises: Train model simulation calculation module, used to simulate the train running status and send the train location information; a power supply system proxy model simulation calculation module, connected to the train model simulation calculation module, for determining and sending a node conductance matrix and a node current matrix based on the location information, pre-configured circuit netlist information, and operating line information, wherein the node conductance matrix includes the conductances of multiple nodes, and the node current matrix includes the currents of multiple nodes; a power supply system simulation calculation module, connected to the power supply system proxy model simulation calculation module, configured to perform adaptive calculations based on the node conductance matrix and the node current matrix to obtain calculation results, and send the calculation results to the power supply system proxy model simulation calculation module, wherein the adaptive calculations include parallel calculations, and the calculation results include a node voltage matrix, which includes voltages of multiple nodes; The power supply system proxy model simulation calculation module is also used to send the calculation result to the train model simulation calculation module, so that the train model simulation calculation module updates the train operation status.
2. The device according to claim 1, characterized in that The power supply system simulation calculation module includes: an ARM-based network interface calculation unit and an FPGA-based parallel simulation calculation unit. The network interface calculation unit is provided between the power supply system proxy model simulation calculation module and the parallel simulation calculation unit, and the network interface calculation unit is used to realize information exchange between the power supply system proxy model simulation calculation module and the parallel simulation calculation unit; The parallel simulation calculation unit is used to perform adaptive calculations based on the node conductance matrix and the node current matrix to obtain calculation results, and the parallel calculations include parallel LU decomposition calculations.
3. The device according to claim 2, characterized in that The power supply system agent model simulation calculation module is connected to the network interface calculation unit through an Ethernet interface, wherein, The node conductance matrix, node current matrix, and node voltage matrix stored in the power supply system proxy model simulation calculation module and the parallel simulation calculation unit are all floating-point data, wherein the power supply system proxy model simulation calculation module and the parallel simulation calculation unit are both used to: convert the data to be sent into character data and send it, and convert the received data into floating-point data and store it; The network interface calculation unit is used to send at least one row of data of the node conductance matrix or the node current matrix to the parallel simulation calculation unit each time.
4. The device according to claim 2 or 3, characterized in that The network interface computing unit is used for: When detecting that the first sending flag is a first preset value, sending a line of data, setting the first sending flag to a second preset value, wherein the first sending flag indicates whether the network interface computing unit has completed sending a line of data; If it is detected that the first receiving flag is a second preset value, the count value is increased by 1, wherein the first sending flag, the first receiving flag, and the initial value of the count value are all first preset values, and the first receiving flag indicates whether the parallel simulation calculation unit has completed receiving a row of data; The parallel simulation calculation unit is used for: receiving the current row data sent by the network interface computing unit when detecting that the first sending flag is a second preset value; If the current row of data is received, the first reception flag is set to a second preset value; If the count value is less than the preset value, the first receiving flag is set to the first preset value.
5. The device according to claim 4, characterized in that The parallel simulation calculation unit is also used for: When the count value reaches a preset value, stopping data reception; An adaptive operation is performed according to the received node conductance matrix and node current matrix to obtain an operation result.
6. The device according to claim 2 or 3, characterized in that The parallel simulation calculation unit is used for: Sending the operation result, and setting a second sending flag bit to a second preset value, where the initial value of the second sending flag bit is the first preset value, and the second sending flag bit indicates whether the parallel simulation calculation unit has completed data sending; The network interface computing unit is used for: receiving the operation result when detecting that the second sending flag is a second preset value; When receiving the operation result, setting a second reception flag bit to a second preset value, where the initial value of the second reception flag bit is the first preset value, and the second reception flag bit indicates whether the network interface computing unit has completed receiving the data; The calculation result is sent to the power supply system agent model simulation calculation module.
7. The device according to claim 6, characterized in that The parallel simulation calculation unit is also used for: When it is detected that the second receiving flag is changed from the initial value to the second preset value, the second receiving flag and the second sending flag are both set to the first preset value, and the sending process of the operation result is ended.
8. The device according to claim 1, characterized in that The power supply system agent model simulation calculation module includes: Circuit simulation class, including simulation circuit structures composed of multiple passive component simulation models and multiple active component simulation models; The rail transit line class includes a line information model and an update model. The line information model stores rail transit line information. The update model is used to update the simulation circuit structure and update the node conductance matrix and node current matrix.
9. The device according to claim 8, characterized in that The power supply system simulation calculation module includes at least one of a DC simulation type, an AC simulation type, and an AC / DC simulation type, wherein: The DC simulation class includes a simulation initialization class and an equivalent capacitance class, which are used to assist simulation calculations, wherein the equivalent capacitance class is an equivalent model of a capacitor under a set simulation step size; The AC simulation class is used to calculate the power flow equation on the AC side and obtain the voltage of the traction substation; The AC / DC simulation class is used for iterative joint simulation of AC and DC.
10. The device according to claim 1, characterized in that The device further comprises: Signal system model simulation calculation module, used to simulate and model the signal transmission system of train operation and signal transmission in rail transit; Infrastructure simulation calculation module, used to simulate and model the infrastructure in the rail transit system; The dispatching system model simulation calculation module is used for simulation modeling of the dispatching system of the rail transit system; Passenger flow simulation calculation module is used to simulate and model the passenger flow of rail transit systems. The train model simulation calculation module is further configured to send the calculation result to the dispatching system model simulation calculation module so that the dispatching system model simulation calculation module updates the dispatching information of the train.