Simulation calculation method and device for traction power supply system
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 国能新朔铁路有限责任公司
- Filing Date
- 2026-04-13
- Publication Date
- 2026-08-07
AI Technical Summary
[0004]本申请实施例的目的是提供一种牵引供电系统的仿真计算方法及装置,以解决相关技术中仿真计算时不收敛或收敛速度慢的问题
本申请实施例通过构建牵引供电系统对应的牵引所状态迭代图,并以深度优先算法为基础,引入强化学习算法构建自适应控制策略,在仿真迭代过程中能够自适应选择状态切换策略,从而提高了迭代收敛的概率,缩短了收敛时间。
Smart Images

Figure CN122528592A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of traction power supply technology, and in particular relates to a simulation calculation method and device for a traction power supply system. Background Technology
[0002] As urban traffic pressure increases, subways have become an important mode of transportation to alleviate urban traffic congestion. Subway stations are close together and trains operate at high density. During operation, the frequent starting and braking of trains generates a large amount of regenerative braking energy. If this energy is not effectively managed, it may lead to safety accidents.
[0003] Currently, energy storage devices are commonly used to absorb excess regenerative braking energy. However, the installation of energy storage devices significantly increases the state dimension of the traction power supply system, leading to frequent state transitions. This results in non-convergence or slow convergence during simulation calculations, especially in the iterative state solution process, severely impacting simulation efficiency and engineering application value. To ensure the reliability and practicality of simulation calculations, the industry urgently needs a simulation scheme that can improve iterative convergence. Summary of the Invention
[0004] The purpose of this application is to provide a simulation calculation method and apparatus for a traction power supply system to solve the problems of non-convergence or slow convergence speed in related technologies.
[0005] To achieve the above objectives, the embodiments of this application adopt the following technical solutions: In a first aspect, embodiments of this application provide a simulation calculation method for a traction power supply system, comprising: constructing a state iteration diagram of a traction substation corresponding to the traction power supply system, wherein the state iteration diagram includes different states of each traction substation in the traction power supply system and iterative relationships between the different states of each traction substation; traversing the different states of each traction substation in the state iteration diagram based on a depth-first search algorithm; and determining the optimal iteration path in the state iteration diagram of the traction substation based on a reinforcement learning algorithm during the traversal process.
[0006] Secondly, embodiments of this application provide a simulation computing device for a traction power supply system, comprising: a construction module for constructing a state iteration diagram of a traction substation corresponding to the traction power supply system, wherein the state iteration diagram includes different states of each traction substation in the traction power supply system and iterative relationships between the different states of each traction substation; a traversal module for traversing the different states of each traction substation in the state iteration diagram based on a depth-first search algorithm; and a determination module for determining the optimal iteration path in the state iteration diagram of the traction substation based on a reinforcement learning algorithm during the traversal process.
[0007] The above-described technical solutions adopted in the embodiments of this application can achieve the following beneficial effects: This application embodiment constructs a state iteration diagram of the traction power supply system corresponding to the traction substation, and introduces a reinforcement learning algorithm to construct an adaptive control strategy based on the depth-first search algorithm. During the simulation iteration process, it can adaptively select the state switching strategy, thereby improving the probability of iteration convergence and shortening the convergence time. Attached Figure Description
[0008] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart illustrating a simulation calculation method for a traction power supply system provided in one embodiment of this application; Figure 2 A schematic diagram of the external characteristics of an energy storage device provided for one embodiment of this application; Figure 3 A schematic diagram illustrating the switching of traction station states according to one embodiment of this application; Figure 4 A schematic diagram of a maze provided for one embodiment of this application; Figure 5 A schematic diagram of traction station state iteration provided for one embodiment of this application; Figure 6 A schematic diagram of the circuit structure of a traction power supply system provided in one embodiment of this application; Figure 7 A schematic diagram of the circuit structure of a traction power supply system provided in another embodiment of this application; Figure 8 A schematic diagram of the structure of a simulation computing device for a traction power supply system provided in one embodiment of this application; Figure 9 This is a schematic diagram of the structure of an electronic device provided in one embodiment of this application. Detailed Implementation
[0009] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0010] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein. Furthermore, "and / or" in this application indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship. It should be noted that all data involved in this application was obtained with the user's authorization.
[0011] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.
[0012] Figure 1 This is a flowchart illustrating a simulation calculation method for a traction power supply system, provided as an embodiment of this application. Figure 1 As shown, the simulation calculation method for the traction power supply system in this application embodiment may specifically include the following steps: S101, construct the state iteration diagram of the traction substation corresponding to the traction power supply system. The state iteration diagram of the traction substation includes the different states of each traction substation in the traction power supply system and the iterative relationship between the different states of each traction substation.
[0013] In this embodiment of the application, the execution entity of the simulation calculation method for the traction power supply system is a simulation calculation device for the traction power supply system, which can be located in an electronic device. This electronic device can be a terminal device or a server. The terminal device can be a mobile phone, tablet computer, desktop computer, laptop, vehicle-mounted device, etc.; the server can be an independent server or a server cluster composed of multiple servers. In the field of traction power supply technology, the simulation calculation device for the traction power supply system (i.e., the traction network) can be located in the intelligent agent of the traction power supply control system.
[0014] In subway traction power supply systems, there are many different types of energy storage devices, each with different properties. However, in the simulation of subway traction power supply systems, it is not necessary to study each type individually; a general black-box model can be used. For example... Figure 2 The image shows a standard black-box model for energy storage devices.
[0015] Based on the collaborative working mode of the energy storage device and the rectifier unit in the traction substation, the different states of each traction substation and the iterative relationship between the different states of each traction substation can be determined, and different states correspond to different circuit models.
[0016] Specifically, based on the collaborative working method of the energy storage device and the rectifier unit in the traction substation, the traction substation can be divided into: Figure 2The six operating states shown (referred to as the six states) each correspond to a different circuit model. Figure 2 middle: The segment curve represents the maximum power discharge state (EDP). The segment curve represents the constant voltage discharge state (EDU). The segment curve represents the rectified state (REC). The segment curve represents the hold state (TOFF); The segment curve represents the constant voltage charging state (ECU). The curve represents the maximum power charging state (ECP). U E The voltage of the energy storage device; I E The current of the energy storage device; P dmax This represents the maximum discharge power of the energy storage device. I dmax This is the maximum discharge current of the energy storage device; I dlim This is the critical current for constant power discharge of the energy storage device; P cmax This is the maximum charging power of the energy storage device. I cmax This is the maximum charging current for the energy storage device, and also the critical current for constant power charging of the energy storage device. U min This is the minimum operating voltage for the energy storage device; U d This refers to the discharge threshold voltage of the energy storage device. U nl This refers to the no-load voltage of the rectifier unit in the traction substation. U c The threshold voltage for charging the energy storage device; U max This is the maximum operating voltage of the energy storage device; The switching patterns between the various states can be summarized as follows: Figure 3The diagram shown illustrates the state switching of the traction substation. The iterative calculation of the traction substation state in the simulation calculation of the subway DC traction power supply system uses the switching rule for iteration.
[0017] In the simulation calculation of the DC traction power supply system of the subway, the state iteration solution process of the traction substation after the addition of energy storage device can be compared with the "maze problem". Figure 4 A diagram illustrating maze navigation is provided. For maze navigation, assume the agent starts at the maze's starting point and continuously tries different paths: if the path is passable (stepping onto a white square), it continues exploring; if it encounters an obstacle (stepping onto a black square), it returns to the starting point and chooses a new path. Through continuous exploration, the agent eventually finds one or more paths that can navigate the maze, and these paths essentially correspond to a series of combinations of white squares.
[0018] Applying this idea to the iterative calculation of traction substation states, instead of improving upon existing state switching rules (which are based on fixed thresholds or empirical rules), it directly breaks the existing state switching rules and adopts a completely new state iteration strategy. In the simulation of a subway DC traction power supply system, each traction substation state corresponds to a different circuit model. Solving the simulation is essentially solving for different circuit structures. Different traction substations on a line have different states, and each state combination corresponds to a solution. Assuming a line contains 10 traction substations, and each traction substation has 5 possible operating states, then the total number of state combinations for the entire line is 510. The final solution that allows the iterative calculation to converge must be contained within these finite number of state combinations.
[0019] Traditional state iteration methods rely on preset state switching rules to find convergent combinations, but because their strategies cannot exhaust all possible combinations, some convergent state combinations may be missed; at the same time, the limitations of traditional iteration methods themselves may also cause the iteration process to fall into a non-convergent loop.
[0020] This application's embodiments employ a depth-first search algorithm to traverse the traction station's state. Prior to this, it is necessary to construct, as follows: Figure 5 The diagram shown is the traction substation state iteration diagram corresponding to the traction power supply system. (See diagram for example.) Figure 5 As shown, the traction substation state iteration diagram includes the different states of each traction substation and the iterative relationships between these states, for example, A1-B2-C3. Assume there are three traction substations in the traction power supply system, and each substation has three states. Figure 5 In this diagram, A, B, and C represent the three states of a traction substation, and 1, 2, and 3 represent different traction substations. Therefore, A1 represents traction substation 1 in state A, and B2 represents traction substation 2 in state B. Figure 5 It is a one-way graph. Figure 5The elements within the game need to move in the direction indicated by the arrows.
[0021] Furthermore, the state of the traction substation can be represented by the traction network admittance matrix, which is updated during each iteration. Correspondingly, before step S101 "constructing the traction substation state iteration diagram corresponding to the traction power supply system", the simulation calculation method of the traction power supply system in this embodiment may also include the following steps for constructing the traction network admittance matrix: constructing the traction network admittance matrix corresponding to the traction power supply system using the nodal voltage method, with different states of each traction substation corresponding to different traction network admittance matrices, and updating the traction network admittance matrix during each iteration.
[0022] Specifically, Figure 6 The diagram shows a circuit structure of a DC traction power supply system (i.e., DC traction network) consisting of two traction substations and two trains. The traction substations use the Thevenin equivalent model (a series connection of voltage source and impedance), while the trains use a power source model. The DC traction network is divided into four sections.
[0023] The nodal voltage method was used for simulation calculations. Y-array, U-array and I-array were constructed respectively. Finally, the nodal voltage equation YU= I was written as shown in the following equation (1): (1) In formula (1): Y The admittance matrix of the network is constructed using a block matrix notation, since the segmented structure of each section is basically the same. Y , Y ii Indicates cross-section i Self-guided absorbance, Y ij Indicates cross-section i and j Mutual admittance, i , j = 1, 2, 3, 4. U =[ U 1 U 2 U 3 U 4] T For voltage vectors, I = [ I 1 I 2 I 3 I 4] T For current phasors, all are constructed using block matrix notation. Figure 6 The circuit structure diagram shown is mainly updated during iterative calculations based on the different states of the traction substation and the circuit model. Y In the formation Y 11 and Y44 The specific forms of the two matrices are shown in equations (2) and (3) below: (2) (3) S102, based on the depth-first algorithm, traverses the different states of each traction station in the traction station state iteration diagram.
[0024] In the embodiments of this application, for Figure 5 In the example shown, assuming that the path formed by the three states A1-B2-C3 can make the computation converge, the depth-first algorithm starts searching from A1-A2-A3, then to A1-A2-B3, then to A1-A2-C3, and finally can find the path A1-B2-C3 to make the computation converge.
[0025] S103, during the traversal process, the optimal iteration path is determined in the traction state iteration diagram based on the reinforcement learning algorithm.
[0026] In this embodiment, regarding the aforementioned depth-first search algorithm, although it can solve the problem of non-convergence of traction substation state iteration in the simulation of metro DC traction networks, the search cost is extremely high when the line scale is large, requiring a lot of time to find the state combination that makes the iteration converge. To improve the solution efficiency, this embodiment introduces a reinforcement learning algorithm, replacing exhaustive search with an agent interacting with the environment, thereby achieving adaptive learning and rapid decision-making on state switching strategies.
[0027] During the traversal process, the state-action value function can be used as a basis. The optimal iteration path is determined in the state iteration diagram of the traction station.
[0028] Furthermore, the above steps are based on the state-action value function. Determining the optimal iteration path in the traction station state iteration diagram may include the following steps: based on the current convergence state... s t Select the corresponding action a t And simulate the execution of actions. a t The action is the state of the traction unit; obtain the simulated execution action. a t The reward after r t and the next convergence state s t+1 ;according to r t and the next convergence state s t+1Update the state-action value function Corresponding state-action value When the iteration termination condition is met, the iteration process ends, and the final list of state-action values, i.e., the Q-table, is generated. This serves as a list of target state action values; based on this list, the optimal iteration path is determined.
[0029] Specifically, as mentioned earlier, the problem of navigating a maze can be abstracted into a path planning problem, that is, finding the optimal path to navigate the maze. The Q-learning algorithm is often used for path planning problems. In the state iteration problem of the traction device, the Q-learning algorithm can be used for training.
[0030] Q-learning is a value-function-based reinforcement learning method applicable to decision optimization problems with discrete state and action spaces. The Q-learning algorithm constructs a state-action value function... The agent continuously iterates and updates during its interaction with the environment, enabling it to learn strategies for selecting the optimal action under different convergence states.
[0031] The basic principle of the Q-learning algorithm is to abstract the iterative calculation process of the traction power supply system into a reinforcement learning environment. The environment's operation result only focuses on whether the iteration converges, defined as a binary output: converged or not converged. In each interaction step, the agent determines the convergence state based on the current convergence state. s t Select Action a t Instant rewards for environmental feedback r t and the next convergence state s t+1 The Q value is then corrected according to the following formula (4): (4) in, s t This represents the current convergence state. a t The action selected in the current convergence state, the one to the left of the arrow. For the updated state-action value, the one to the right of the arrow The value of the action before the update. α The learning rate is used to control the update magnitude. r t To simulate the execution of actions a t The reward afterwards γ A discount factor is used to balance immediate rewards with long-term returns.s t+1 For the next convergence state, This is the optimal action in the next convergence state. The optimal action in the next convergence state The estimated value of state-action.
[0032] Through multiple iterations, the Q-table gradually converges, storing the value of different actions chosen in the "non-converged" state. During online execution, the agent selects the optimal action based on the Q-table, achieving adaptive control of the iterative process and ensuring that it enters the "converged" state as quickly as possible.
[0033] It should be noted that the agent training can be completed offline, and the time cost of training is consumed before deployment. Once online, the agent already possesses decision-making capabilities, directly addressing traction station state iteration problems, significantly reducing online computational overhead and avoiding simulation interruptions due to iteration non-convergence. For trained lines, the agent can adaptively adjust state switching strategies based on real-time operating conditions, automatically searching for and selecting traction station state combinations that enable iteration convergence, thereby improving the convergence probability and shortening the convergence time. By replacing empirical rules and manual parameter tuning with learned strategies, manual intervention by engineers under multiple operating conditions is reduced, improving the automation and repeatability of the simulation process. Combining offline training and rapid online decision-making, this solution ensures the reliability of simulation results while also meeting the real-time requirements of engineering applications, facilitating its widespread application in design, evaluation, and operational support.
[0034] Among them, the above step "based on the current convergence state" s t Select the corresponding action a t Specifically, this may include the following steps: based on the current convergence state s t And exploration rate greed ( -greedy) Strategy selection action a t In a greedy exploration strategy, the exploration rate... The cost is gradually reduced during the iteration process, thus realizing the transformation from exploration to utilization.
[0035] Definition of convergent state space: The set of convergent states is simplified to s ={converged, not converged}, the current convergence state is determined by the convergence criterion of the iterative calculation, such as whether the power flow calculation error is less than a set threshold.
[0036] Action space definition: The action set consists of the six working states of the traction substation: rectification (REC), constant voltage charging (ECU), maximum power charging (ECP), constant voltage discharging (EDU), maximum power discharging (EDP), and shutdown (TOFF). Each action corresponds to a different power supply strategy. The environment updates the power flow distribution according to the action. The action space is a finite discrete set, which is suitable for Q-learning algorithms.
[0037] The above steps "obtain simulation execution actions" a t The reward after r t Specifically, this may include the following steps: If the simulation converges iteratively after executing the action, then obtain the preset first positive reward + R conv If the simulation fails to converge after the iterations have been executed and the preset maximum number of iterations has not been exceeded, then a preset second positive reward or first negative reward will be obtained. R step The second positive reward is less than the first positive reward; if the simulation fails to converge after the action is executed and the maximum number of iterations is exceeded, a preset second negative reward is obtained. R fail Second negative reward - R fail Less than the first negative reward - R step .
[0038] Specifically, the reward function focuses on convergence. If the convergence state transitions to "convergence", a high positive reward is given; if the convergence state remains "non-converged" and exceeds the preset maximum number of iterations, a low reward or penalty is given; if the number of iterations exceeds the maximum number of iterations and convergence is still not achieved, a strong negative reward is given. The reward function can be expressed as shown in equation (5) below: (5) The formulas mentioned here need to be initialized before the iteration process: initialize the Q-table. , Set all values to zero and set the learning rate. α Discount Factor γ Exploration rate , Reset the environment and set the initial convergence state to "non-converged".
[0039] The aforementioned "iteration termination condition" may include, but is not limited to: if the system enters a "convergence" state or reaches the upper limit of the number of iterations, it indicates that the iteration termination condition is met and the iteration process ends.
[0040] In addition, during the simulation calculation, the traditional iterative method is used first. If the calculation converges, the result is output. If the calculation does not converge, the convergence enhancement strategy of this application embodiment (i.e., the simulation calculation method based on the depth-first algorithm and the reinforcement learning algorithm) is used to perform the calculation.
[0041] To demonstrate the effectiveness of the simulation calculation method for the traction power supply system in this application for solving the problem of non-convergence in state iteration calculation, a practical example is given below.
[0042]
Calculation Example Parameters
[0043] The parameters of the traction sub-model in the model are shown in Table 1 below: Table 1 Traction Station Model Parameters
[0044] Traction network parameters: contact network unit resistance 0.0136Ω / km, rail unit resistance 0.02Ω / km, ground leakage resistance 15Ω / km.
[0045] The energy storage device used in the example has an equivalent resistance of 0.01Ω.
[0046] The parameters of the energy storage device model are shown in Table 2 below: Table 2 Energy Storage Device Model Parameters
[0047] The power of each train used in the calculation example is as follows: -4320 kW, -4130 kW, 2532 kW, 1821 kW, 1130 kW, 1310 kW, and 1213 kW.
[0048] [Calculation Results] When using the traditional iterative method, once the iteration process is determined to have entered a loop state in the fourth iteration, it can be concluded that the combination of traction state states leads to non-convergence, and the simulation outputs a "non-convergence" result. In actual simulation software use, the software will output a "calculation failed" result at this time. To make the calculation converge, it is necessary to change the input simulation parameters or manually adjust them.
[0049] When adopting the convergence enhancement strategy of this application embodiment, when the iteration process is determined to be trapped in a loop state for the fourth time, the convergence enhancement mechanism is immediately triggered, and the traction combination that can make the state iteration calculation converge is automatically searched. After finding a new traction state combination, the simulation outputs the "convergence" result, which solves the defect of the traditional state iteration algorithm that does not converge in the iterative calculation and cannot obtain a usable solution.
[0050] Table 3. Results of the example
[0051]
Summarize
[0052] Overall, the improved convergence strategy not only maintains the computational speed advantage of the traditional method, but also achieves significant improvements in convergence, adaptability, and robustness, demonstrating greater value for engineering applications. A comparison of the two methods is shown in Table 4 below: Table 4 Comparison of the two calculation methods
[0053] In summary, the simulation calculation method for the traction power supply system in this application constructs an iterative state diagram of the traction substation corresponding to the traction power supply system. Based on the depth-first search algorithm, it introduces a reinforcement learning algorithm to construct an adaptive control strategy. During the simulation iteration process, it can adaptively select the state switching strategy, thereby improving the probability of iteration convergence and shortening the convergence time. The iterative solution problem of the traction substation state is abstracted into a combination search problem of several traction substation states, thus overcoming the limitations of traditional iterative strategies and discovering convergent state combinations that are difficult to cover or easily overlooked by traditional methods. Based on this idea, a depth-first search is introduced for the system to traverse state combinations, and a reinforcement learning algorithm is combined to train the agent, enabling it to adaptively select or adjust the traction substation state configuration during traversal and decision-making. This method can exhaustively explore potential convergent solutions and gradually optimize the search strategy through learning, significantly improving the convergence of simulation iteration, reducing the need for manual parameter tuning, and enhancing robustness and applicability under complex operating conditions and large-scale lines.
[0054] Figure 8 This is a schematic diagram of the structure of a simulation calculation device for a traction power supply system provided in one embodiment of this application. Figure 8 As shown, the simulation calculation device 800 for the traction power supply system in this embodiment of the application may specifically include: a construction module 801, a traversal module 802, and a determination module 803. Wherein: Module 801 is used to construct the state iteration diagram of the traction substation corresponding to the traction power supply system. The state iteration diagram of the traction substation includes the different states of each traction substation in the traction power supply system and the iterative relationship between the different states of each traction substation. Traversal module 802 is used to traverse the different states of each traction station in the traction station state iteration diagram based on the depth-first search algorithm. The determination module 803 is used to determine the optimal iteration path in the traction state iteration graph based on the reinforcement learning algorithm during the traversal process.
[0055] In this embodiment of the application, the specific process by which each module in the simulation calculation device of the traction power supply system implements its function can be found in the relevant description in the above embodiment of the simulation calculation method of the traction power supply system, and will not be repeated here.
[0056] In summary, the simulation computing device for the traction power supply system in this application constructs an iterative state diagram of the traction substation corresponding to the traction power supply system. Based on the depth-first search algorithm, it introduces a reinforcement learning algorithm to construct an adaptive control strategy. During the simulation iteration process, it can adaptively select the state switching strategy, thereby improving the probability of iteration convergence and shortening the convergence time. The iterative solution problem of the traction substation state is abstracted into a combination search problem of several traction substation states, thus overcoming the limitations of traditional iterative strategies and discovering convergent state combinations that are difficult to cover or easily overlooked by traditional methods. Based on this idea, depth-first search is introduced for the system to traverse state combinations, and a reinforcement learning algorithm is combined to train the agent, enabling it to adaptively select or adjust the traction substation state configuration during traversal and decision-making. This not only exhaustively explores potential convergent solutions but also gradually optimizes the search strategy through learning, significantly improving the convergence of simulation iteration, reducing the need for manual parameter tuning, and enhancing robustness and applicability under complex operating conditions and large-scale lines.
[0057] This application also provides an electronic device. For example... Figure 9 As shown, the electronic device 900 can vary considerably due to differences in configuration or performance. It may include one or more processors 901 and memory 902, with memory 902 storing one or more programs or instructions. Memory 902 may be temporary or persistent storage. The application programs stored in memory 902 may include one or more modules (not shown), each module including a series of computer-executable instructions for the electronic device 900. Furthermore, processor 901 may be configured to communicate with memory 902 and execute the series of computer-executable instructions in memory 902 on the electronic device 900. The electronic device 900 may also include one or more power supplies 903, one or more wired or wireless network interfaces 904, one or more input / output interfaces 905, and one or more keyboards 906.
[0058] Specifically, in the embodiments of this application, the electronic device includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor. When the program or instructions are executed by the processor, they implement the steps of any of the above-described simulation calculation method embodiments for traction power supply systems.
[0059] The electronic device in this application constructs an iterative state diagram of the traction substation corresponding to the traction power supply system. Based on the depth-first search algorithm, it introduces a reinforcement learning algorithm to build an adaptive control strategy. During the simulation iteration process, it can adaptively select the state switching strategy, thereby improving the probability of iterative convergence and shortening the convergence time. The iterative solution problem of the traction substation state is abstracted into a combination search problem of several traction substation states, thus overcoming the limitations of traditional iterative strategies and discovering convergent state combinations that are difficult to cover or easily overlooked by traditional methods. Based on this idea, a depth-first search is introduced for the system to traverse state combinations, and a reinforcement learning algorithm is combined to train the agent, enabling it to adaptively select or adjust the traction substation state configuration during traversal and decision-making. This approach can exhaustively explore potential convergent solutions and gradually optimize the search strategy through learning, significantly improving the convergence of simulation iterations, reducing the need for manual parameter tuning, and enhancing robustness and applicability under complex operating conditions and large-scale lines.
[0060] This application also proposes a readable storage medium storing one or more computer programs or instructions, which, when executed by a processor in an electronic device, enable the processor in the electronic device to perform the steps of any of the above-described embodiments of the simulation calculation method for the traction power supply system.
[0061] The readable storage medium of this application constructs an iterative state diagram of the traction substation corresponding to the traction power supply system. Based on the depth-first search algorithm, it introduces a reinforcement learning algorithm to build an adaptive control strategy. During the simulation iteration process, it can adaptively select the state switching strategy, thereby improving the probability of iterative convergence and shortening the convergence time. The iterative solution problem of the traction substation state is abstracted into a combination search problem of several traction substation states, thus overcoming the limitations of traditional iterative strategies and discovering convergent state combinations that are difficult to cover or easily overlooked by traditional methods. Based on this idea, a depth-first search is introduced for the system to traverse state combinations, and a reinforcement learning algorithm is combined to train the agent, enabling it to adaptively select or adjust the traction substation state configuration during traversal and decision-making. This not only exhaustively explores potential convergent solutions but also gradually optimizes the search strategy through learning, significantly improving the convergence of simulation iterations, reducing the need for manual parameter tuning, and enhancing robustness and applicability under complex operating conditions and large-scale lines.
[0062] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0063] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.
[0064] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0065] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0066] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0067] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0068] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0069] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0070] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0071] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0072] This application can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0073] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0074] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A simulation calculation method for a traction power supply system, characterized in that, include: Construct a state iteration diagram of the traction substation corresponding to the traction power supply system. The state iteration diagram of the traction substation includes the different states of each traction substation in the traction power supply system and the iterative relationship between the different states of each traction substation. Based on the depth-first search algorithm, the different states of each traction station in the traction station state iteration diagram are traversed. During the traversal, the optimal iteration path is determined based on the reinforcement learning algorithm in the state iteration diagram of the traction station.
2. The method according to claim 1, characterized in that, Before constructing the traction substation state iteration diagram corresponding to the traction power supply system, the following steps are also included: Based on the collaborative working mode of the energy storage device and the rectifier unit in the traction substation, the different states of each traction substation and the switching relationship between the different states of each traction substation are determined, and different states correspond to different circuit models.
3. The method according to claim 2, characterized in that, Before constructing the traction substation state iteration diagram corresponding to the traction power supply system, the following steps are also included: The traction network admittance matrix corresponding to the traction power supply system is constructed using the nodal voltage method. Different states of each traction station correspond to different traction network admittance matrices, and the traction network admittance matrix is updated in each iteration.
4. The method according to claim 1, characterized in that, During the traversal process, the optimal iteration path is determined based on a reinforcement learning algorithm in the traction station state iteration graph, including: During the traversal, the optimal iteration path is determined based on the state-action value function in the state iteration diagram of the traction station.
5. The method according to claim 4, characterized in that, During the traversal process, the optimal iteration path is determined based on the state-action value function in the traction station state iteration graph, including: Select the corresponding action based on the current convergence state, and simulate the execution of the action, where the action is the state of the traction station; Obtain the reward and the next convergence state after the simulation executes the action; Update the state action value corresponding to the state action value function based on the reward and the next convergence state; When the iteration termination condition is met, the iteration process ends, and the final list of state-action values is used as the target list of state-action values. The optimal iteration path is determined based on the target state action value list.
6. The method according to claim 5, characterized in that, The step of selecting the corresponding action based on the current convergence state includes: The action is selected based on the current convergence state and the greedy exploration rate strategy.
7. The method according to claim 6, characterized in that, In the greedy exploration strategy, the exploration rate gradually decreases during the iteration process.
8. The method according to claim 5, characterized in that, The process of obtaining the reward after the simulation performs the action includes: If the simulation converges iteratively after performing the action, a preset first positive reward is obtained; If the simulation fails to converge after the action is performed and the preset number of iterations is not exceeded, then a preset second positive reward or a first negative reward is obtained, wherein the second positive reward is less than the first positive reward. If the simulation fails to converge after the action is performed and the number of iterations exceeds the upper limit, a preset second negative reward is obtained, which is less than the first negative reward.
9. The method according to claim 5, characterized in that, The step of updating the state action value corresponding to the state action value function based on the reward and the next convergence state includes: The state action value is updated using the following formula: ; Among them, the s t The current convergence state is the state in which the... a t The action selected for the current convergence state, the one to the left of the arrow The value of the updated state action, as stated by the arrow to the right. The value of the state action before the update, the α The learning rate, the r t To simulate the execution of the action a t The subsequent reward, the γ Discount factor, the s t+1 For the next convergence state, the The optimal action in the next convergence state. The optimal action in the next convergence state The estimated value of state-action.
10. A simulation calculation device for a traction power supply system, characterized in that, include: A construction module is used to construct a state iteration diagram of the traction substation corresponding to the traction power supply system. The state iteration diagram of the traction substation includes the different states of each traction substation in the traction power supply system and the iterative relationship between the different states of each traction substation. The traversal module is used to traverse the different states of each traction station in the traction station state iteration diagram based on the depth-first search algorithm. The determination module is used to determine the optimal iteration path in the state iteration graph of the traction station based on a reinforcement learning algorithm during the traversal process.