Energy-saving control method and equipment for freight train

By applying simulated operating environment, reward function and reinforcement learning model in freight trains, combined with the depth of attention mechanism, the residual neural network can be separated, which solves the problem of high energy consumption of freight trains and achieves efficient and energy-saving operation.

CN120029066APending Publication Date: 2025-05-23CRRC YANGTZE GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510172561.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing freight train control methods have large energy consumption, which is difficult to effectively reduce energy consumption and cannot adapt to the trend of energy conservation development.

Method used

By establishing a simulated operation environment, determining the reward function, and using a preset deep Q network to build an energy-saving operation control model based on reinforcement learning, using a deep separable residual neural network based on attention mechanism, and formulating the control strategy for freight trains.

Benefits of technology

It realizes the generation of a high-performance and efficient energy-saving operation control model, can make decisions that optimize energy consumption, reduces the energy consumption of freight trains, and conforms to the trend of train energy conservation development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029066A_ABST
    Figure CN120029066A_ABST
Patent Text Reader

Abstract

The invention discloses an energy-saving control method and equipment for a freight train, and belongs to the technical field of freight trains, and the method comprises the steps: building a simulation operation environment according to the technical parameters, line parameters and operation modes of the freight train; determining a reward function based on the operation energy consumption of the freight train; constructing an energy-saving operation control model based on reinforcement learning by using the simulated operation environment, the reward function and a preset deep Q network; wherein the Q network in the preset depth Q network adopts a depth separable residual neural network based on an attention mechanism; and determining a control strategy of the freight train according to the energy-saving operation control model. Through the technical scheme provided by the invention, the high-performance and high-efficiency energy-saving operation control model can be generated, the energy-saving operation control model can make a decision making the optimal energy consumption, the energy consumption of the freight train is reduced, and the method conforms to the trend of energy-saving development of the train.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of freight trains, and particularly relates to an energy-saving control method and device for freight trains. Background Art

[0002] A freight train is a train for transporting goods and is one of the common equipment in rail transit transportation. The Outline for the Construction of a Transportation Powerhouse proposes that by 2035, China will basically form a transportation powerhouse, laying a solid foundation for the great development of rail transit transportation. To meet the needs of the development of modern rail transit logistics transportation, rail transit transportation equipment needs to continuously innovate in technology to significantly improve the modernization and intelligence levels of railway freight transportation, promote the transformation and upgrading of China's rail transit transportation to intelligent freight transportation, and achieve an overall technical level reaching the world's advanced level.

[0003] Currently, rail transit transportation is one of the industries with relatively high energy consumption. As an important rail transit equipment, a freight train is usually controlled by a driver based on experience, resulting in relatively high energy consumption. Based on this, how to reduce the energy consumption of freight trains to conform to the trend of energy-saving development is an urgent problem to be solved. Summary of the Invention

[0004] The embodiments of this application provide an energy-saving control method and device for freight trains, which can, to at least a certain extent, reduce the energy consumption of freight trains to conform to the trend of energy-saving development.

[0005] Other features and advantages of this application will become apparent through the following detailed description, or be learned in part through the practice of this application.

[0006] According to the first aspect of the embodiments of this application, an energy-saving control method for a freight train is provided, including:

[0007] Establish a simulated operation environment according to the technical parameters, line parameters, and operation mode of the freight train;

[0008] Determine a reward function based on the operation energy consumption of the freight train;

[0009] Construct an energy-saving operation control model based on reinforcement learning by using the simulated operation environment, the reward function, and a preset deep Q-network; wherein, the Q-network in the preset deep Q-network adopts a depthwise separable residual neural network based on an attention mechanism;

[0010] Determine the control strategy of the freight train according to the energy-saving operation control model.

[0011] In some embodiments, a deep separable residual neural network based on an attention mechanism includes: a first convolutional layer, a second convolutional layer and an attention mechanism unit, wherein the second convolutional layer includes a deep convolutional layer and a point-by-point convolutional layer in sequence, and the attention mechanism unit includes a global pooling layer, a fully connected layer and an activation layer.

[0012] In some embodiments, before establishing the simulated operation environment according to the technical parameters, line parameters and operation mode of the freight train, the method further includes:

[0013] Carry out single-mass point modeling on the freight train to obtain a single-mass point model;

[0014] According to the maximum traction force of the freight train and the length of the line where the freight train is to run, the energy consumption optimization objective function is obtained;

[0015] According to the single-particle model and the energy consumption optimization objective function, the target Hamiltonian function is obtained;

[0016] The operation mode is determined based on maximizing the value of the objective Hamiltonian function.

[0017] In some embodiments, the single particle model is:

[0018]

[0019] Where v is the running speed of the freight train, x is the running distance of the freight train, t is the running time of the freight train, and D max (v) is the maximum traction force of the freight train at the running speed v, B max (v) is the maximum braking force of the freight train at the running speed v, W 0 (v) is the basic resistance of the freight train, G(x) is the slope resistance of the freight train, m is the mass of the freight train, μ d is the coefficient of maximum traction, μ b is the coefficient of maximum braking force.

[0020] In some embodiments, the energy consumption optimization objective function is:

[0021]

[0022] Among them, E is the operating energy consumption of freight trains, and S is the line length.

[0023] In some embodiments, the target Hamiltonian function is:

[0024] H=(λ-1)μ d D max (v) -λμ b B max (v)+C(v);

[0025]

[0026] Among them, λ is the accompanying variable of the freight train during the operation phase, λ 1 is a constant and M(x) is a slack variable.

[0027] In some embodiments, the operation mode is determined based on the value of the maximized objective Hamiltonian function, including:

[0028] When the operation stage is startup, λ is greater than 1, and the operation mode is to use the maximum traction force;

[0029] When the operation stage is cruising, λ is equal to 1, and the operation mode is to use partial traction;

[0030] In the operation stage, when the operating speed reaches the speed limit, λ is equal to 0, and the operation mode is to use partial braking force;

[0031] When the operation stage is coasting, λ is greater than 0 and less than 1, and the operation mode is to stop using traction and braking force;

[0032] When the operation phase is braking, λ is less than 0, and the operation mode is a mode of using the maximum braking force.

[0033] In some embodiments, determining a reward function based on the operating energy consumption of a freight train includes:

[0034] Determine the total time required for a freight train to travel the operating route;

[0035] The reward function is determined based on the energy consumption and total operating time of the freight train.

[0036] In some embodiments, before constructing the energy-saving operation control model based on reinforcement learning by using the simulated operation environment, the reward function and the preset deep Q network, the method further includes:

[0037] Determine multiple action values ​​according to a preset acceleration increment;

[0038] Determine the action space based on multiple action values;

[0039] Accordingly, the energy-saving operation control model based on reinforcement learning is constructed by using the simulated operation environment, reward function and preset deep Q network, including:

[0040] An energy-saving operation control model based on reinforcement learning is constructed by utilizing the simulated operating environment, reward function, preset deep Q network and action space.

[0041] In some embodiments, controlling the movement of the freight train according to the energy-saving operation control model includes:

[0042] Select the action value selected by the energy-saving operation control model from the action space as the initial action value;

[0043] Perform exponential smoothing on the initial action value to obtain the target action value;

[0044] Control the action of the freight train according to the target action value.

[0045] In some embodiments, the target action value is obtained through the following formula:

[0046] a t = α × a 0 + (1 - α) × a t-1 ;

[0047] Wherein, a t is the target action value corresponding to the running time t, a 0 is the initial action value, α is the smoothing factor, and a t-1 is the target action value corresponding to the running time t - 1.

[0048] According to the second aspect of the embodiments of the present application, there is provided an energy-saving control device for a freight train, including:

[0049] An environment determination module, configured to establish a simulated operation environment according to the technical parameters, line parameters, and operation mode of the freight train;

[0050] A reward function determination module, configured to determine a reward function based on the running energy consumption of the freight train;

[0051] A model determination module, configured to construct an energy-saving operation control model based on reinforcement learning by using the simulated operation environment, the reward function, and a preset deep Q network; wherein, the Q network in the preset deep Q network adopts a depthwise separable residual neural network based on an attention mechanism;

[0052] A train control module, configured to determine a control strategy for the freight train according to the energy-saving operation control model.

[0053] According to the third aspect of the embodiments of the present application, there is provided an energy-saving control device for a freight train, including a processor and a memory, the memory stores computer program instructions that can be executed by the processor, and when the processor executes the computer program instructions, the steps of the method according to any one of the first aspects are implemented.

[0054] According to the fourth aspect of the embodiments of the present application, there is provided a computer-readable storage medium, in which computer program instructions are stored, and when the computer program instructions are executed by a processor, the processor is caused to implement the steps of the method according to any one of the first aspects.

[0055] According to a fifth aspect of an embodiment of the present application, a computer program product is provided, including a computer program. When the computer program is executed by a processor, it prompts the processor to implement the steps of any method of the first aspect described above.

[0056] In this application, a simulated operation environment is established according to the technical parameters, line parameters and operation mode of the freight train; a reward function is determined based on the operating energy consumption of the freight train; an energy-saving operation control model based on reinforcement learning is constructed using the simulated operation environment, reward function and preset deep Q network; wherein the Q network in the preset deep Q network adopts a deep separable residual neural network based on the attention mechanism; and the control strategy of the freight train is determined according to the energy-saving operation control model. The technical solution provided by this application can generate a high-performance and efficient energy-saving operation control model, which can make decisions that optimize energy consumption, reduce the energy consumption of freight trains, and conform to the trend of energy-saving development of trains.

[0057] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and together with the specification, are used to explain the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings:

[0059] Figure 1 A schematic flow chart of an energy-saving control method for a freight train in one embodiment is shown;

[0060] Figure 2 A schematic diagram of the structure of a deep separable residual neural network based on an attention mechanism in one embodiment is shown;

[0061] Figure 3 A schematic diagram of the interaction between the simulation running environment and the intelligent agent in one embodiment is shown;

[0062] Figure 4 A block diagram of an energy-saving control device for a freight train in one embodiment is shown;

[0063] Figure 5 A schematic structural diagram of an energy-saving control device for a freight train in one embodiment is shown. DETAILED DESCRIPTION

[0064] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0065] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present application. However, those skilled in the art will appreciate that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, known methods, devices, realizations or operations are not shown or described in detail to avoid blurring the various aspects of the application.

[0066] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0067] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.

[0068] Figure 1 FIG. 1 is a flow chart showing a method for controlling energy saving of a freight train in one embodiment. Figure 1 As shown, a method for energy-saving control of a freight train is provided, and the method may include the following steps 101 to 104.

[0069] In step 101, a simulated operation environment is established according to technical parameters, line parameters and operation mode of the freight train.

[0070] Among them, technical parameters are parameters defined according to the attributes of the freight train itself, which may include the mass of the freight train, the relevant parameters of the traction motor, etc.; line parameters are the relevant parameters of the line on which the freight train is to run, which may include speed limit, line length, slope of the ramp and radius of the curve, etc. The operation mode refers to the way in which the freight train uses traction or braking force in different operation stages.

[0071] It is understandable that reinforcement learning is a machine learning method in which the agent learns how to perform the next action by interacting with the environment and based on the feedback from the environment. In an embodiment of the present application, the simulated operating environment is a computer-simulated operating environment of a freight train. The simulated operating environment interacts with the agent (e.g., a train controller). When the simulated operating environment receives an action from the agent, it automatically updates to the next state based on the current state and current action, and passes the next state to the reward function to obtain the corresponding reward value, and then puts the reward value and the next state into the experience storage pool. The agent obtains the reward value and the next state from the experience storage pool, and outputs the next action based on the reward value and the next state.

[0072] In some embodiments, the operation mode may be determined first according to the technical parameters and the line parameters, and then a simulated operation environment may be established according to the technical parameters, the line parameters and the operation mode.

[0073] The steps for determining the operation mode can refer to the following steps: single-particle modeling of the freight train to obtain a single-particle model; according to the maximum traction of the freight train and the line length of the freight train to be operated, an energy consumption optimization objective function is obtained; according to the single-particle model and the energy consumption optimization objective function, a target Hamiltonian function is obtained; based on maximizing the value of the target Hamiltonian function, the operation mode is determined.

[0074] It is understandable that the energy consumption of freight trains in actual operation will be affected by many factors. Through analysis and summary, the energy consumption of freight trains is mainly composed of technical parameters, line parameters, operation mode, etc. Considering that the operating energy consumption determined by the traction of freight trains accounts for more than half of the total energy consumption of freight trains, and it effectively reflects the energy consumption under different operating conditions of freight trains, the formula for the work done by the traction force of freight trains is:

[0075]

[0076] Where E is the running energy consumption of the freight train, F(v) is the traction force of the freight train at the running speed v, ΔS i is the running distance of the freight train in the section of the line to be operated, and n is the number of sections of the line to be operated.

[0077] Since the actual operating environment of freight trains is very complex, in order to simplify the calculation, freight trains can be modeled as a single particle.

[0078] From Newton's second law, the differential equation for the motion of a freight train is:

[0079]

[0080] Among them, v is the running speed of the freight train, x is the running distance of the freight train, t is the running time of the freight train, and a is the acceleration of the freight train.

[0081] Assume that the traction force function of the freight train is D(v) and the braking force function is B(v), which satisfies the following formula:

[0082]

[0083] Among them, D max (v) is the maximum traction force of the freight train at the running speed v, B max (v) is the maximum braking force of the freight train at the running speed v, μ d is the coefficient of maximum traction, μ b is the coefficient of maximum braking force.

[0084] Coefficient of maximum traction μ d and the coefficient of maximum braking force μ b The following constraints are met:

[0085]

[0086] According to the force conditions of the freight train, the acceleration of the freight train during operation is obtained as follows:

[0087]

[0088] Among them, W 0 (v) is the basic resistance of the freight train, G(x) is the slope resistance of the freight train, and m is the mass of the freight train.

[0089] Combining formulas 2, 3, and 4, we can get the single particle model:

[0090]

[0091] According to the running characteristics of freight trains, when a freight train runs on the line to be run, the starting speed and the terminal speed are both zero, the running time cannot exceed the total time required for the freight train to run the line length, and the running speed cannot exceed the maximum speed limit of the line to be run. Therefore, the constraints are:

[0092]

[0093] Among them, v max (x) is the maximum speed limit, S is the line length, and T is the total time required for a freight train to travel the length of the line.

[0094] According to formula 1 and the above constraints, the traction energy consumption of freight trains is optimized, and the following energy consumption optimization objective function can be obtained:

[0095]

[0096] Combining Formula 5 and Formula 6, according to the maximum principle, the initial Hamiltonian function is:

[0097]

[0098] Among them, λ 1 and λ 2 are all accompanying variables and satisfy the following constraints:

[0099]

[0100] Combining the above formula, we can solve for λ 1 is a constant, and the 2 The co-state equation is:

[0101]

[0102] According to the speed constraints, the slack variable M(x) is set to meet the following conditions:

[0103]

[0104] After sorting, we can get the following formula:

[0105]

[0106] in,

[0107]

[0108] Combining Formula 9 and Formula 10, we can get the target Hamiltonian function:

[0109] H=(λ-1)μ d D max (v) -λμ b B max (v)+C(v) Formula 11;

[0110] Among them, λ is the accompanying variable of the freight train during the operation stage.

[0111] According to the maximum principle, to minimize the traction energy consumption, the target Hamiltonian function should be maximized. When λ is greater than or equal to 1, the traction force of the freight train exists and the braking force is 0. Among them, when λ is greater than 1, the freight train is in full traction state, and when λ is equal to 1, 0<μ d <1, the freight train is in a partial traction state, and the target Hamiltonian function can be simplified to:

[0112] H=(λ-1)μd D max (v)+C(v) Formula twelve.

[0113] When λ is less than or equal to 0, the braking force of the freight train exists and the traction force is 0. b =1, the freight train is in full braking state, when λ is equal to 0, 0<μ b <1, the freight train is in a partial braking state, and the target Hamiltonian function can be simplified to:

[0114] H=-λμ b B max (v)+C(v) Formula 13.

[0115] When λ is greater than 0 and less than 1, the traction and braking force of the freight train do not exist. At this time, the target Hamiltonian function needs to obtain the maximum value, μ d =μ b =0, the freight train is in an inert state.

[0116] According to the above kinematic equations of the freight train and the maximum principle, it can be deduced that the operation mode of the freight train can be: when the operation stage is starting, the maximum traction force is used to make the freight train increase the operation speed as quickly as possible. During this process, the energy conversion mode of the freight train is the conversion of electrical energy into kinetic energy; when the operation stage is cruising, partial traction force is used. At this time, the resultant force on the freight train is 0, and the freight train enters a uniform speed operation state. During this process, the energy conversion mode of the freight train is the conversion of electrical energy into kinetic energy. When the running speed of the freight train reaches the maximum speed limit, the freight train adopts partial braking force to keep the speed of the freight train stable and enter cruising operation. This process does not consume electrical energy; when the operation stage is coasting, the freight train is only subjected to running resistance, and no electrical energy is consumed at this time; when the operation stage is braking, the maximum braking force is used to make the speed of the freight train drop sharply to 0, and this process does not consume electrical energy.

[0117] In step 102, a reward function is determined based on the operating energy consumption of the freight train.

[0118] It is understandable that, according to the energy loss of freight trains during operation and the execution of waybills, when determining the reward function, various factors can be considered for design. For example, the lower the energy consumption and the more accurate the time punctuality, the greater the value of the reward function.

[0119] In some embodiments, the total time required for a freight train to travel the operating route may be determined; and a reward function may be determined based on the operating energy consumption and the total time of the freight train.

[0120] It is understood that the total time required for a freight train to travel the length of the line is:

[0121]

[0122] From the energy consumption analysis of freight trains, it can be seen that by converting the optimization target of running energy consumption into a form with a penalty function, the following formula can be obtained:

[0123]

[0124] Among them, μ 2 is the coefficient corresponding to the total duration.

[0125] Expanding the penalty function in Formula 15, we can get the following formula:

[0126]

[0127] In the formula, is the actual running time of the freight train. Substituting into Formula 16, the simplified reward function can be obtained as follows:

[0128]

[0129] Through the design of the above reward function, the simulated operation environment can accurately calculate the reward value of the action, which provides a prerequisite for the subsequent energy-saving operation control model to select appropriate actions and obtain correct convergence.

[0130] In step 103, an energy-saving operation control model based on reinforcement learning is constructed using a simulated operating environment, a reward function and a preset deep Q network; wherein the Q network in the preset deep Q network adopts a deep separable residual neural network based on an attention mechanism.

[0131] It is understandable that the Q network in the preset deep Q network usually adopts CNN or RNN architecture, which requires a large number of model parameters, not only has a large amount of calculation but also has poor performance. In the embodiment of the present application, the Q network adopts a deep separable residual neural network based on the attention mechanism, which can effectively solve this problem.

[0132] Figure 2 FIG. 1 shows a schematic diagram of the structure of a deep separable residual neural network based on an attention mechanism in one embodiment. Figure 2 As shown, in some embodiments, a deep separable residual neural network based on an attention mechanism includes: a first convolutional layer, a second convolutional layer and an attention mechanism unit, wherein the second convolutional layer sequentially includes a deep convolutional layer and a point-by-point convolutional layer, and the attention mechanism unit includes a global pooling layer, a fully connected layer and an activation layer.

[0133] It should be noted that the core of the attention mechanism is to enhance useful features and suppress unimportant features. The attention mechanism unit performs global average pooling on each feature channel, learns the activation of specific samples, uses the Relu function to reduce the complexity of the network, and uses the Sigmoid function to learn the weight of each channel, so as to learn to use global information to selectively emphasize information features and suppress less useful features. In the deep separable residual neural network based on the attention mechanism of the embodiment of the present application, the second convolutional layer is a deep separable convolutional layer, which can be divided into two steps: deep convolution (Dw) and point-by-point convolution (Pw). Such a design can reduce network parameters and computational complexity without significantly reducing performance. Residual connections can enable the entire network to maintain a strong ability to learn feature representations, so that the network can learn more complex feature representations and improve the performance of the network.

[0134] In some embodiments, the activation layer uses Relu as the activation function. When the input value of this function is greater than zero, the gradient value is a constant, the convergence speed of the network is accelerated, the network training time is reduced, and the calculation amount of parameters is reduced, avoiding the problem of gradient disappearance.

[0135] In some embodiments, Adam can be used as an optimizer, which combines the momentum and adaptive learning rate algorithms, improves the training effect by adjusting the learning rate and momentum parameters, optimizes and updates the parameters in the neural network, and has a faster convergence speed and better performance.

[0136] It is understandable that to build an energy-saving operation control model based on reinforcement learning, in addition to designing a suitable reward function and neural network, it is also necessary to design the action space and state space.

[0137] When designing the state space, according to the analysis of the actual operation of freight trains, it can be seen that the state information of freight trains mainly includes running speed v, running distance x, running time t and acceleration a. Therefore, the state information can be designed as a four-dimensional vector and the state information can be normalized to make the algorithm better trained in the neural network.

[0138] When designing the action space, multiple action values ​​can be determined according to the preset acceleration increment; and the action space can be determined according to the multiple action values.

[0139] After determining the state space, action space, reward function and preset deep Q network, we can use the simulated operating environment, reward function, preset deep Q network, action space and state space to build an energy-saving operation control model based on reinforcement learning, so that the intelligent agent and the simulated operating environment can interact to obtain rewards, and continuously optimize its own decisions to output the optimal decision.

[0140] Figure 3 FIG. 1 shows a schematic diagram of the interaction between the simulation running environment and the intelligent agent in one embodiment. Figure 3 As shown, the preset deep Q network includes a current value network (ie, Q network) and a target value network. Both the current value network and the target value network can adopt a deep separable residual neural network based on an attention mechanism, but the parameters of the current value network and the target value network are different.

[0141] As mentioned above, the experience storage pool stores reward values ​​and corresponding states. The intelligent agent can randomly extract a batch of samples from the experience storage pool to train the current value network. The current value network outputs the next action to the simulated operation environment. The simulated operation environment stores the reward value and state corresponding to the next action in the experience storage pool. This cycle is repeated until the function loss gradient between the value predicted by the current value network and the value predicted by the target network reaches the set condition, and the training is terminated to obtain a trained energy-saving operation control model.

[0142] In step 104, a control strategy for the freight train is determined according to the energy-saving operation control model.

[0143] It is understandable that the control strategy may be an action that the freight train needs to take, and the action may be achieved through indicators such as the acceleration of the freight train.

[0144] In some embodiments, the action value selected from the action space by the energy-saving operation control model can be used as the initial action value; the initial action value is subjected to exponential smoothing to obtain a target action value; and the action of the freight train is controlled according to the target action value.

[0145] The target action value is obtained by the following formula:

[0146] a t =α×a 0 +(1-α)×a t-1 Formula XVIII;

[0147] Among them, a t is the target action value corresponding to the running time t, a 0 Initial action value, α is the smoothing factor, a t-1 is the target action value corresponding to the running time t-1.

[0148] It is understandable that, considering that the action of the freight train will change according to the operation stage of the freight train, in order to avoid excessive fluctuations in the state of the freight train, the selected action value can be exponentially smoothed to make the speed curve of the freight train during operation smoother.

[0149] During actual use, the actual operation information of the freight train can be input into the energy-saving operation control model, and the energy-saving operation control model will output the action of the freight train, which can not only ensure the stability and speed limit conditions of the freight train, but also follow the principle of minimum energy consumption.

[0150] The embodiment of the present application establishes a simulated operation environment according to the technical parameters, line parameters and operation mode of the freight train; determines the reward function based on the operation energy consumption of the freight train; uses the simulated operation environment, reward function and preset deep Q network to construct an energy-saving operation control model based on reinforcement learning; wherein the Q network in the preset deep Q network adopts a deep separable residual neural network based on the attention mechanism; and determines the control strategy of the freight train according to the energy-saving operation control model. The technical solution provided by the present application can generate a high-performance and high-efficiency energy-saving operation control model, which can make decisions that optimize energy consumption, reduce the energy consumption of freight trains, and conform to the trend of energy-saving development of trains.

[0151] The following describes an embodiment of the device of the present application, which can be used to execute the energy-saving control method for freight trains in the above embodiment of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the energy-saving control method for freight trains in the above embodiment of the present application.

[0152] Figure 4 FIG. 1 is a block diagram of an energy-saving control device for a freight train in one embodiment. Figure 4 As shown, the energy-saving control device of the freight train in the embodiment of the present application includes: an environment determination module 401, which is used to establish a simulated operating environment according to the technical parameters, line parameters and operating mode of the freight train; a reward function determination module 402, which is used to determine the reward function based on the operating energy consumption of the freight train; a model determination module 403, which is used to use the simulated operating environment, the reward function and the preset deep Q network to construct an energy-saving operation control model based on reinforcement learning; wherein, the Q network in the preset deep Q network adopts a deep separable residual neural network based on the attention mechanism; a train control module 404, which is used to determine the control strategy of the freight train according to the energy-saving operation control model.

[0153] In some embodiments, a deep separable residual neural network based on an attention mechanism includes: a first convolutional layer, a second convolutional layer and an attention mechanism unit, wherein the second convolutional layer includes a deep convolutional layer and a point-by-point convolutional layer in sequence, and the attention mechanism unit includes a global pooling layer, a fully connected layer and an activation layer.

[0154] In some embodiments, the energy-saving control device of the freight train may also include an operation mode determination module (not shown), which is used to perform single-particle modeling on the freight train to obtain a single-particle model; obtain an energy consumption optimization objective function based on the maximum traction force of the freight train and the line length of the freight train to be operated; obtain a target Hamiltonian function based on the single-particle model and the energy consumption optimization objective function; and determine the operation mode based on maximizing the value of the target Hamiltonian function.

[0155] In some embodiments, in some embodiments, the single particle model is:

[0156]

[0157] Where v is the running speed of the freight train, x is the running distance of the freight train, t is the running time of the freight train, and D max (v) is the maximum traction force of the freight train at the running speed v, B max (v) is the maximum braking force of the freight train at the running speed v, W 0 (v) is the basic resistance of the freight train, G(x) is the slope resistance of the freight train, m is the mass of the freight train, μ d is the coefficient of maximum traction, μ b is the coefficient of maximum braking force.

[0158] In some embodiments, the energy consumption optimization objective function is:

[0159]

[0160] Among them, E is the operating energy consumption of freight trains, and S is the line length.

[0161] In some embodiments, the target Hamiltonian function is:

[0162] H=(λ-1)μ d D max (v) -λμ b B max (v)+C(v);

[0163]

[0164] Among them, λ is the accompanying variable of the freight train during the operation phase, λ 1 is a constant and M(x) is a slack variable.

[0165] In some embodiments, the operating mode module is also used for: when the operating stage is starting, λ is greater than 1, and the operating mode is to use the maximum traction force; when the operating stage is cruising, λ is equal to 1, and the operating mode is to use partial traction force; when the operating speed reaches the speed limit, λ is equal to 0, and the operating mode is to use partial braking force; when the operating stage is coasting, λ is greater than 0 and less than 1, and the operating mode is to stop using traction and braking force; when the operating stage is braking, λ is less than 0, and the operating mode is to use the maximum braking force.

[0166] In some embodiments, the reward function determination module 402 is further used to determine the total time required for the freight train to complete the operation route; and determine the reward function according to the operation energy consumption and the total operation time of the freight train.

[0167] In some embodiments, the energy-saving control device of the freight train may also include an action space determination module (not shown), which determines multiple action values ​​according to a preset acceleration increment; determines the action space according to the multiple action values; accordingly, the model determination module 403 can also be used to construct an energy-saving operation control model based on reinforcement learning by using a simulated operating environment, a reward function, a preset deep Q network and an action space.

[0168] In some embodiments, the train control module 404 can also be used to use the action value selected by the energy-saving operation control model from the action space as the initial action value; perform exponential smoothing on the initial action value to obtain the target action value; and control the action of the freight train according to the target action value.

[0169] In some embodiments, the target action value is obtained by the following formula:

[0170] a t =α×a 0 +(1-α)×a t-1 ;

[0171] Among them, a t is the target action value corresponding to the running time t, a 0 Initial action value, α is the smoothing factor, a t-1 is the target action value corresponding to the running time t-1.

[0172] Based on the same inventive concept, the embodiment of the present application also provides an energy-saving control device for a freight train, referring to Figure 5, showing a schematic structural diagram of an energy-saving control device for a freight train in an embodiment of the present application, wherein the control device for the freight train includes one or more memories 504, one or more processors 502, and at least one computer program (computer program instruction) stored in the memories 504 and executable on the processors 502, and the processor 502 implements the method described above when executing the computer program.

[0173] Among them, Figure 5 In the embodiment of the present invention, a bus architecture (represented by bus 500) is shown, which may include any number of interconnected buses and bridges, and bus 500 links various circuits including one or more processors represented by processor 502 and memory represented by memory 504. Bus 500 may also link various other circuits such as peripherals, voltage regulators, and power management circuits, which are well known in the art and are therefore not further described herein. Bus interface 505 provides an interface between bus 500 and receiver 501 and transmitter 503. Receiver 501 and transmitter 503 may be the same element, namely a transceiver, providing a unit for communicating with various other devices over a transmission medium. Processor 502 is responsible for managing bus 500 and general processing, while memory 504 may be used to store data used by processor 502 when performing operations.

[0174] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, in which computer program instructions are stored. When the computer program instructions are executed by a processor, the processor is prompted to implement the steps of the method as described above.

[0175] Based on the same inventive concept, an embodiment of the present application provides a computer program product, including a computer program. When the computer program product is executed by a processor, it prompts the processor to implement the steps of the method as described above.

[0176] The functions described herein may be implemented in hardware, software executed by a processor, firmware, or any combination thereof. If implemented in software executed by a processor, the functions may be stored as one or more instructions or codes on a computer-readable medium or transmitted via a computer-readable medium. Other examples and implementations are within the scope and spirit of the present application and the appended claims. For example, due to the nature of software, the functions described above may be implemented using software executed by a processor, hardware, firmware, hard wiring, or a combination of any of these. In addition, each functional unit may be integrated into a processing unit, each unit may exist physically separately, or two or more units may be integrated into one unit.

[0177] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units can be a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0178] The units described as separate components may or may not be physically separated, and the components of the control device may or may not be physical units, that is, they may be located in one place or distributed in multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0179] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or all or part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk, etc., various media that can store computer program instructions.

[0180] The above description is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present application shall be included in the scope of the claims of the present application.

Claims

1. A method for energy-saving control of a freight train, characterized in that: include: Establish a simulated operation environment based on the technical parameters, line parameters and operation mode of freight trains; Determining a reward function based on the operating energy consumption of the freight train; Using the simulated operating environment, the reward function and the preset deep Q network, an energy-saving operation control model based on reinforcement learning is constructed; wherein the Q network in the preset deep Q network adopts a deep separable residual neural network based on an attention mechanism; A control strategy for the freight train is determined according to the energy-saving operation control model.

2. The energy-saving control method for freight trains according to claim 1, characterized in that: The deep separable residual neural network based on the attention mechanism includes: a first convolutional layer, a second convolutional layer and an attention mechanism unit, wherein the second convolutional layer includes a deep convolutional layer and a point-by-point convolutional layer in sequence, and the attention mechanism unit includes a global pooling layer, a fully connected layer and an activation layer.

3. The energy-saving control method for freight trains according to claim 1, characterized in that: Before establishing the simulated operation environment according to the technical parameters, line parameters and operation mode of the freight train, the method further includes: Performing single-particle modeling on the freight train to obtain a single-particle model; Obtaining an energy consumption optimization objective function according to the maximum traction force of the freight train and the length of the line on which the freight train is to run; Obtaining a target Hamiltonian function according to the single-particle model and the energy consumption optimization objective function; The operation mode is determined based on maximizing the value of the objective Hamiltonian function.

4. The energy-saving control method for freight trains according to claim 3, characterized in that: The single particle model is: Wherein, v is the running speed of the freight train, x is the running distance of the freight train, t is the running time of the freight train, and D max (v) is the maximum traction force of the freight train at the running speed v, B max (v) is the maximum braking force of the freight train at the running speed v, W0(v) is the basic resistance of the freight train, G(x) is the slope resistance of the freight train, m is the mass of the freight train, μ d is the coefficient of the maximum traction force, μ b is the coefficient of the maximum braking force.

5. The energy-saving control method for freight trains according to claim 4, characterized in that: The energy consumption optimization objective function is: Wherein, E is the operating energy consumption of the freight train, and S is the length of the line.

6. The energy-saving control method for freight trains according to claim 5, characterized in that: The target Hamiltonian function is: H=(λ-1)μ d D max (v)-lm b B max (v)+C(v); Among them, λ is the accompanying variable of the freight train in the operation stage, λ1 is a constant, and M(x) is a slack variable.

7. The energy-saving control method for freight trains according to claim 6, characterized in that: The step of determining the operation mode based on maximizing the value of the target Hamiltonian function includes: When the operation stage is startup, λ is greater than 1, and the operation mode is a mode of using maximum traction; When the operation stage is cruising, λ is equal to 1, and the operation mode is a mode of using partial traction; In the operation phase, when the operation speed reaches the speed limit, λ is equal to 0, and the operation mode is a mode of using partial braking force; When the operation stage is coasting, λ is greater than 0 and less than 1, and the operation mode is to stop using traction and braking force; When the operation phase is braking, λ is less than 0, and the operation mode is a mode using maximum braking force.

8. The energy-saving control method for freight trains according to claim 1, characterized in that: The step of determining a reward function based on the running energy consumption of the freight train comprises: Determine the total time required for the freight train to complete the operation route; The reward function is determined according to the operating energy consumption of the freight train and the total duration.

9. The energy-saving control method for freight trains according to claim 1, characterized in that: Before constructing the energy-saving operation control model based on reinforcement learning by using the simulated operation environment, the reward function and the preset deep Q network, the method further includes: Determine multiple action values ​​according to a preset acceleration increment; Determining an action space according to the multiple action values; Accordingly, the energy-saving operation control model based on reinforcement learning is constructed by utilizing the simulated operation environment, the reward function and the preset deep Q network, including: An energy-saving operation control model based on reinforcement learning is constructed by utilizing the simulated operation environment, the reward function, the preset deep Q network and the action space.

10. The energy-saving control method for freight trains according to claim 9, characterized in that: Determining the control strategy of the freight train according to the energy-saving operation control model includes: using the action value selected by the energy-saving operation control model from the action space as the initial action value; Performing exponential smoothing on the initial action value to obtain a target action value; The movement of the freight train is controlled according to the target movement value.

11. The energy-saving control method for a freight train according to claim 1, characterized in that: The target action value is obtained by the following formula: a t =α×a0+(1-α)×a t-1 ; Among them, a t is the target action value corresponding to the running time t, a0 is the initial action value, α is the smoothing factor, and a t-1 is the target action value corresponding to the running time t-1.

12. An energy-saving control device for a freight train, characterized in that: include: An environment determination module is used to establish a simulated operation environment according to the technical parameters, line parameters and operation mode of the freight train; A reward function determination module, used to determine a reward function based on the running energy consumption of the freight train; A model determination module, used to construct an energy-saving operation control model based on reinforcement learning by using the simulated operation environment, the reward function and a preset deep Q network; wherein the Q network in the preset deep Q network adopts a deep separable residual neural network based on an attention mechanism; The train control module is used to determine the control strategy of the freight train according to the energy-saving operation control model.

13. An energy-saving control device for a freight train, comprising a processor and a memory, characterized in that: The memory stores computer program instructions that can be executed by the processor, and when the processor executes the computer program instructions, the steps of the method according to any one of claims 1 to 11 are implemented.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, which, when executed by a processor, prompt the processor to implement the steps of the method according to any one of claims 1 to 11.

15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the processor is prompted to implement the steps of the method according to any one of claims 1 to 11.