Train speed tracking control method, device, equipment, storage medium and product

By combining an improved GRU network and reinforcement learning algorithm with a model predictive controller, the adaptability problem of traditional train speed control methods in complex systems is solved, achieving more precise train speed tracking control and improving stability and the ability to cope with disturbances.

CN121516077APending Publication Date: 2026-02-13QINGDAO UNIV OF TECH +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511980847.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Traditional train speed control methods struggle to achieve adaptive optimization control when faced with complex nonlinear systems and external disturbances, resulting in compromised control performance.

Method used

An improved GRU network is used to train a train data-driven model. Combined with a model predictive controller and reinforcement learning algorithm, the train speed is predicted through multiple time steps. Reinforcement learning is used to correct the optimal control sequence, thereby achieving precise train speed tracking control.

Benefits of technology

It improves the stability and accuracy of train control, effectively responds to delays or disturbances, and enhances the overall control performance of train operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121516077A_ABST
    Figure CN121516077A_ABST
Patent Text Reader

Abstract

The invention provides a train speed tracking control method, device and equipment, a storage medium and a product, and relates to the technical field of train automatic control. The method comprises the following steps: acquiring a basic data set of a target train, wherein the basic data set is used for representing nonlinear dynamic information of the target train; using the basic data set to train and generate a train data driving model based on the improved GRU network; a train data driving model is used as a prediction model of a model prediction controller, and train speed prediction sequences of multiple time steps are obtained based on the basic data set; establishing a train state space model based on the kinetic equation, taking the train state space model as a controlled object of model prediction control and an environment of reinforcement learning, and outputting an optimal control quantity sequence based on a train speed prediction sequence by using a model prediction controller; correcting the optimal control quantity sequence output by the model prediction controller by using a reinforcement learning algorithm to obtain a corrected optimal control quantity sequence; and performing tracking control on the target train based on the corrected optimal control quantity sequence.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of train automatic control, and in particular, to a train speed tracking control method, device, equipment, storage medium and product. BACKGROUND

[0002] As an important part of modern rail transportation, the safety, stability and energy consumption optimization of high-speed trains are the hot issues of current research. Train speed tracking control is a key link to ensure the planned operation of trains, improve the comfort of passengers and reduce energy consumption. The traditional train speed control method mainly relies on proportional-integral-derivative control (PID), fuzzy control or optimization control and other classical control theories. These methods can achieve good results under certain conditions, but when facing complex nonlinear systems, external disturbances and uncertain factors, their control performance is easily affected, and it is difficult to achieve adaptive optimal control. SUMMARY

[0003] The present disclosure relates to the field of train automatic control, and in particular, to a train speed tracking control method, device, equipment, storage medium and product.

[0004] In a first aspect, the present disclosure provides a train speed tracking control method, comprising: obtaining a basic data set of a target train, the basic data set being used to represent the nonlinear dynamics information of the target train; using the basic data set to train a train data-driven model based on an improved GRU network; using the train data-driven model as a prediction model of a model predictive controller, obtaining a train speed prediction sequence at multiple time steps based on the basic data set; establishing a train state space model based on a dynamics equation, as a controlled object of the model predictive control and an environment of reinforcement learning, using the model predictive controller to output an optimal control amount sequence based on the train speed prediction sequence; correcting the optimal control amount sequence using a reinforcement learning algorithm to obtain a corrected optimal control amount sequence; and tracking the target train based on the corrected optimal control amount sequence.

[0005] In a second aspect, an embodiment of the present application provides a train speed tracking control device, comprising: a data acquisition module configured to acquire a basic data set of a target train, the basic data set being used to represent nonlinear dynamic information of the target train; a train data-driven model construction module configured to generate a train data-driven model based on an improved GRU network training using the basic data set; a speed prediction module configured to use the train data-driven model as a prediction model of a model predictive controller, and obtain a train speed prediction sequence at multiple time steps based on the basic data set; a control quantity determination module configured to establish a train state space model based on a dynamic equation, as a controlled object of the model predictive control and an environment of reinforcement learning, and output an optimal control quantity sequence based on the train speed prediction sequence using the model predictive controller; a control quantity sequence correction module configured to correct the optimal control quantity sequence using a reinforcement learning algorithm to obtain a corrected optimal control quantity sequence; and a control module configured to perform tracking control on the target train based on the corrected optimal control quantity sequence.

[0006] In a third aspect, an embodiment of the present application provides an electronic device, comprising: at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to implement the train speed tracking control method as described in any of the implementation manners of the first aspect.

[0007] In a fourth aspect, an embodiment of the present application provides a non-transitory computer-readable storage medium storing computer instructions for enabling a computer to implement the train speed tracking control method as described in any of the implementation manners of the first aspect.

[0008] In a fifth aspect, an embodiment of the present application provides a computer program product comprising a computer program, which, when executed by a processor, can implement the train speed tracking control method as described in any of the implementation manners of the first aspect.

[0009] The train speed tracking control method, device, equipment, storage medium and product provided by the embodiments of the present application can train a train data-driven model based on data of a train running process, can more accurately reflect complex characteristics of an actual system, can predict a running speed of the train at a future time through the train data-driven model, and can track and control the running of the train at the future time through a model predictive controller in combination with the predicted running speed of the train, can take into account future states through multi-time-step prediction, can effectively cope with delays or disturbances, and can improve stability and accuracy of train control.

[0010] It should be understood that the contents described in this part are not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent through the following description. BRIEF DESCRIPTION OF DRAWINGS

[0011] Other features, objects, and advantages of the present disclosure will become more apparent from the following detailed description of non-limiting embodiments made with reference to the drawings:

[0012] Figure 1 is an exemplary system architecture to which the present disclosure can be applied;

[0013] Figure 2 A flow chart of a train speed tracking control method provided by an embodiment of the present disclosure;

[0014] Figure 3 and Figure 4 A flow chart of another train speed tracking control method provided by an embodiment of the present disclosure;

[0015] Figure 5 A flow chart of another train speed tracking control method provided by an embodiment of the present disclosure;

[0016] Figure 6 A structural block diagram of a train speed tracking control device provided by an embodiment of the present disclosure;

[0017] Figure 7 A structural schematic diagram of an electronic device suitable for executing a train speed tracking control method provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0018] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to assist in understanding, which should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description. It should be noted that the embodiments in the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0019] In the technical solutions of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in the technical solutions comply with the relevant legal regulations and do not violate public order and good customs.

[0020] Figure 1 An exemplary system architecture 100 to which embodiments of a train speed tracking control method, device, electronic device and computer readable storage medium of the present disclosure can be applied is shown.

[0021] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0022] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various applications for enabling information communication between the terminal devices 101, 102, and 103 and server 105 can be installed. These applications include instant messaging applications.

[0023] Terminal devices 101, 102, and 103 and server 105 can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with displays, including but not limited to smartphones, tablets, laptops, and desktop computers. When terminal devices 101, 102, and 103 are software, they can be installed in the aforementioned electronic devices, and can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here. When server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When server 105 is software, it can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here.

[0024] Server 105 can provide various services through its built-in applications. It should be noted that the data or information required to provide these services can be obtained from terminal devices 101, 102, and 103 via network 104, or it can be pre-stored locally on server 105 through various means. Therefore, when server 105 detects that this data is already stored locally, it can choose to retrieve it directly from the local storage. In this case, the exemplary system architecture 100 may not include terminal devices 101, 102, and 103 and network 104.

[0025] Since processing the corresponding data or information may require significant computing resources and power, the train speed tracking control method provided in the subsequent embodiments of this disclosure is generally executed by a server 105 with strong computing power and abundant computing resources. Correspondingly, the train speed tracking control device is also generally located within the server 105. However, it should also be noted that when terminal devices 101, 102, and 103 also possess sufficient computing power and resources, they can also perform the aforementioned calculations performed by the server 105 through their installed applications, thereby outputting the same results as the server 105. Especially when multiple terminal devices with different computing capabilities exist simultaneously, but the relevant application determines that the terminal device has strong computing power and abundant remaining computing resources, the terminal device can perform the aforementioned calculations, thereby appropriately reducing the computing pressure on the server 105. Accordingly, the train speed tracking control device can also be located within the terminal devices 101, 102, and 103. In this case, the exemplary system architecture 100 may also exclude the server 105 and the network 104.

[0026] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0027] Please refer to Figure 2 , Figure 2 A flowchart of a train speed tracking control method provided in an embodiment of the present invention, wherein process 200 includes the following steps:

[0028] Step 201: Obtain the basic dataset of the target train, which is used to characterize the nonlinear dynamic information of the target train.

[0029] This step is intended for the implementer of the train speed tracking control method (e.g., Figure 1 The server 105 shown obtains the basic dataset of the target train. Exemplarily, this basic dataset includes, but is not limited to, speed information, position information, acceleration information, commands corresponding to traction and braking forces, gradient, and gradient length data. This basic dataset can be used to comprehensively characterize the train's nonlinear dynamics information. This data is typically stored in the onboard controller, which stores a large amount of data including time, system number, data integrity, gradient, analog output, target speed, load, network current, and network voltage. The onboard controller stores data at 200ms intervals. Analog signals refer to the analog output of the onboard controller. For example, when a driver applies the brakes, corresponding digital and analog signals are generated. In this embodiment, the analog signal, i.e., the analog output of the onboard controller, is collected.

[0030] Step 202: Train a train data-driven model based on the basic dataset using an improved GRU network.

[0031] This step aims to train a train data-driven model using an improved GRU-based deep learning method as a prediction model for model predictive control. This method integrates four consecutive modules, namely a CNN-based data feature learning module, a gated recurrent unit (GRU)-based encoder module, an attention mechanism module, and a GRU-based decoder module, to realize multi-step train speed prediction.

[0032] Step 203: Obtain a train speed prediction sequence for multiple time steps based on the basic dataset using the train data-driven model as a model predictive controller.

[0033] This step aims to input the data in the basic dataset into the model predictive controller embedded with the train data-driven model by the aforementioned executing subject to obtain a speed prediction sequence for the target train at the next multiple time steps. In this embodiment, the data in the basic dataset can be divided into a training set and a test set to train the train data-driven model and obtain a model for train speed prediction.

[0034] For example, real-time speed information, position information, and acceleration information in the basic dataset can be extracted first. Then, the current state quantity is obtained based on the real-time speed information, position information, and acceleration information and the control quantity at the previous time. Then, the current state quantity is input into the train data-driven model to obtain a train speed prediction sequence for multiple time steps.

[0035] Step 204: Establish a train state space model based on the dynamic equation as the controlled object of model predictive control and the environment of reinforcement learning, and output an optimal control quantity sequence based on the train speed prediction sequence using the model predictive controller.

[0036] This step aims to establish a train state space model based on the dynamic equation by the aforementioned executing subject as the controlled object of model predictive control and the environment of reinforcement learning, and then use the model predictive controller (MPC) to iteratively predict the system state at future time steps based on the train speed prediction sequence obtained in the previous step, and output an optimal control quantity sequence, thereby forming a rolling horizon control. The optimal control quantity in the optimal control quantity sequence is used to drive the operating speed of the target train, etc.

[0037] Step 205: Correct the optimal control quantity sequence using a reinforcement learning algorithm to obtain a corrected optimal control quantity sequence.

[0038] In the embodiment, the model predictive controller is combined with the reinforcement learning algorithm, and the obtained optimal control quantity sequence is further corrected to improve the accuracy of the control of the target train.

[0039] Step 206: The target train is tracked and controlled based on the optimal control quantity sequence.

[0040] The step is designed to implement the optimal control quantity in the optimal control quantity sequence in the target train by the above execution subject, and the speed of the target train can be tracked and controlled again by the above process at the next time to realize real-time adjustment of the optimal control quantity.

[0041] The train speed tracking control method provided by the embodiment can train a train data-driven model based on the data of the train running process, can more accurately reflect the complex characteristics of the actual system, can predict the running speed of the train at the future time through the train data-driven model, and can track and control the running of the train at the future time through the model predictive controller in combination with the predicted train running speed, can consider the future state through multi-time step prediction, can effectively cope with delay or disturbance, and can improve the stability and accuracy of train control.

[0042] In some optional embodiments of the embodiment, as shown in Figure 3 and Figure 4 , the process of correcting the optimal control quantity sequence by the reinforcement learning algorithm in the above step 205 mainly includes:

[0043] Step 301: A single-particle model is used to describe the target train, the target train is analyzed under force, and the dynamic equation of the target train is constructed.

[0044] In the embodiment, first, a reinforcement learning environment is established, a single-particle model is used to describe the train, and the train is analyzed under force as follows:

[0045] ,

[0046] Among them, is the mass of the train; is the acceleration of the train; is the traction or braking force of the train; is the basic resistance of the train; is the additional resistance of the train.

[0047] The basic resistance of the train refers to the force that always hinders the running of the train, such as wheel-rail friction, impact vibration resistance, air resistance, and pressure difference between the train head and the train tail. Davis formula is used for approximation, and the formula is as follows:

[0048] ,

[0049] where, is the Davis coefficient, The value of is affected by the train friction and the wheel-rail wear condition, etc. is greatly affected by the air resistance; is the train speed; is the gravity acceleration. Since the basic resistance of the train is nonlinear, the basic resistance is linearized by performing a first-order Taylor expansion, and the formula is:

[0050] ,

[0051] where, is the current train speed; is the remainder of the Taylor formula, which is a high-order infinitesimal of .

[0052] The additional resistance of the train is , where, is the slope additional resistance; is the curve additional resistance; is the tunnel additional resistance. The tunnel additional resistance of the train is not considered, i.e. . The slope additional resistance is the component of the train gravity in the slope direction, and its calculation formula is , where, is the slope value of the slope. The curve additional resistance is the resistance generated by the friction between the train wheels and the track when the train runs on the curve, and its formula is:

[0053] ,

[0054] where, is the curve radius of the line.

[0055] Based on the above force molecules, the dynamic equation of the train can be constructed as:

[0056] ,

[0057] where, are the position, speed, and actual acceleration of the train at time , respectively; is the expected acceleration of the train at time ; is the lag coefficient, which represents that the train traction / braking controller can only reach the expected acceleration after a delay time .

[0058] ​​​Step 302: generating a state space model of the target train based on the dynamic equation.

[0059] After obtaining the dynamic equation of the target train, a state space model of the target train can be generated based on the dynamic equation, specifically:

[0060] ,

[0061] wherein, is the control input of the train, is the state of the train at time .

[0062] The train state space model of the above formula can be discretized according to the sampling time ,

[0063] ,

[0064] .

[0065] Step 303: using the actor-critic algorithm to obtain the expected value of the action quantity and the cumulative reward based on the state space model.

[0066] In the embodiment of the present application, the process realized in combination with the actor-critic algorithm mainly includes:

[0067] Step one: input the state quantity of the state space model into the actor network to obtain the action quantity; the state quantity includes the speed error.

[0068] In this embodiment, the reinforcement learning algorithm is realized based on the critic-actor network. The state quantity of the state space model is input into the actor network to obtain the action quantity. The actor network input is the state , and the output is the action , which is composed of a multi-layer fully connected neural network and activation function. The state is set to the speed error , and the current control input .

[0069] Step two: determining the reward function based on the optimal control quantity sequence and the speed error.

[0070] The reward function is set to , is a hyperparameter for measuring the tracking error and energy consumption, which controls the relative importance of the two.

[0071] Step three: using the critic network to obtain the expected value of the cumulative reward based on the state-action pair composed of the state quantity and the action quantity and the reward function.

[0072] critic network The input is a state-action pair , and the output is the expected value of cumulative reward, i.e. value , which is composed of a multi-layer fully connected neural network, and outputs a scalar. In addition, the respective target networks and are also needed to stabilize training.

[0073] Step 304: Correct the optimal control sequence based on the action amount to obtain the actual control input sequence. Step 305: Adjust the actual control input sequence of the action amount based on the expected value to obtain the corrected optimal control sequence.

[0074] In this embodiment, first, initialize the actor network parameters and the critic network parameters ; copy the parameters to the target networks , ; initialize the experience replay buffer . For each time step , get the current state , calculate the action , where is the exploration noise; based on the control input obtained by the MPC module in the foregoing process, combined with the correction term output by the actor network, the actual control input is obtained.

[0075] Further, apply to the high-speed train system, and observe the next state and obtain the reward . Convert to the experience replay buffer . Randomly sample a small batch of samples from the buffer; for each sample, calculate the target value , where is the discount factor, between 0 and 1, measuring the influence of future rewards on current decisions, and the value is closer to 1, the higher the weight of future rewards; is the target critic network used to calculate a more stable target value.

[0076] Then minimize the mean square error loss:

[0077] ,

[0078] where is the number of small batch samples (sampled from the experience replay buffer); is the target value.

[0079] Update actor network parameters using policy gradient:

[0080] ,

[0081] Through the update process, the actor network is moving towards a higher level. The direction of the value is adjusted by the parameter.

[0082] Furthermore, soft updates can be used to update network parameters to improve training stability. , ,in This is the soft update coefficient, typically set to 0.001, ensuring that the target network slowly follows the changes in the main network, thus smoothing the training process. Repeat this process until training converges or the predetermined training period is reached.

[0083] The train speed tracking control method provided in this invention integrates model predictive control and reinforcement learning models, combining the advantages of both to improve the overall performance of high-speed train speed tracking control. Reinforcement learning primarily focuses on data-driven policy adaptation capabilities, while model predictive control emphasizes efficient control through rolling time-domain optimization and precise handling of system constraints. The tracking control scheme based on this fusion of methods can utilize reinforcement learning to learn the dynamic characteristics under complex operating conditions in advance to initialize the control strategy, while simultaneously leveraging the predictive mechanism of model predictive control to precisely adjust the speed response during actual operation. Furthermore, the online update capability of reinforcement learning helps to continuously optimize control parameters, further improving the system's robustness and avoiding getting trapped in local optima. This fusion is expected to demonstrate superior response speed, stability, and overall control performance in high-speed train speed tracking control.

[0084] In some optional embodiments of the present invention, the aforementioned basic dataset is subsequently used as input to a neural network model for corresponding processing and calculation. The neural network model is sensitive to the scale of the input data; therefore, the data can be preprocessed. This preprocessing process may include cleaning, identifying which features or samples in the dataset have missing values, and, based on the proportion of missing data and the business context, selecting to fill in or directly delete samples or features containing a large number of missing values; and using statistical methods to identify outliers and remove abnormal data. Then, baseline calibration can be performed on the basic dataset of the target train to obtain calibrated data. For example, a Butterworth filter can be used to filter the data in the basic dataset, and its transfer function formula is:

[0085] ,

[0086] in, is the signal frequency, is the cutoff frequency, is the filter order.

[0087] The continuous-time Butterworth filter is converted into a difference equation using the bilinear transformation method: , where: is the filter coefficient calculated by the Butterworth filter; is the original speed data; is the filtered data; t represents time, and is a positive integer.

[0088] Then, the calibrated data is normalized to obtain the standardized basic data set. Exemplarily, the Min-Max normalization method can be used for data standardization, and the formula is as follows:

[0089] ,

[0090] where, is the original data; and is the minimum and maximum value of the feature in the basic data set, is the normalized data, ranging from [0, 1].

[0091] In some optional embodiments of the embodiment of the application, the train data-driven model is a deep learning method based on a gated recurrent unit (GRU) for speed prediction of the target train. The train data-driven model mainly includes: a CNN-based data feature learning module, a GRU-based encoder module, an attention mechanism module, and a GRU-based decoder module. The data feature learning module is configured to extract features from the original data in the basic data set to obtain a feature array. In order to make the model better learn the influence degree of each feature on the model output, the multi-layer structure and local receptive field advantage of CNN are used to automatically extract, learn and abstract effective features from the original data.

[0092] The core calculation process of the CNN-based data feature learning module is represented by the following formula:

[0093] ,

[0094] where, is the input matrix; is a set of weights to be learned; represents the weighted sum result in the local area; is the index inside the convolution kernel; is the size of the convolution kernel; is an activation function, using ReLU; denotes each value of the activation function output within the pooling window; is the output value of the region after pooling; is the flattened feature vector; is the weight matrix of the fully connected layer; is the bias vector; is an activation function, using linear activation; is the predicted output of the model.

[0095] The encoder module is configured to output a hidden state vector of the current time step based on the feature array. In this embodiment, the encoder module can be formed by a series of GRU units connected in sequence, assuming that the features of the first n time steps are used as model inputs in total, so a layer of n GRU units should be constructed to encode the input information. The input feature array, i.e. , will be composed of train operation data of time . Each GRU unit receives a feature vector and the previous hidden state, and calculates through the formula, outputting a series of hidden states, i.e. .

[0096] The core calculation process of the GRU-based encoder module is represented by the following formula:

[0097] ,

[0098] wherein, is the output of the reset gate, is the weight matrix of the reset gate, is the hidden state of the previous time step, is the input feature of the current time step, is a sigmoid activation function, defined as , which maps any real number to the range of (0, 1), is the output of the update gate, is the weight matrix of the update gate, is the hidden state of the current time step, is the candidate hidden state, is the weight matrix of the candidate hidden state, is a hyperbolic tangent activation function, and the symbol * represents element-wise multiplication of matrices.

[0099] The attention mechanism module is configured to obtain attention values based on the hidden state vector and the last hidden state of the decoder through the attention mechanism. In this embodiment, the attention unit is repeatedly used The length of the train speed sequence to be predicted is used to construct the attention mechanism module. The input of the attention unit consists of two parts: one part is the hidden state output by the encoder. This is called the "key"; the other part is the decoder's last hidden state. This is called a "query". The output is an attention value. The core calculation process of the attention module is represented by the following formula:

[0100] ,

[0101] in, It is a query vector; It is the first One key vector; , These are the matrices and vectors that need to be trained; yes Compared to query Normalized weights; It's the attention value.

[0102] The decoder module is configured to output a train speed prediction sequence for multiple time steps based on the attention value and the prediction value from the previous time step. In this embodiment, the decoder module also uses a GRU as the basic unit of model output, but there are two main differences compared to the encoder module: in the decoder module, the input of each GRU unit includes not only the current attention value but also the prediction value from the previous time step; and a fully connected layer is added after the GRU layer to reduce the hidden state. The dimension is then determined. Finally, the Corrected Linear Unit (ReLU) is used as the activation function. Similar to the attention unit, the basic unit of the decoder module will be reused N times to generate the predicted train speed for the next N time steps.

[0103] Through the above process, the data-driven model based on GRU deep learning serves as the prediction model for the model prediction controller. It is used to predict the system state at the next moment.

[0104] Please refer to Figure 5 , Figure 5 A flowchart of a train speed tracking control method provided in this disclosure embodiment, namely for... Figure 2 Step 203 in process 200 provided a specific implementation. Other steps in process 200 are not adjusted; a new complete embodiment is obtained by replacing step 203 with the specific implementation provided in this embodiment. Process 400 includes the following steps:

[0105] Step 401: Extract reference velocity information from the basic dataset.

[0106] Step 402: Construct a target function based on the reference speed information and the train speed prediction sequence.

[0107] In the prediction time domain , the constructed target function makes the train speed track the reference trajectory as much as possible while penalizing the changes in control input. The target function is set as:

[0108]

[0109] where denotes the current sampling time; , denotes the actual speed and the reference speed; denotes the control input; is a weight matrix.

[0110] Step 403: Based on the preset constraints and the target function, use the gradient optimization algorithm to obtain the optimal control quantity sequence with the train real-time speed at the current time as the initial condition.

[0111] To ensure the safety and efficiency of the train, the following constraints must be met. Due to the physical characteristics of the traction motor, the traction force and the braking force are bounded, so the control input satisfies where is the maximum traction force, is the maximum braking force; the maximum speed of the train is first limited by the maximum speed allowed by the train itself; secondly, as the line conditions and operating conditions change, the maximum allowable speed of the train may also change, so the train speed needs to satisfy ; at the same time, to ensure the safety of train operation, the coupling force between trains should be kept within the interval , i.e. , the coupling force is determined by the train formation type; to ensure the comfort of passengers, the acceleration change of the train in the traction and braking phases should also be limited, i.e. where is given by the ATP (Automatic Train Protection).

[0112] In summary, the preset constraints can be expressed as .

[0113] Then, the real-time speed of the target train at the current time is obtained, and the reference trajectory is also obtained. With the current time as the initial condition, the following optimization problem is solved:

[0114] ,​

[0115] ,

[0116] by automatic differentiation of the gradient , update using gradient descent , repeat iteration until convergence. In practical applications, only the first control quantity in the resulting optimal control sequence is input to the target train to achieve tracking control of the target train. Further, at the next time , the above process can also be performed to re-solve, forming a rolling horizon control.

[0117] Further referring to Figure 6 , as an implementation of the method shown in the above figures, the present disclosure provides an embodiment of a train speed tracking control device, which corresponds to the method embodiment shown in Figure 2 , and the device can be specifically applied to various electronic devices.

[0118] As shown in Figure 6 , the train speed tracking control device 500 of the present embodiment can include a data acquisition module 501, a train data-driven model construction module 502, a speed prediction module 503, a control quantity determination module 504, a control quantity sequence correction module 505, and a tracking control module 506. The data acquisition module 501 is configured to acquire a basic data set of a target train, the basic data set being used to represent the nonlinear dynamics information of the target train; the train data-driven model construction module 502 is configured to use the basic data set to generate a train data-driven model based on an improved GRU network training; the speed prediction module 503 is configured to use the train data-driven model as a model predictive controller to obtain a train speed prediction sequence at multiple time steps based on the basic data set; the train data-driven model is generated based on the basic data set; the control quantity determination module 504 is configured to establish a train state space model based on a dynamics equation, as a controlled object of model predictive control and an environment of reinforcement learning, and use the model predictive controller to output an optimal control quantity sequence based on the train speed prediction sequence; the control quantity sequence correction module 505 is configured to correct the optimal control quantity sequence using a reinforcement learning algorithm to obtain a corrected optimal control quantity sequence; and the tracking control module 506 is configured to perform tracking control on the target train based on the corrected optimal control quantity sequence.

[0119] In the train speed tracking control device 500 in this embodiment, the specific processing of the data acquisition module 501, the train data driven model construction module 502, the speed prediction module 503, the control quantity determination module 504, the control quantity sequence correction module 505, and the tracking control module 506 and the technical effects brought by the same can be referred to the corresponding descriptions of steps 201-206 in the method embodiment. Figure 2 The relevant descriptions of steps 201-206 in the corresponding embodiment are not repeated here.

[0120] This embodiment exists as a device embodiment corresponding to the above-mentioned method embodiment. The train speed tracking control device provided in this embodiment trains a train data driven model based on the data of the train running process, which can more accurately reflect the complex characteristics of the actual system. The train data driven model is used to predict the running speed of the train at a future time, and a model predictive controller is used to track and control the running of the train at the future time in combination with the predicted running speed of the train. The future state is taken into account through multi-time step prediction, which can effectively cope with delays or disturbances and improve the stability and accuracy of train control.

[0121] According to embodiments of the present disclosure, the present disclosure further provides an electronic device, comprising: at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to implement the train speed tracking control method described in any of the above embodiments when executed.

[0122] According to embodiments of the present disclosure, the present disclosure further provides a readable storage medium, which stores computer instructions for enabling a computer to implement the train speed tracking control method described in any of the above embodiments when executed.

[0123] According to embodiments of the present disclosure, the present disclosure further provides a computer program product, which enables the train speed tracking control method described in any of the above embodiments when executed by a processor.

[0124] Figure 7 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present disclosure described and / or claimed in this document.

[0125] like Figure 7 As shown, device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 602 or a computer program loaded into random access memory (RAM) 603 from storage unit 608. RAM 603 may also store various programs and data required for the operation of device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.

[0126] Multiple components in device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of monitors, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0127] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as the train speed tracking control method. For example, in some embodiments, the train speed tracking control method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the train speed tracking control method described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform the train speed tracking control method by any other suitable means (e.g., by means of firmware).

[0128] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a load programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0129] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or the block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, or entirely on a remote machine or server.

[0130] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical conductors, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0131] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0132] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0133] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. A server can be a cloud server, also known as cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of large management difficulty and weak business scalability in traditional physical host and virtual private server (VPS, Virtual Private Server) services.

[0134] According to the technical scheme of the embodiment of the application, the beneficial effects are repeated.

[0135] It should be understood that various forms of flow shown above can be used to reorder, add, or delete steps. For example, the steps described in the present disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical scheme of the present disclosure can be achieved, and the present disclosure is not limited herein.

[0136] The above detailed description does not limit the scope of the disclosure. Various modifications, combinations, sub-combinations and alternatives can be made to the detailed description. Any modification, equivalent replacement and improvement etc. made within the spirit and principle of the disclosure shall be included in the scope of the disclosure.

Claims

1. A train speed tracking control method, characterized in that, include: Obtain a basic dataset of the target train, which is used to characterize the nonlinear dynamic information of the target train; Train data-driven models are generated by training an improved GRU network using the aforementioned basic dataset. Using the train data-driven model as the prediction model for the model prediction controller, a train speed prediction sequence for multiple time steps is obtained based on the basic dataset. A train state-space model is established based on the dynamic equations, which serves as the controlled object of model predictive control and the environment for reinforcement learning. The model predictive controller outputs the optimal control quantity sequence based on the train speed prediction sequence. The optimal control quantity sequence is corrected using a reinforcement learning algorithm to obtain the corrected optimal control quantity sequence. The target train is tracked and controlled based on the corrected optimal control sequence.

2. The method according to claim 1, characterized in that, The step of using a reinforcement learning algorithm to correct the optimal control input sequence includes: The target train is described using a single-mass model, and the force analysis of the target train is performed to construct the dynamic equations of the target train. A state-space model of the target train is generated based on the dynamic equations. Using the actor-commentator algorithm, the expected values ​​of action quantity and cumulative reward are obtained based on the state space model. The optimal control input sequence is corrected based on the action quantity to obtain the actual control input sequence; The actual control input sequence of the action quantity is adjusted based on the expected value to obtain the corrected optimal control quantity sequence.

3. The method according to claim 2, characterized in that, The method of using the actor-commentator algorithm to obtain the expected values ​​of action quantity and cumulative reward based on the state space model includes: The state variables of the state space model are input into the actor network to obtain the motion variables; the state variables include velocity errors. The reward function is determined based on the optimal control sequence and velocity error. Using a commentator network, the expected value of the cumulative reward is obtained based on the state-action pairs formed by the state and action quantities and the reward function.

4. The method according to claim 1, characterized in that, Before establishing the train data-driven model based on the basic dataset, the following is also included: Baseline calibration is performed on the data in the basic dataset to obtain calibrated data; The calibrated data is then normalized to obtain a standardized basic dataset.

5. The method according to claim 1, characterized in that, The train data-driven model includes: The data feature learning module is configured to extract features from the raw data in the basic dataset to obtain a feature array; The encoder module is configured to output the hidden state vector of the current time step based on the feature array; The attention mechanism module is configured to obtain an attention value based on the hidden state vector and the last hidden state of the decoder through an attention mechanism; The decoder module is configured to output a train speed prediction sequence for multiple time steps based on the attention value and the prediction value of the previous time step.

6. The method according to claim 1, characterized in that, The prediction model using the train data-driven model as the model prediction controller obtains a train speed prediction sequence for multiple time steps based on the basic dataset, including: Extract real-time velocity information, position information, and acceleration information from the basic dataset; The current state quantity is obtained based on the real-time velocity information, position information, acceleration information, and control quantity from the previous moment. The current state quantity is input into the train data-driven model to obtain the train speed prediction sequence for the multiple time steps.

7. The method according to claim 1, characterized in that, The step of using the model predictive controller to output the optimal control quantity sequence based on the train speed prediction sequence includes: Extract the reference velocity information from the basic dataset; A target function is constructed based on the reference speed information and the train speed prediction sequence. Based on the preset constraints and the objective function, and using the current real-time train speed as the initial condition, the optimal control quantity sequence is obtained using a gradient optimization algorithm.

8. The method according to claim 1, characterized in that, The tracking control of the target train based on the optimal control sequence includes: The first control quantity in the optimal control quantity sequence is sent to the target train to perform tracking control on the target train.

9. A train speed tracking control device, comprising: The data acquisition module is configured to acquire a basic dataset of the target train, which is used to characterize the nonlinear dynamic information of the target train. The train data-driven model building module is configured to use the base dataset to train and generate a train data-driven model based on an improved GRU network. The speed prediction module is configured to use the train data-driven model as the prediction model of the model prediction controller to obtain a train speed prediction sequence for multiple time steps based on the basic dataset. The control quantity determination module is configured to establish a train state-space model based on the dynamic equations, which serves as the controlled object of model predictive control and the environment for reinforcement learning. The model predictive controller is used to output the optimal control quantity sequence based on the train speed prediction sequence. The control quantity sequence correction module is configured to use a reinforcement learning algorithm to correct the optimal control quantity sequence to obtain the corrected optimal control quantity sequence. The tracking control module is configured to track and control the target train based on the modified optimal control sequence.

10. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the train speed tracking control method according to any one of claims 1-8.