Formation cooperative tracking control method, device, equipment, medium and product

Through deep neural network and reinforcement learning, the optimization of train spacing and speed control is solved, and the problem of coordinated tracking of virtual marshalling trains under uncertain interference is achieved, achieving higher accuracy and stability to ensure the safe operation of the train.

CN120397038AActive Publication Date: 2025-08-01EAST CHINA JIAOTONG UNIVERSITY

Patent Information

Application Number
CN202510920529.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-08-01
Estimated Expiration
2045-07-04

AI Technical Summary

Technical Problem

In virtual marshalling trains, the operation of the train is affected by a variety of uncertain interference factors, causing the operating state to deviate from the expected trajectory, affecting the coordinated tracking performance of the marshalling trains. How to improve accuracy and stability is an important issue.

Method used

The deep neural network dynamic model and reinforcement learning method are adopted, combined with the elastic interval evaluation model and the cross-entropy algorithm, and the elastic optimal control objective function of the virtual marshalling train is determined, and the train spacing and speed control are optimized through data-driven methods to ensure the safe and coordinated operation of the train.

Benefits of technology

The accuracy and stability of the formation collaborative tracking control of virtual marshalling trains is improved, prevents collisions between adjacent trains and ensures the safe operation of the train.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120397038A_ABST
    Figure CN120397038A_ABST
Patent Text Reader

Abstract

The invention discloses a formation cooperative tracking control method and device, equipment, a medium and a product, and relates to the field of tracking control. The method comprises the following steps: acquiring operation data of a virtual marshalling train; inputting the operation data into the deep neural network dynamic model, and updating dynamic model parameters; determining an elastic optimal control objective function of the virtual marshalling train based on an elastic interval evaluation model according to the train operation data and the updated dynamic model by adopting a reinforcement learning method; the elastic interval evaluation model is a mathematical model for calculating a safe following distance based on a relative braking distance between two adjacent trains so as to set interval following distance cooperative operation; and solving the elastic optimal control objective function by adopting a cross entropy algorithm to obtain a control sequence so as to perform formation cooperative tracking control on the virtual marshalling trains. The invention aims to improve the accuracy and stability of formation cooperative tracking control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of tracking control, and in particular, to a formation cooperative tracking control method, device, equipment, medium and product. Background Art

[0002] Virtual formation trains are a new type of railway transportation organization method. By using communication and control technologies, multiple trains are virtually combined to run together to improve transportation efficiency, reduce energy consumption and operating costs. However, in actual operation, trains are affected by various uncertain interference factors, such as track irregularities, wind force, vehicle failures, communication delays, etc. These interference factors may cause the operating state of the train to deviate from the expected trajectory, thus affecting the cooperative tracking performance of the formation trains.

[0003] Due to the existence of various interference factors during the train operation process, it is crucial to improve the accuracy and stability for the problem of cooperative tracking control of virtual formation trains considering uncertain interference factors. Summary of the Invention

[0004] The purpose of the present application is to provide a formation cooperative tracking control method, device, equipment, medium and product, which can improve the accuracy and stability of formation cooperative tracking control.

[0005] To achieve the above purpose, the present application provides the following solutions: In the first aspect, the present application provides a formation cooperative tracking control method, including: Obtain the operation data of the virtual formation train; the virtual formation train is a coupled train cluster composed of multiple single-mass trains; Input the operation data into the deep neural network dynamics model to obtain output state data; the deep neural network dynamics model is obtained by training the deep neural network using historical operation data and corresponding output state data as an offline data set; Adopt the method of reinforcement learning to determine the elastic optimal control objective function of the virtual formation train based on the elastic interval evaluation model according to the output state data; the elastic interval evaluation model is a mathematical model that calculates the safe following distance based on the relative braking distance between two adjacent trains and runs in a set interval following distance. Solve the elastic optimal control objective function by using the cross-entropy algorithm to obtain a control sequence; the control sequence is used for formation cooperative tracking control of the virtual formation train.

[0006] In the second aspect, the present application provides a formation cooperative tracking control device, including: A data acquisition module, configured to acquire the operation data of a virtual formation train; the virtual formation train is a coupled train cluster composed of multiple single-mass trains; An output module, configured to input the operation data into a deep neural network dynamics model to obtain output state data; the deep neural network dynamics model is obtained by training a deep neural network using historical operation data and corresponding output state data as an offline data set; A target function determination module, configured to use the method of reinforcement learning to determine an elastic optimal control target function of the virtual formation train based on the elastic interval evaluation model according to the output state data; the elastic interval evaluation model is a mathematical model that calculates a safe following distance based on the relative braking distance between two adjacent trains and runs in coordination with a set interval following distance; A solution module, configured to solve the elastic optimal control target function using a cross-entropy algorithm to obtain a control sequence; the control sequence is used for formation cooperative tracking control of the virtual formation train.

[0007] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the above-mentioned formation cooperative tracking control method.

[0008] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned formation cooperative tracking control method is implemented.

[0009] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the above-mentioned formation cooperative tracking control method is implemented.

[0010] According to the specific embodiments provided by the present application, the following technical effects are disclosed in the present application: The present application provides a formation cooperative tracking control method, device, equipment, medium and product. By acquiring the operation data of virtual formation trains; inputting the operation data into a deep neural network dynamics model to obtain output state data; adopting a reinforcement learning method, based on the output state data, an elastic optimal control objective function of the virtual formation trains is determined based on an elastic interval evaluation model; through the deep neural network dynamics model and the reinforcement learning method, it can ensure that the virtual formation trains reach the desired formation operation state and prevent collisions between adjacent trains to ensure the safe operation of the trains. Since the elastic interval evaluation model calculates the safe following distance based on the relative braking distance between two adjacent trains to set a mathematical model for following at an interval distance; the elastic interval evaluation model can ensure the dynamic tracking distance adjustment between adjacent trains. Due to the non-linear characteristics of train dynamics and the reward function, the accuracy is not high. The present application uses the cross-entropy algorithm to solve the elastic optimal control objective function, improve the accuracy of formation cooperative tracking control, obtain a control sequence, and thus perform formation cooperative tracking control on the virtual formation trains. Therefore, the present application can improve the accuracy and stability of formation cooperative tracking control. Description of the Drawings

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0012] Figure 1 It is a flowchart of the formation cooperative tracking control method; Figure 2 It is a schematic diagram of the virtual formation train tracking interval mechanism based on the relative braking distance; Figure 3 It is a schematic diagram of the dynamic formation process; [[ID=!7]] Figure 4 It is a schematic diagram of the cooperative operation process; Figure 5 It is a schematic diagram of the dynamic decoupling process; Figure 6 It is a block diagram of the deep reinforcement learning data-driven model predictive control; Figure 7 It is a block diagram of the DNN; Figure 8 It is a schematic diagram of the RL control process; Figure 9 It is a curve of the train historical operation data; Figure 10 It is a schematic diagram of the speed prediction result; Figure 11 Schematic diagram of speed prediction error Figure 12 Schematic diagram of displacement prediction result Figure 13 Schematic diagram of displacement prediction error Figure 14 Speed curve of leader-follower trains under Scenario 1 Figure 15 Distance curve of leader-follower trains under Scenario 1 Figure 16 State tracking error curve of leader-follower trains under Scenario 1 <000007 Figure 17 Acceleration curve of leader-follower trains under Scenario 1 Figure 18 Speed curve of leader-follower trains under Scenario 2 Figure 19 Distance curve of leader-follower trains under Scenario 2 Figure 20 State tracking error curve of leader-follower trains under Scenario 2 Figure 21 Acceleration curve of leader-follower trains under Scenario 2 Detailed implementation manners

[0013] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0014] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0015] In an exemplary embodiment, as Figure 1 shown, a formation cooperative tracking control method is provided, and the formation cooperative tracking control method includes: <000 Step 100: Obtain the operation data of the virtual formation train. The virtual formation train is a coupled train cluster composed of multiple single-mass trains.

[0016] Step 200: Input the operation data into the deep neural network dynamics model to obtain the output state data. The deep neural network dynamics model is trained by using the historical operation data and the corresponding output state data as the offline data set.

[0017] Step 300: Using the method of reinforcement learning, based on the output state data, determine the elastic optimal control objective function of the virtual formation train according to the elastic interval evaluation model. The elastic interval evaluation model is a mathematical model that calculates the safe following distance based on the relative braking distance between two adjacent trains and runs in coordination with the set interval following distance.

[0018] The elastic optimal control objective function specifically includes: .

[0019] Among them, is the elastic optimal control objective function; is the prediction horizon; is the sampling time; is any time within the prediction horizon; is the reward function of train at time is the operating state of train at time is the control input of train at time

[0020] Step 400: Use the cross-entropy algorithm to solve the elastic optimal control objective function to obtain the control sequence. The control sequence is used for formation cooperative tracking control of the virtual formation train.

[0021] Among them, the method for determining the deep neural network dynamics model specifically includes: Obtain the offline data set; perform normalization processing on the offline data set to obtain the processed data set; construct a deep neural network; input the processed data set into the deep neural network, and use the ADAM algorithm to train the parameters of the deep neural network with the goal of minimizing the loss function to obtain the trained deep neural network; the parameters include: weights.

[0022] Determine the trained deep neural network as the deep neural network dynamics model.

[0023] In one embodiment, use the min-max normalization method to perform normalization processing on the offline data set to obtain the processed data set; the expression of the processed data set is: .

[0024] Among them, is the processed data set; is the offline data set; is the minimum value in the offline data set; and are both the sample numbers in the offline dataset; is the total number of samples in the offline dataset; is the maximum value in the offline dataset.

[0025] The expression of the loss function is: .

[0026] Among them, is the loss function; is the sampling time; is the sampling moment; is the train at the operating state at the sampling moment; is the train at the operating state at the sampling moment; is the train at the control input at the sampling moment; is the learning model; , are both the weight coefficients of the deep neural network.

[0027] In one embodiment, the cross-entropy algorithm is used to solve the elastic optimal control objective function to obtain a control sequence, which specifically includes: Determine the initial parameters; the initial parameters include: setting the iteration time, ratio, and N train samples; the train samples are collected based on the probability distribution function through parametric distribution.

[0028] Determine the fitness function according to the elastic optimal control objective function.

[0029] Based on the parameter combination at the current iteration time, determine the fitness function value according to the fitness function; the parameter combination includes: ratio and train samples.

[0030] Judge whether the iteration stop condition is satisfied; the iteration stop condition is that the current iteration time reaches the set iteration time or the fitness function value at the current iteration time is within the set elite threshold interval.

[0031] If not, then determine the elite sample set based on the fitness function value, update the parameters of the probability distribution function based on the set learning step size parameter, and return "Based on the parameter combination at the current iteration time, determine the fitness function value according to the fitness function".

[0032] If so, determine the control sequence according to the parameter combination at the current iteration time.

[0033] In practical applications, the operation steps corresponding to the solution proposed in this application are as follows: First, an elastic interval evaluation model for virtual formation trains is established. Through this elastic interval evaluation model, it is determined whether adjacent trains are in a coupled operation state, thereby enabling online adjustment of the train operation state.

[0034] Although the control method designed in this application does not require a mathematical model of virtual formation trains, for a better understanding of its operating principle, the dynamic mechanism of trains is briefly introduced here. A virtual formation train is described as a coupled train cluster composed of multiple single-particle trains, that is, the forces between carriages in a unit train are ignored. According to Newton's second law, the dynamic model of a unit train can be described as: (1) where and are the position and velocity of the train at the current moment, respectively; and are the mass of the train and the coefficient of gyratory mass, respectively; is the first derivative of ; is the first derivative of . is the serial number.

[0035] is the self-input control force (including traction force and braking force) received by the train at ; is the basic resistance received by the train, mainly generated by the running resistance of internal components of the train, air resistance, friction between wheels and rails, etc. Its characteristics are related to the train speed, and can be approximately expressed as: (2) where , , are all basic running resistance coefficients, which can be obtained based on past experience; is the additional resistance received by the train, and its characteristics are closely related to the environment where the train operates on the line. Usually, it includes three parts: additional resistance on slopes, additional resistance on curves, and additional resistance in tunnels. Therefore, can be further expressed as: (3) where is the gravitational acceleration of the train; , , They are the track gradient, the track curve radius, and the tunnel length, respectively. is an empirical constant.

[0036] Elastic interval evaluation model: In the virtual formation operation mode, each unit train runs in coordination with a small following distance. The optimal safe following distance is solved by calculating the relative braking distance between two adjacent trains in real time. As Figure 2 shown, the actual running interval of the train and the minimum safe interval are calculated as follows: (4) where is the body length of the leading train; is the normal braking distance of the trailing train after time, expressed as ; is the emergency braking distance of the leading train at time, usually expressed as ; is the distance traveled by the train during the reaction time at time; is the displacement of the train ; is the normal braking control acceleration of the train ; is the normal braking control acceleration of the train ; at time.

[0037] Furthermore, the optimal tracking interval is designed as: (5) where is the optimal interval adjustment coefficient. Then the elasticity between adjacent trains is designed as: (6) When the elasticity of the train is small, the actual interval between the two trains is large, and the trains may be in a non-coupled state; when the elasticity of the train approaches "1", the actual interval between the two trains gradually approaches the optimal interval, and only at this time does the train meet one of the conditions for entering the coupled state; when the elasticity of the train is large, the actual interval between the two trains is small, and the trains are in a dangerous state, which may lead to safety accidents such as collisions.

[0038] "Three-stage" formation process: Virtual formation consists of three stages, namely dynamic formation, coordinated operation, and dynamic dissolution. To ensure the punctuality and safety of operation, both the dynamic formation and dynamic dissolution stages need to be completed within the specified section of the route.

[0039] Dynamic formation means that the train transitions from the independent operation state to the coupled operation state. Taking the formation of two trains as an example, as Figure 3 shown, where "LT" represents the leader train (the front train); "FT" represents the follower train (the rear train); represents the desired spacing between two trains in the virtual formation mode; represents the desired spacing in the independent operation stage. During the dynamic formation process, the train needs to reduce the tracking spacing with the front convoy to achieve vehicle-to-vehicle communication. When the communication range is satisfied, it needs to maintain the same speed as the front convoy. There are two dynamic formation methods as follows: In the first case, assume that initially both trains are operating independently on the line at the maximum operating speed . When the rear train completely enters the formation section, the rear train will send a formation request signal to the front train. After the front train receives the signal, at this time, the front train will reduce its speed to a certain value so that the rear train can catch up with the front train. When the distance between the two trains is about to reach the desired spacing in the coordinated operation stage, the front train starts to accelerate to and maintain the same speed as the rear train. In the second case, assume that initially both trains are operating independently on the line at the operating speed . When the rear train completely enters the formation section, the rear train will also send a formation request signal to the front train to prepare the front train for formation. At this time, the rear train will increase its speed to to catch up with the front train. When the distance between the two trains is about to reach the desired spacing in the coordinated operation stage, the rear train starts to decelerate to and maintain the same speed as the front train. After the dynamic formation is completed, the train group immediately enters the coordinated operation stage. In this stage, the train control system needs to achieve the speed coordination of adjacent trains and ensure the goal of the desired operating spacing. This process is as Figure 4 shown. When the virtual formation train enters the dynamic dissolution section, it is necessary to increase the operating spacing between the dissolution train and the preceding train until the spacing required for the independent operation state mode is reached, as Figure 5 shown. This process is actually an inverse process of dynamic formation. Its dissolution process is similar to the dynamic formation process and will not be elaborated here.

[0040] Secondly, a more accurate data-driven deep neural network (DNN) was constructed. By using the recorded train operation input and output data, a train dynamics model was trained, and a reinforcement learning distributed predictive control method was designed. This control method introduced the idea of reinforcement learning on the basis of distributed model predictive control, combined with the elastic interval evaluation model to design the elastic optimal control objective function for the virtual formation train, constructed the train optimal control problem according to the trained DNN and the objective function, and used the Cross-Entropy Method (CEM) to solve this optimal control problem.

[0041] As Figure 6 shown is the block diagram of the deep reinforcement learning control based on the data-driven model designed in this application. The system offline learns the train model through the collected historical data and, during the actual operation process, online updates the train model through the collected real-time data

[0042] Deep neural network dynamics model.

[0043] For the virtual formation train, its unit train state vector can be defined as , and the control input is defined as . Then the train dynamics model is expressed as follows: (9) where is the derivative of the train i state quantity; is a function related to , specifically expressed as .

[0044] Since the data-driven model designed in this application is trained according to the discrete sampling data, therefore, the discretized train dynamics model is expressed as: (10) where is the running state of the train at the sampling moment; is the running state of the train at the sampling moment; is the train At the control input at the sampling moment; is the state transition function; is the sampling time interval.

[0045] According to the current operating state information of the high - speed train , a deep neural network is used to learn the state transition function of the train, and the learning model can be expressed as , where are the weight coefficients of the DNN model, and the model output is the predicted state deviation at the next moment . Therefore, the prediction model can be expressed as: (11) DNN usually consists of three parts: the input layer, the hidden layer, and the output layer. As Figure 7 shown, DNN forwards the information received by the input layer together with the weight coefficients , , and then combines with the activation function , distributes the obtained results to each neuron in the hidden layer, and forwards them sequentially until the output layer. This network uses the historical operation data of the train to approximate the real model of the train, so as to achieve an accurate prediction of the train's operating state at the next moment.

[0046] The sampled historical operation data of the train is used as the offline data set for training DNN, which is expressed as: (12) where is the sampling data set of the train ; is the initial state data set of the train , including the initial speed information and initial position information of the train at each sampling moment, that is, the state information of the train at the previous moment; is the control input data set of the train , including the train control input information at each sampling moment; is the output state data set of the train , including the final speed information and final position information of the train at each sampling moment, that is, the actual state information of the train at the sampling moment. To ensure the accuracy of the model, during the train operation, a new data set is sampled and collected in real - time to online - train DNN. Therefore, the training data set of DNN is composed of two parts: the offline data set and the online data set .

[0047] As described above, the input of DNN includes speed (Unit: ), position (Unit: ) and control acceleration (Unit: ). Since these three types of input data are in different scale ranges, this may lead to low efficiency in training the DNN model. Therefore, to accelerate the model training speed, it is necessary to normalize the input data. This application selects the min-max normalization method, and its mathematical expression is described as follows: (13) Input the normalized data into the DNN for training. During the DNN training process, this application uses the ADAM algorithm to minimize the loss function: (14) Therefore, the training objective of the DNN model is to find the optimal weight coefficients to minimize the loss function .

[0048] Reinforcement learning distributed predictive control.

[0049] To achieve the cooperative formation control of virtual formation trains, this application proposes a distributed model predictive control algorithm based on reinforcement learning. This method can ensure that the virtual formation trains reach the desired formation operation state and prevent collisions between adjacent trains to ensure the safe operation of the trains.

[0050] The control objective of this application is to achieve the cooperative control of dynamic interval adjustment of virtual formation trains, that is, the controller not only needs to make the operation states of each unit train converge and stabilize within a certain time, but also for problems such as uncertain disturbances existing during the operation process, it can ensure that the dynamic tracking distance between adjacent trains is adjusted to avoid safety accidents. At the same time, the acceleration output by the controller should meet the requirements of the high-speed train operation control system. Therefore, the control objective can be specifically described as: (15) where, is the maximum braking deceleration of the train; is the acceleration of the train at the current moment; is the maximum traction acceleration of the train. is the speed of the train at the current moment.

[0051] Reinforcement learning (RL) is a machine learning method for sequential decision-making. As Figure 8As shown, during the reinforcement learning control process, the agent and the environment are always in an interactive state. The agent obtains state information from the environment and uses this state information to apply an action to the environment. Then, the environment outputs the next state and the corresponding action reward obtained according to this applied action. The agent, or the decision maker, is usually used to explore appropriate actions to maximize the total reward obtained from the environment. In this application, the controller represents the agent and the train represents the environment. During the actual operation control process, assume is the state vector of the train at the current moment, is the reward value of the train at the current moment, and the controller obtains a control action to act on the train to obtain the state vector of the train at the next moment and the reward function , and repeat this process.

[0052] The reward function constructed in this application includes two parts: collaborative reward and collision avoidance penalty.

[0053] First, according to formula (15), to maintain a stable formation of the train formation, three reward functions need to be set, including consistent speed reward, elastic interval reward, and control variable reward, which are specifically described as follows: (16) (17) (18) Among them, , , are the consistent speed reward, elastic interval reward, and control variable reward respectively; is the speed deviation of the trailing train from the leading train ; is the deviation of the elasticity of the train from the optimal elasticity; , , , are the weight coefficients of the corresponding reward functions respectively. By designing the magnitudes of the four, the importance degrees of speed, spacing, control amount, and control change amount are reflected. is the elasticity of the train i at t moment.

[0054] Secondly, to avoid collisions between adjacent trains, a collision avoidance penalty function needs to be set, which is specifically described as follows: (19) Among them, is the optimal spacing adjustment coefficient; is the weight coefficient of the collision avoidance penalty function.

[0055] Therefore, after sorting, the final reward function of the virtual formation train is designed as: (20) By combining the learned model, that is, formula (11) and the final reward function, that is, formula (20), a cooperative tracking controller based on a data-driven model is proposed to achieve the cooperative control task of the virtual formation train. Specifically, it is to solve the finite prediction time domain according to formula (11) and each operation constraint condition within the control sequence , and then use the learned neural network dynamics model to predict the future state. This problem is specifically described in the following form: (21) (21a) (21b) (21c) (21d) (21e) (21f) (21g) is the parameter and terminal constraint set.

[0056] Due to the non-linear characteristics of the train dynamics and the reward function, it is difficult to accurately calculate the above equations, and approximate solutions can be obtained by using other optimization algorithms. This application will use the cross-entropy algorithm (Cross Entropy Method, CEM) to solve this non-linear problem. Such optimization algorithms can randomly generate a better predictive control sequence within the prediction time domain , use the learned neural network model to predict the corresponding state sequence. Calculate the reward values obtained for all state sequences and select the action sequence corresponding to the highest expected cumulative reward. In this operation process, the basic idea of model predictive control is realized, that is, the algorithm only executes the first action at each time step and receives updated state information in a timely manner. If the next time step is reached, the algorithm recalculates the best action sequence and repeats the above process.

[0057] CEM is an evolutionary strategy method that can perform random sampling from a parameterized probability distribution and update the parameters in the distribution based on the results of each sampling, so that the probability of action sequences with higher cumulative rewards in the distribution is relatively high.

[0058] This method aims to find the most (sub)optimal solution for the cumulative reward function, which is specifically described as follows: (22) Among them, is the selected most (sub)optimal solution vector; is the fitness function. In this application, the fitness function selects the cumulative reward function shown in formula (21), and the most (sub)optimal solution vector corresponds to the control sequence . is an arbitrary solution set.

[0059] Suppose is a sample collected from the parameterized distribution , where , and are independent of each other; is the total number of samples collected. Then there exists such that there is a set of high-value samples , is an arbitrary real number. The specific description is as follows: (23) Among them, is the th collected sample; is the corresponding fitness function.

[0060] Based on this set of high-value samples, a new probability distribution function can be learned. However, if is too large, will only contain a few samples, or even none, which makes learning impossible. And if is set too small, it will cause the optimization process to converge slowly to the most (sub)optimal solution. Therefore, to ensure the algorithm's solution efficiency, a ratio is selected, and is adjusted to the set of the best samples. This corresponds to setting , provided that the samples are sorted in descending order of their values. The best samples are called elite samples. In the actual solution process, usually selects a certain value within the range of [0.01, 0.1]. Therefore, the probability distribution parameters and The calculation formula can be described as follows: (24) are all sets of probability distribution parameters of the elite sample set; 、 、 are all subsets thereof; is the i th solution of the j th collected sample. i is the serial number.

[0061] (25) To ensure the accuracy of the probability distribution function for learning, a learning step parameter is introduced, and the probability distribution parameters and for learning are expressed as: (26) (27) are all sets of probability distribution parameters for learning; 、 、 are all subsets thereof; is the j th subset of the corresponding set of probability distribution parameters.

[0062] For clarity, the specific algorithm steps are summarized as follows: Parameters: , ,iteration time ; Initialize the probability distribution parameters and : , ; Steps: For do; For do; Collect samples from the parametric distribution ; Evaluate the fitness function of each sample ; Sort the fitness function values of each sample; Set the elite threshold ; is the value corresponding to the elite threshold.

[0063] Obtain the elite sample set formula (23); Obtain the elite probability distribution parameters according to formula (24) and formula (25); Update the parameters according to formula (26) and formula (27) and the parameters ; End the loop.

[0064] and are both subsets of the probability distribution parameters ; are both subsets of the probability distribution parameters ;

[0065] Finally, a hardware-in-the-loop simulation platform equipped with a high-speed multiple unit train for tracking operation in the laboratory is used for simulation experiments. The experimental results show that the learned data-driven DNN has a high matching accuracy. At the same time, the reinforcement learning predictive control method can achieve the cooperative tracking operation of the virtual formation train.

[0066] To verify the effectiveness of the deep reinforcement learning data-driven model predictive algorithm proposed in this application for the cooperative operation control of virtual formation trains, this application uses a hardware-in-the-loop simulation platform equipped with a high-speed multiple unit train for tracking operation in the laboratory for simulation experiments. The experiment includes two parts: First, in the first group of experiments, a DNN model of the train is learned based on the sampled historical operation data of the train, and the effectiveness of the data-driven model is verified; in the second group of experiments, elastic interval formation control is performed on two unit trains under different operation scenarios to verify the effectiveness of the control algorithm for the formation control of virtual formation trains. The specific parameters of the control system are shown in Table 1.

[0067] Table 1 Control System Parameter Table Experiment 1: DNN training.

[0068] The experiment uses a DNN model with two fully connected layers, and each layer contains 500 neurons. During the training process of the DNN model, the rectified linear units (RELU) are selected as the activation function , and its representation is as follows: (28) where is the independent variable of the function and has no specific meaning.

[0069] According to the sampled historical operation data of the train running between a certain place and another area, such as Figure 9As shown, the DNN model is "offline" learned, and the effectiveness of the trained data-driven model is verified.

[0070] Figure 10 and Figure 11 are the speed prediction result and the predicted speed error respectively. It can be seen that the error between the predicted speed output by the DNN model designed in this application and the actual train speed can be guaranteed to be within the range of [-0.0706, 0.0749]. Figure 12 and Figure 13 are the displacement prediction result and the predicted displacement error respectively. It can be seen that the error between the predicted displacement output by the model and the actual train displacement can be guaranteed to be within the range of [-14.1046, 9.6591], verifying the effectiveness of the model.

[0071] To more intuitively analyze the performance of the DNN, the mean absolute error of speed (MAEv), the mean absolute error of displacement (MAEp), the root mean square error of speed (RMSEv), and the root mean square error of displacement (RMSEp) of the trained model are calculated according to formulas (29)-(32) to verify the effectiveness of the model. The calculations are as follows ; ; ; . The results show that the proposed DNN model can well estimate the high-speed train model, verifying the effectiveness of the model.

[0072] (29) (30) (31) (32) is the train speed output by the DNN; is the train displacement output by the DNN; is the actual running speed of the train; is the actual running displacement of the train.

[0073] This group of experiments "offline" learned the DNN model of the train. In subsequent experimental simulations, the DNN model will be further "online" learned by combining the online operation data of the train. Since the learning method is the same, it will not be described further later.

[0074] Experiment 2: Elastic formation control.

[0075] After verifying the prediction performance of the DNN, the performance of the train controller will be simulated and analyzed next. It is assumed that the train fleet consists of two trains and the impact of communication delay is not considered. Two train operation scenarios are set in this section of the experiment: (1) Train 1 runs at a constant speed of 250 km / h on a certain section of the road, and Train 2 collaboratively tracks Train 1 with an initial speed of 260 km / h to reach the formation operation state. The initial distance between the two trains is 1.2 km; (2) Train 1 runs at a variable speed along a set reference speed curve on a certain section of the road, and Train 2 collaboratively tracks Train 1 with an initial speed of 310 km / h and conducts dynamic spacing formation operation control. The initial distance between the two trains is 1.6 km. The specific description of the initial state parameters of the train is shown in Table 2.

[0076] Table 2 Train Initial Condition Table

[0077] Figures 14 to 21 They are respectively the operation simulation results under Scenario 1 and Scenario 2. Figures 14 - 15 、 Figures 18 - 19 They are respectively the operation state tracking situations of the two trains in Scenario 1 and Scenario 2. It can be seen that in both scenarios, the trailing train can track the leading train in a short time and always maintain the coupled operation state, verifying the effectiveness of the algorithm mentioned in this application. Figure 16 、 Figure 20 They are respectively the state tracking error situations of the two trains in Scenario 1 and Scenario 2. DAS and DAO are respectively the errors between the actual spacing and the safety spacing and the optimal spacing, and SD is the speed error between the two trains. It can be seen that after 600 s, the tracking error between the two trains in Scenario 1 is approximately 0, and the coupled operation of the trains can be achieved; while in Scenario 2, due to the influence of the frequent speed change of the train, the tracking error will have some small fluctuations, but the coupled operation of the trains can still be maintained.

[0078] Figure 17 、 Figure 21 They are respectively the acceleration change situations of the two trains in Scenario 1 and Scenario 2. It can be seen that their change ranges are guaranteed within the specified values, meeting the requirements of passenger comfort. In the initial stage, due to the limitations of the optimization algorithm, it will affect the sub-optimality of the solution and cause some fluctuations, but the control quantity is still guaranteed to be between [-1, 1].

[0079] In an exemplary embodiment, a formation collaborative tracking control device is provided, including: A data acquisition module, configured to acquire the operation data of the virtual formation train; the virtual formation train is a coupled train cluster composed of multiple single-particle trains.

[0080] An output module, configured to input operation data into a deep neural network dynamics model to obtain output state data; the deep neural network dynamics model is obtained by training a deep neural network using historical operation data and corresponding output state data as an offline data set.

[0081] A target function determination module, configured to use a reinforcement learning method to determine an elastic optimal control target function of a virtual formation train based on the output state data according to an elastic interval evaluation model; the elastic interval evaluation model is a mathematical model that calculates a safe following distance based on the relative braking distance between two adjacent trains and operates in coordination with a set following distance.

[0082] A solving module, configured to solve the elastic optimal control target function using a cross-entropy algorithm to obtain a control sequence; the control sequence is used to perform formation cooperative tracking control on the virtual formation train.

[0083] In an exemplary embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0084] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0085] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0086] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0087] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0088] In this application, specific examples are used to illustrate the principles and implementation manners of the application. The description of the above embodiments is only for helping to understand the method and its core idea of the application; meanwhile, for those of ordinary skill in the art, according to the idea of the application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the application.

Claims

1. A formation cooperative tracking control method, characterized in that, The formation cooperative tracking control method includes: Obtaining the operation data of the virtual formation train; the virtual formation train is a coupled train cluster composed of multiple single-particle trains; Inputting the operation data into the deep neural network dynamics model to obtain output state data; the deep neural network dynamics model is obtained by training the deep neural network using historical operation data and corresponding output state data as an offline data set; Using the method of reinforcement learning, based on the output state data, determining the elastic optimal control objective function of the virtual formation train according to the elastic interval evaluation model; the elastic interval evaluation model is a mathematical model that calculates the safe following distance based on the relative braking distance between two adjacent trains and runs in coordination with a set interval following distance; Solving the elastic optimal control objective function by using the cross-entropy algorithm to obtain a control sequence; the control sequence is used for the formation cooperative tracking control of the virtual formation train.

2. The formation cooperative tracking control method according to claim 1, characterized in that The method for determining the deep neural network dynamics model specifically includes: Obtaining an offline data set; Performing normalization processing on the offline data set to obtain a processed data set; Constructing a deep neural network; Inputting the processed data set into the deep neural network, and using the ADAM algorithm to train the parameters of the deep neural network with the goal of minimizing the loss function to obtain a trained deep neural network; the parameters include: weights; Determining the trained deep neural network as the deep neural network dynamics model.

3. The formation cooperative tracking control method according to claim 2, characterized in that Using the min-max normalization method to perform normalization processing on the offline data set to obtain a processed data set; the expression of the processed data set is: ; Among them, is the processed dataset; is the offline dataset; is the minimum value in the offline dataset; and are both the sample numbers in the offline dataset; is the total number of samples in the offline dataset; is the maximum value in the offline dataset.

4. The formation cooperative tracking control method according to claim 2, characterized in that, The expression of the loss function is: ; Among them, is the loss function; is the sampling time; is the sampling moment; is the operation state of the train at the sampling moment; is the operation state of the train at the sampling moment; is the control input of the train at the sampling moment; and are both weight coefficients of the deep neural network.

5. The formation cooperative tracking control method according to claim 1, wherein The elastic optimal control objective function specifically includes: ; Among them, is the optimal elastic control objective function; is the prediction horizon; is the sampling time; is any time within the prediction horizon; is the reward function of the train at time ; is the operating state of the train at time ; is the control input of the train at time .

6. The formation cooperative tracking control method according to claim 1, characterized in that Solving the elastic optimal control objective function by using the cross-entropy algorithm to obtain a control sequence, specifically including: Determining initial parameters; the initial parameters include: setting the iteration time, ratio, and N train samples; the train samples are obtained by parametric distribution sampling based on the probability distribution function; Determining the fitness function according to the elastic optimal control objective function; Based on the parameter combination at the current iteration time, determining the fitness function value according to the fitness function; the parameter combination includes: ratio and train samples; Judging whether the iteration stop condition is satisfied; the iteration stop condition is that the current iteration time reaches the set iteration time or the fitness function value at the current iteration time is within the set elite threshold interval; If not, then determining the elite sample set based on the fitness function value, updating the parameters of the probability distribution function based on the set learning step size parameter, and returning "Based on the parameter combination at the current iteration time, determining the fitness function value according to the fitness function"; If so, then determining the control sequence according to the parameter combination at the current iteration time.

7. An formation cooperative tracking control device, characterized in that The formation cooperative tracking control device includes: A data acquisition module, configured to acquire the operation data of the virtual formation train; the virtual formation train is a coupled train cluster composed of multiple single-particle trains; An output module, configured to input the operation data into a deep neural network dynamics model to obtain output state data; the deep neural network dynamics model is obtained by training a deep neural network using historical operation data and corresponding output state data as an offline data set; A target function determination module, configured to use a reinforcement learning method to determine an elastic optimal control target function of a virtual formation train based on the elastic interval evaluation model according to the output state data; the elastic interval evaluation model is a mathematical model that calculates a safe following distance based on the relative braking distance between two adjacent trains and runs in coordination with a set interval following distance; A solution module, configured to solve the elastic optimal control target function by using a cross-entropy algorithm to obtain a control sequence; the control sequence is used to perform formation cooperative tracking control on the virtual formation train.

8. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the formation cooperative tracking control method according to any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the formation cooperative tracking control method according to any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the formation cooperative tracking control method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Train autonomous scheduling deep reinforcement learning method and module

    CN111369181A

  • Multi-user terahertz array safety modulation method based on cross entropy iteration

    CN113242073A

  • Virtual marshalling train control method based on elastic tracking model

    CN114834503A

  • Virtual marshalling train reference curve calculation method based on improved reinforcement learning algorithm

    CN116090336A

  • Cross entropy reinforcement learning variable speed limit control method based on refined return mechanism

    CN116189464A

Cited By

  • Virtual marshalling train spacing dynamic control method, medium and equipment

    CN121469675A