A TACS Network Resource Allocation Method Based on Deep Reinforcement Learning

By adopting a multi-agent model based on deep reinforcement learning in rail transit TACS communication, the convergence problems and signaling priorities in the prior art are solved, and efficient network resource allocation and more accurate decision-making are achieved.

CN116405904BActive Publication Date: 2025-05-30SHANGHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310358415.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-06
Publication Date
2025-05-30
Estimated Expiration
2043-04-06

AI Technical Summary

Technical Problem

The prior art has convergence problems in the allocation of rail transit TACS communication resources, cannot handle continuous operations, and does not consider the signaling priority of communication between trains.

Method used

A multi-agent deep reinforcement learning model based on deep reinforcement learning is adopted, combined with deep deterministic strategy gradient algorithms with deep learning, attention mechanism and priority experience replay, a multi-agent deep reinforcement learning model in a tunnel environment is built, and trained to achieve efficient allocation of network resources.

Benefits of technology

It improves the efficiency of network resource allocation in a tunnel environment, enhances the efficiency of information interaction and the accuracy, stability and convergence speed of decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116405904B_ABST
    Figure CN116405904B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for allocating TACS network resources based on deep reinforcement learning, which includes the following steps: constructing a communication scenario in a tunnel environment for the train autonomous operation system, and dividing signaling priorities for the communication services within the communication scenario; on the premise of constraining the maximum delay of the communication link between trains in the communication scenario, constructing a multi-agent deep reinforcement learning model with the goal of maximizing the throughput of the train autonomous operation system; using the deep deterministic policy gradient algorithm based on deep learning fitting, attention mechanism and prioritized experience replay to train the multi-agent deep reinforcement learning model; based on the trained multi-agent deep reinforcement learning model, obtaining the optimal action policy, and realizing the allocation of network resources based on the optimal action policy. Compared with the prior art, the present invention effectively improves the total capacity of the TACS system and reduces the transmission delay of the T2T link.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of rail transit, and in particular to a method for allocating TACS network resources based on deep reinforcement learning. Background Technique

[0002] With the acceleration of the urbanization development process and the enhancement of people's awareness of green travel, the proportion of rail transit in public transportation is getting higher and higher. It is becoming more and more urgent to ensure the sustainable development of rail transit, improve the operation quality and efficiency, and achieve a more safe, efficient and high-tech rail transit. The train control system based on communication in rail transit is widely used. With the continuous upgrading of the train control system, the TACS communication method can significantly improve the train control safety through direct information exchange between the front and rear trains.

[0003] In the Train Autonomous Circumambulation System (TACS) mode of train-to-train communication in rail transit, there is communication between train-to-train (T2T) and train-to-wayside (T2W). While ensuring the transmission of important information, the utilization rate of the entire communication resources should be improved. With the popularization of machine learning algorithms, in the TACS communication resource allocation scheme, a relatively advanced resource allocation scheme based on multi-agent deep reinforcement learning (MADRL) proposed by using the Deep Q-Network (DQN) algorithm. However, the DQN algorithm cannot guarantee convergence all the time, and the DQN algorithm cannot solve the prediction problem of continuous actions. The power in the simulation process is generally a discrete value, which does not conform to the real tunnel environment; at the same time, the signaling priority problem of communication between trains is not considered in the existing work. Summary of the Invention

[0004] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art and provide a method for allocating TACS network resources based on deep reinforcement learning. A multi-agent deep reinforcement learning model is established based on the communication scenario considering signaling priority in the tunnel environment. The multi-agent deep reinforcement learning model is trained using the deep deterministic policy gradient algorithm based on deep learning fitting, attention mechanism and prioritized experience replay. The network resource allocation result is obtained based on the trained multi-agent deep reinforcement learning model, which improves the network resource allocation efficiency in the tunnel environment.

[0005] The purpose of the present invention can be achieved by the following technical solutions:

[0006] The present invention provides a method for allocating TACS network resources based on deep reinforcement learning, including the following steps:

[0007] Construct a communication scenario in a tunnel environment for the train autonomous operation system, and divide signaling priorities for communication services within the communication scenario;

[0008] On the premise of constraining the maximum delay of the communication link between trains in the communication scenario, a multi-agent deep reinforcement learning model is constructed with the goal of maximizing the throughput of the train autonomous operation system;

[0009] Use the deep deterministic policy gradient algorithm based on deep learning fitting, attention mechanism, and prioritized experience replay to train the multi-agent deep reinforcement learning model;

[0010] Based on the trained multi-agent deep reinforcement learning model, obtain the optimal action policy, and realize the allocation of network resources based on the optimal action policy.

[0011] As a preferred technical solution, the construction process of the multi-agent deep reinforcement learning model includes the following steps:

[0012] Take the communication link between each pair of trains as an agent. Each agent obtains a state in the state space and correspondingly takes an action in the action space, and allocates network resources according to the policy, where the policy is determined by the action-value function representing the probability of an agent executing a certain action in a certain state;

[0013] Each agent interacts with the communication scenario to obtain a reward, which is used to determine the action to be executed when selecting the next new state, where the reward is obtained based on the capacity of the communication link between trains and the communication link between the train and the trackside, and the delay of the communication link between trains.

[0014] As a preferred technical solution, the reward is obtained by the following formula:

[0015]

[0016] where T 0 is the maximum delay of the communication link between trains, λ c , λ d , λ p are weights, (T 0 -U t ) is the time used for transmission, C c [m] is the channel capacity of the m-th communication link between the train and the trackside, C d [n] is the channel capacity of the n-th communication link between trains, N is the number of communication links between the train and the trackside and the number of communication links between trains, respectively.

[0017] As a preferred technical solution, the process of training the multi-agent deep reinforcement learning model includes the following steps:

[0018] Initialize the Actor deep learning neural network of each agent in the multi-agent deep reinforcement learning model, and the Critic deep learning neural network based on the multi-head attention mechanism and decentralized.

[0019] Reset the vehicle networking environment, update the train position, obtain the updated large-scale fading and updated small-scale fading according to the updated train position, obtain training samples and store them in a preset experience replay pool.

[0020] Randomly select small-sized training samples from the experience replay pool to form a data set, and train the Actor deep learning neural network and Critic deep learning neural network of each agent.

[0021] As a preferred technical solution, the training samples include the input state and the corresponding output action, the state at the next moment, and the common reward obtained after all agents execute the actions. The input state and the state at the next moment both include the channel gain of the communication link between trains in the current time slot, the channel gain of the communication link between the train and the trackside in the current time slot, link interference, the selection of adjacent sub-channels in the previous time slot, transmission load, and the remaining time satisfying the delay constraint.

[0022] As a preferred technical solution, the channel gains of the communication link between trains and the communication link between the train and the trackside are: g n [m] = α n h n [m], where h n [m] is the small-scale fading power component of the nth communication transmitter between trains on the mth sub-band, and α n is the large-scale fading component of the nth communication transmitter between trains.

[0023] As a preferred technical solution, the Actor deep learning neural network and Critic deep learning neural network of the agent are trained based on the loss function, and the loss function is obtained based on the rewards and probability action value functions obtained by each agent through interacting with the communication scenario.

[0024] As a preferred technical solution, the process of allocating network resources based on the optimal action strategy includes the following steps:

[0025] Based on the above-mentioned optimal action strategy, obtain the optimal transmit power and channel allocation data of the communication link between trains, and based on the trained multi-agent deep reinforcement learning model, obtain the capacity of the communication link between trains and between trains and the trackside, as well as the transmission delay of the communication link between trains.

[0026] As a preferred technical solution, in the communication scenario, it includes a core network, a base station, and a backbone network based on LTE-M, as well as an on-vehicle controller and a target controller. The on-vehicle controller is used to realize communication between adjacent trains according to the line plan, exchange train position and resource information, and generate a movement authority. The target controller is used to register and unlock the occupancy of line resources, track non-communication trains, recycle train operation resources, and degrade the route safety protection.

[0027] As a preferred technical solution, the process of dividing signaling priorities for communication services in the communication scenario includes the following steps:

[0028] Divide the train automatic operation, train automatic protection, and train automatic monitoring services into the first priority level, and divide the passenger information system, video surveillance system, and high data rate services into the second priority level.

[0029] Compared with the prior art, the present invention has the following advantages:

[0030] (1) Good allocation effect for tunnel environment: Based on the communication scenario considering signaling priorities in the tunnel environment, this method establishes a multi-agent deep reinforcement learning model, and uses the deep deterministic policy gradient algorithm based on deep learning fitting, attention mechanism, and prioritized experience replay to train the multi-agent deep reinforcement learning model. Based on the trained multi-agent deep reinforcement learning model, the network resource allocation result is obtained. Compared with the existing method, this method divides signaling priorities for each communication service in the communication scenario and uses the deep deterministic policy gradient algorithm based on deep learning fitting, attention mechanism, and prioritized experience replay, improving the network resource allocation efficiency in the tunnel environment.

[0031] (2) Improve the efficiency of information interaction, the accuracy of decision-making, and the stability of training: In the TACS scenario, each agent needs to interact with other agents to complete tasks collaboratively. The traditional method is to equally consider the information of all agents and then make decisions. However, in complex scenarios, this method may lead to information redundancy and unnecessary communication overhead. This method enables agents to pay more attention to important information by introducing an attention mechanism, reducing unnecessary communication, thereby improving the efficiency of information interaction and more accurately predicting future states and behaviors. In addition, the performance of the MADDPG algorithm in multi-agent games is affected by the stability of training. The attention mechanism can enable the model to learn and converge more stably, thereby improving the stability and convergence speed of training. Brief Description of the Drawings

[0032] Figure 1 It is a flowchart of the TACS network resource allocation method based on deep reinforcement learning in Embodiment 1;

[0033] Figure 2 It is a schematic diagram of each part of the communication scenario in Embodiment 1;

[0034] Figure 3 It is a schematic diagram during the model optimization process. Detailed Embodiments

[0035] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0036] Embodiment 1

[0037] As Figure 1 described, in response to the problems mentioned above, this embodiment first divides different priorities for the content of different communications in the train control system, and then models the TACS communication mode as a MADRL problem, and proposes a multi-agent deep deterministic policy gradient (MADDPG) algorithm based on distributed execution, which is applied to the communication scenario as Figure 2 described to solve the power control problem of the continuous actions of each agent. This method includes the following steps:

[0038] S1: Construct a T2T system based on LTE-M communication. This step includes the following sub-steps:

[0039] S11: The TACS system based on LTE-M communication is built on the existing communication-based train control system consisting of a core network, base stations, and a backbone network. However, the functions of the original computer interlocking and area controller are integrated into the VOBC. The VOBC can achieve direct communication between adjacent trains according to the line plan, complete the exchange of key information such as train positions and resources, and generate a moving authorization in a timely manner.

[0040] S12: An OC is added to register and unlock the occupancy of line resources. The OC realizes system degradation functions such as non-communication train tracking, train operation resource recovery, and route safety protection, achieving the compatibility and easy deployment of the signal system.

[0041] S2: According to different communication contents, the priorities of different signaling are divided to simulate the real scenario of TACS communication. This step includes the following sub-steps:

[0042] S21: Divide the priorities of different signaling according to different communication contents;

[0043] S22: 5 MHz in the 20 MHz bandwidth of LTE-M is set aside specifically for the transmission of services with higher priorities (such as ATO instructions), and the remaining 15 MHz bandwidth participates in the spectrum sharing resource allocation scheme, and the allocation method is as described in Table 1.

[0044] Table 1 Spectrum Sharing Resource Allocation

[0045]

[0046] S3: Build different resource allocation frameworks for the TACS scenario, and derive the signal-to-interference-plus-noise ratio formula and channel capacity formula for the T2T link and the T2W link. The derivation process is as follows:

[0047] The signal gain is: g n [m] = α n h n [m]

[0048] In the formula, h n [m] is the frequency-dependent small-scale fading power component, and α n is the large-scale fading effect, mainly the path loss.

[0049] The signal-to-interference-plus-noise ratio of the T2W link:

[0050] Among them, is the transmission power of the m-th T2W link, σ 2 is the noise power, is the transmission power of the n-th T2T transmitter on the m-th sub-band, m ∈ {1,..., M}, n ∈ {1,..., N}. ρ n[m] is the binary spectrum allocation metric, ρ n [m]=1 means that the nth T2T link uses the mth subband, otherwise ρ n [m]=0. Assume that each T2T link only accesses one subband, i.e., ∑ n ρ n [m]≤1.

[0051] Signal-to-interference-plus-noise ratio of the T2T link:

[0052] Among them, represents the transmission power of the nth T2T transmitter on the mth subband, represents the interference channel from the mth T2T transmitter on the mth subband to the nth T2T receiver, g n′,n [m] is similar.

[0053] Channel capacity of the T2W link:

[0054] Channel capacity of the T2T link:

[0055] Among them, B is the bandwidth of each spectrum subband.

[0056] S4: Model the spectrum sharing problem in TACS communication as a MADRL problem. This step includes the following substeps:

[0057] S41: Combine the reinforcement learning algorithm to model the spectrum sharing scenario as a MADRL problem. Combine Figure 3 , obviously, each T2T link acts as an agent, observes a state S t in the state space S, and accordingly takes an action A t in the action space A, and selects the subband and transmission power according to the policy π. The policy v can be determined by the action value function Q(S t ,A t ), which represents the probability that the agent executes a certain action A t in a certain state S t , that is, π(a∣s)=P(A t =a∣S t =s).

[0058] S42: Each agent interacts with the tunnel communication environment to obtain a reward R t , so as to guide itself to select the action A t+1 to be executed in the next new state S t+1 . The reward R t is determined by the system capacity of the T2T and T2W links and the delay constraint of the corresponding T2T link.

[0059] S43: Multiple T2T agents represent multiple trains in the tunnel scenario. All trains jointly explore the environment and improve the spectrum allocation and power control strategies according to their observations of the environmental state, as Figure 3 described.

[0060] S5: Apply the DDPG algorithm to the MADRL model to propose the MADDPG algorithm, and give the state space, reward function, value function, algorithm flow, etc. of the algorithm to design a distributed execution algorithm for reinforcement learning. This step specifically includes the following sub-steps:

[0061] S51: State space: S t ={G t , H t , I t-1 , E t-1 , F t , U t}

[0062] Among them, the channel gain G of the T2T link in the current time slot t , the channel gain H of the T2W link in the current time slot t , the interference I caused by other previous links t-1 , the selection E of the adjacent sub-channel in the previous time slot t-1 , the transmission load F t and the remaining time U that meets the delay constraint t .

[0063] Reward function:

[0064] In the formula, T 0 is the maximum acceptable delay, and λ c , λ d , λ p are the weights of the three parts. (T 0 -U t ) is the time used for transmission.

[0065] Final reward:

[0066] Among them, β ∈ [0, 1] is the attenuation factor.

[0067] Action value function:

[0068] Among them, α represents the learning rate.

[0069] Loss function:

[0070] Among them,

[0071] θ′ is the parameter of the target Q network.

[0072] S52: The specific algorithm process is as follows:

[0073] Initialize the parameters of the real network and the target network of the Actor and Critic for each agent;

[0074] Initialize the size B of the experience pool for each agent k ;

[0075] Reset the vehicle networking environment;

[0076] Update the vehicle position and the large-scale fading α;

[0077] The real Actor policy network outputs an action A according to the input state S t , and after the agent executes the action, it obtains the state S at the next moment t , and after all agents execute the actions and obtain the common reward R t+1 ; t ;

[0078] Update the small-scale fading of the channel;

[0079] Obtain the training data (S t , A t , R t , S t+1 );

[0080] Store the training data in the experience replay pool;

[0081] Randomly sample m training data from the experience replay pool to form a data set and send it to the real Actor network, the real Critic network, the target Actor network, and the target Critic network;

[0082] Set the Q estimate as:

[0083] Define the loss function of the online Critic evaluation network as:

[0084] Update the target Actor network;

[0085] Update the target Critic network;

[0086] Update all parameters δ of the current Actor network through the gradient backpropagation of the neural network;

[0087] If the number of online training times reaches the target network update frequency, update the target network parameters δ ′ and θ ′ .

[0088] S6: Consider the joint optimization problem in the continuous action space, and use the DDPG algorithm that includes three mechanisms: deep learning fitting, attention mechanism, and experience replay to optimize the deep reinforcement learning model. This step specifically includes the following sub-steps:

[0089] S61: Deep learning fitting means that the present invention uses deep neural networks with different parameters to fit the deterministic policy and action value function;

[0090] S62: Since the communication between trains actually only needs to focus on the communication between adjacent trains in the same direction, and there is no need to pay too much attention to trains traveling in different directions, introducing an attention mechanism can improve the efficiency and accuracy of information interaction.

[0091] This method adopts centralized training and distributed execution. The agent cannot observe the complete state of the environment, thus decentralizing.

[0092] Specifically, the following steps can be adopted to introduce the attention mechanism:

[0093] (1) Encode the direction and speed information of each agent into a vector as the input data.

[0094] (2) For each agent, calculate an attention vector based on information such as its direction and relative speed to other agents, which is used to reflect the degree of attention and importance of the current agent to other agents.

[0095] (3) Perform a dot product operation between the attention vector of each agent and the input vector of other agents to obtain the attention-weighted input vector of each agent to other agents.

[0096] (4) Use the attention-weighted input vector of each agent as the input to calculate the action or state of the vehicle.

[0097] S63: Experience replay first collects a sample pool by setting up an experience replay mechanism, and then randomly selects some small-sized samples from the sample pool for training.

[0098] S7: According to the optimized deep reinforcement learning model, obtain the optimal T2T transmission power selection model, channel allocation decision, better T2W link capacity, T2T system capacity, and lower T2T transmission delay. This step specifically includes the following sub-steps:

[0099] S71: Output the best action policy to obtain the optimal T2T transmission power and channel allocation policy;

[0100] S72: Based on the optimized deep reinforcement learning model, obtain a better T2W link capacity, T2T system capacity, and a lower T2T transmission delay.

[0101] As Figure 2 shown, the Vehicle On-Board Controller (VOBC) communicates with the Object Controller (OC), the Passenger Information System (PIS), and the Image Monitoring System (IMS) through the base station - backbone network - core network; among them, step S1 includes:

[0102] Regarding the spectrum sharing problem of multi-agent deep reinforcement learning based on continuous action space in the TACS communication scenario, the simulation power of previous reinforcement learning algorithms is generally discrete and cannot truly simulate the tunnel scenario. At the same time, the signaling priority problem of communication between trains is not considered in the existing work. In this embodiment, the communication methods are divided into two categories: T2T and T2W. According to the train safety importance, the communication content is divided into two different priorities, and the resource sharing is modeled as a multi-agent deep reinforcement learning problem. A Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm based on distributed execution is proposed. Each agent continuously interacts with the tunnel environment and observes its own local state, and all obtain a common reward. By aggregating the action sets of other agents, the Critic network is trained, thereby improving the power control selected by each agent. By designing the reward function and training mechanism, the multi-agent algorithm can achieve distributed resource allocation, effectively improving the total capacity of the TACS system and reducing the transmission delay of the T2T link.

[0103] Embodiment 2

[0104] This embodiment provides an electronic device, including: one or more processors and a memory. The memory stores one or more programs, and the one or more programs include instructions for executing the TACS network resource allocation method based on deep reinforcement learning as described in Embodiment 1.

[0105] Embodiment 3

[0106] This embodiment provides a computer-readable storage medium, including one or more programs for execution by one or more processors of an electronic device. The one or more programs include instructions for executing the TACS network resource allocation method based on deep reinforcement learning as described in Embodiment 1.

[0107] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A method for allocating network resources of TACS based on deep reinforcement learning, characterized in that, it includes the following steps: Construct a communication scenario in a tunnel environment for the train autonomous operation system, and divide signaling priorities for the communication services in the communication scenario; On the premise of constraining the maximum delay of the communication link between trains in the communication scenario, with the goal of maximizing the throughput of the train autonomous operation system, construct a multi-agent deep reinforcement learning model; Use the deep deterministic policy gradient algorithm based on deep learning fitting, attention mechanism and prioritized experience replay to train the multi-agent deep reinforcement learning model; Based on the trained multi-agent deep reinforcement learning model, obtain the optimal action policy, and realize the allocation of network resources based on the optimal action policy, The construction process of the multi-agent deep reinforcement learning model includes the following steps: Regard the communication link between each train and train as an agent. Each agent obtains a state in the state space and correspondingly takes an action in the action space, and allocates network resources according to the policy. Among them, the policy is determined by the action value function representing the probability of an agent executing an action in a certain state; Each agent interacts with the communication scenario to obtain rewards, which are used to determine the actions to be executed when selecting the next new state. Among them, the rewards are obtained based on the capacities of the communication links between trains and between trains and the trackside, and the delay of the communication link between trains; The rewards are obtained by the following formula: , wherein, is the maximum time delay of the communication link between trains, , , are weights, is the time used for transmission, is the th channel capacity of the communication link between the train and the trackside, is the th channel capacity of the communication link between trains, , are respectively the number of communication links between the train and the trackside and the number of communication links between trains.

2. A method for allocating network resources of TACS based on deep reinforcement learning according to claim 1, characterized in that, The process of training the multi-agent deep reinforcement learning model includes the following steps: Initialize the Actor deep learning neural network of each agent in the multi-agent deep reinforcement learning model, and the Critic deep learning neural network based on the multi-head attention mechanism and decentralized; Reset the vehicle networking environment, update the train positions, obtain the updated large-scale fading and updated small-scale fading according to the updated train positions, obtain training samples and store them in a preset experience replay pool; Randomly select small-sized training samples from the experience replay pool to form a data set, and train the Actor deep learning neural network and Critic deep learning neural network of each agent.

3. A method for allocating network resources of TACS based on deep reinforcement learning according to claim 2, characterized in that, The training samples include the input state and the corresponding output actions, the state at the next moment, and the common rewards obtained after all agents execute the actions. The input state and the state at the next moment both include the channel gain of the communication link between trains in the current time slot, the channel gain of the communication link between trains and the trackside in the current time slot, link interference, the selection of adjacent sub-channels in the previous time slot, transmission load, and the remaining time satisfying the delay constraint.

4. A method for allocating network resources of TACS based on deep reinforcement learning according to claim 3, characterized in that, The channel gains of the communication link between trains and the communication link between trains and the trackside are as follows: , where is the small-scale fading power component of the th train-to-train communication transmitter on the th sub-band, and is the large-scale fading component of the th train-to-train communication transmitter.

5. A method for allocating TACS network resources based on deep reinforcement learning according to claim 2, characterized in that, the Actor deep learning neural network and the Critic deep learning neural network of the agent are trained based on a loss function, and the loss function is obtained based on the rewards obtained by each agent through interacting with the communication scenario and the probability action value function.

6. A method for allocating TACS network resources based on deep reinforcement learning according to claim 1, characterized in that, the process of allocating network resources based on the optimal action policy includes the following steps: Based on the optimal action policy, obtain the transmission power and channel allocation data of the optimal train-to-train communication link. Based on the trained multi-agent deep reinforcement learning model, obtain the capacity of the train-to-train communication link and the train-to-wayside communication link, and the transmission delay of the train-to-train communication link.

7. A method for allocating TACS network resources based on deep reinforcement learning according to claim 1, characterized in that, in the communication scenario, it includes a core network, a base station, and a backbone network based on LTE-M, as well as an on-vehicle controller and a target controller. The on-vehicle controller is used to realize communication between adjacent trains according to the line plan, exchange train position and resource information, and generate a movement authorization. The target controller is used to register and unlock the occupancy of line resources, track non-communicating trains, recycle train operation resources, and degrade the route safety protection.

8. A method for allocating TACS network resources based on deep reinforcement learning according to claim 1, characterized in that, the process of dividing signaling priorities for communication services in the communication scenario includes the following steps: Divide the train automatic operation, train automatic protection, and train automatic monitoring services into the first priority, and divide the passenger information system, video surveillance system, and high data rate services into the second priority.