Urban rail transit vehicle-ground wireless communication network coverage optimization method

By constructing a quantitative model and using deep reinforcement learning methods in urban rail transit, the vehicle-to-ground wireless communication network was optimized, solving the problem of unnecessary braking when trains disembark in high-speed moving scenarios and improving network coverage quality and train operation efficiency.

CN121151833AActive Publication Date: 2025-12-16BEIJING JIAOTONG UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511521052.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2025-12-16
Estimated Expiration
2045-10-23

AI Technical Summary

Technical Problem

In urban rail transit, existing technologies for optimizing vehicle-to-ground wireless communication network coverage are ill-suited to high-speed moving scenarios, cannot effectively reduce the probability of unnecessary train braking, are difficult to deploy based on existing infrastructure, and do not adequately consider the impact of network coverage quality on train operating efficiency.

Method used

Based on the train braking model, a quantitative model is constructed to calculate the minimum train tracking interval and track throughput capacity. A deep reinforcement learning state space and reward function are designed, and a multi-agent deep deterministic gradient policy algorithm is used to adjust the base station radio frequency parameters and optimize network coverage quality.

Benefits of technology

It reduces the probability of unnecessary braking by trains, improves network coverage quality, enhances train operation efficiency and communication performance, and dynamically adapts to complex and ever-changing wireless environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121151833A_ABST
    Figure CN121151833A_ABST
Patent Text Reader

Abstract

The invention provides an urban rail transit train-ground wireless communication network coverage optimization method, which comprises the following steps of: based on a train braking model, calculating a train minimum tracking spacing distance and line trafficability by considering unnecessary train braking triggered by communication time delay and packet loss; establishing a quantitative model of the influence of the urban rail vehicle-ground wireless network on the train operation efficiency by considering communication time delay and packet loss and taking the triggering probability of unnecessary service braking and emergency braking of the train as evaluation indexes; for the urban rail vehicle-ground wireless network topology structure in the quantitative model, calculating the train receiving power and the signal to interference plus noise ratio in combination with a parameter-adjustable antenna model; the method comprises the following steps of: designing a state space, an action space and a reward function of deep reinforcement learning by taking each base station in an urban rail vehicle-ground wireless network as an intelligent agent; based on a multi-agent depth deterministic gradient strategy algorithm, a base station agent radio frequency parameter adjustment strategy is obtained through centralized learning and training of a value-strategy network; and executing a radio frequency parameter adjustment strategy through all base station intelligent agents.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of wireless network, and particularly relates to a method for optimizing coverage of train-ground wireless communication network in urban rail transit. BACKGROUND

[0002] The communication-based train control system (CBTC) is widely used in urban rail transit and is the core of ensuring the safe and efficient operation of trains. As an important part of CBTC, the train-ground wireless communication network is responsible for real-time transmission of key data such as train operation status and control information. However, due to factors such as terrain and interference, the transmission quality of signals differs in different areas. At the same time, the high-speed operation of trains makes the wireless channel state change rapidly with the train position and speed. When the communication performance cannot meet the train operation requirements, for example, the train behind cannot receive the status of the train in front in time, it needs to brake to ensure safety, which may trigger a chain reaction of subsequent trains, seriously affecting the train operation efficiency.

[0003] In recent years, reinforcement learning (RL) methods have been applied to solve the problem of wireless network adaptive optimization, which has the ability of dynamic environment perception and adaptive decision-making, can simulate the learning and decision-making process of humans in complex environments, and quickly find network optimization solutions through continuous trial and error and feedback, greatly improving optimization efficiency. Therefore, reinforcement learning methods can be used to solve the train-ground wireless network coverage optimization problem according to the requirements and characteristics of urban rail transit scenarios. Through interaction with complex and variable environments, autonomous learning of parameter adjustment strategies can be achieved to dynamically regulate base station radio frequency parameters and improve network coverage quality.

[0004] However, the current wireless network adaptive optimization method has the following limitations:

[0005] Not applicable to high-speed mobile scenarios: In general public network scenarios, existing wireless network adaptive optimization methods generally train and deploy RL models for low-speed mobile users and macro cell or dense cell scenarios. However, in urban rail transit scenarios, due to the high-speed movement of trains, the channel environment of train-ground communication changes more rapidly, and the path loss and signal fading characteristics are also different from general cellular networks. It is necessary to pay attention to these more rapid and complex dynamic changes.

[0006] Difficult to deploy and apply based on existing urban rail transit infrastructure: Current network optimization research for high-speed mobile scenarios mostly focuses on beam control, antenna deployment or reconfigurable intelligent surface control, and the methods used rely on millimeter wave communication technology or additional hardware facilities, making it difficult to deploy and apply based on existing urban rail transit infrastructure.

[0007] Less attention is paid to the influence of network coverage quality on train operation efficiency: most of the existing researches aim to optimize received signal strength, spectrum efficiency or coverage probability. In the urban rail transit scenario, the train operation density is high, and the wireless network serves the key data transmission of the train operation control system, so the influence of network coverage quality on train operation efficiency needs to be concerned. SUMMARY

[0008] Embodiments of the present application provide a method for optimizing the coverage of a train-ground wireless communication network in urban rail transit, which solves the technical problems existing in the prior art.

[0009] To achieve the above-mentioned purpose, the present application adopts the following technical solutions.

[0010] A method for optimizing the coverage of a train-ground wireless communication network in urban rail transit, comprising:

[0011] S1 Based on the train braking model, considering the train unnecessary braking triggered by communication delay and packet loss, calculating the minimum tracking interval distance of the train and the line passing capacity;

[0012] S2 Considering the communication delay and packet loss, taking the train unnecessary braking and emergency braking triggering probability as the evaluation index, constructing a quantitative model of the influence of the urban rail train-ground wireless network on the train operation efficiency;

[0013] S3 For the topology structure of the urban rail train-ground wireless network in the quantitative model, combining the parameter-adjustable antenna model, calculating the train received power and signal-to-interference-and-noise ratio;

[0014] S4 Taking each base station in the urban rail train-ground wireless network as an intelligent agent, designing the state space, action space and reward function of deep reinforcement learning;

[0015] S5 Based on the multi-agent deep deterministic gradient policy algorithm, obtaining the base station intelligent agent radio frequency parameter adjustment strategy through the centralized learning and training strategy-value network;

[0016] S6 Through the distributed execution of the base station intelligent agent radio frequency parameter adjustment strategy by all base station intelligent agents, the adaptive optimization of the coverage quality of the urban rail train-ground wireless network is carried out.

[0017] Preferably, step S1 specifically comprises:

[0018] Based on the periodically received MA information, the minimum tracking interval distance of the front and rear trains is calculated by the formula

[0019] (1)

[0020] The minimum tracking interval distance of the front and rear trains is calculated; in the formula: is the interval time of two adjacent train MA information reception The distance traveled by subsequent trains during this period, It is the safe braking distance for following trains, where safe braking distance... Depending on the emergency braking trigger speed of the train's current position And the train's trajectory after it triggers emergency braking at that point;

[0021] Through

[0022] (2)

[0023] Calculate the time required for the following train to reach the service braking curve at the current speed. In the formula: The difference between the train's running speed and the normal braking trigger speed, This is the commonly used braking deceleration. If the interval between two consecutive MA messages received by the train... Exceed If so, the train will trigger the service brake;

[0024] set up The longest transmission delay allowed by ATP, if Exceed If this occurs, the train control system will immediately apply emergency braking until the train comes to a complete stop.

[0025] Preferably, the setting The longest transmission delay allowed by ATP, if Exceed If the train comes to a complete stop, the train control system will immediately output an emergency brake, and the process will include the following:

[0026] when If the train does not trigger the service brake, then the train will proceed via a bypass system.

[0027] (3)

[0028] Calculate trains Speed ​​is the distance traveled before receiving the next MA message. If the train triggers the service brake, then...

[0029] (4)

[0030] Calculate trains The speed is the distance traveled before receiving the next MA message;

[0031] when If the train does not trigger service or emergency braking, then the train's braking force is calculated using equation (3). The running distance before the next MA information is received; if the train triggers an emergency brake, then the passing type is

[0032] (5)

[0033] The minimum tracking distance of the train is calculated; if the train runs at the minimum tracking distance, then the passing type is

[0034] (6)

[0035] The time interval of the front train and the rear train continuously passing the same point is calculated In formula (6), is the protection margin and the ranging error, is the train length;

[0036] The passing type is calculated by formula

[0037] (7)

[0038] The line passing capacity is calculated .

[0039] Preferably, step S2 specifically comprises:

[0040] The passing type is calculated by formula

[0041] (8)

[0042] The probability of the train causing unnecessary service braking is calculated; in the formula: represents the th period, represents the number of continuously lost data packets, is the packet loss rate of the th period;

[0043] The passing type is calculated by formula

[0044] (9)

[0045] The minimum number of continuous lost packets causing unnecessary service braking of the train is calculated; in the formula: is the data packet transmission delay of the train in the th period, is the packet sending period;

[0046] The passing type is calculated by formula

[0047] (10)

[0048] The probability of unnecessary emergency braking of the train is calculated; in the formula: represents the number of continuously lost data packets;

[0049] By formula

[0050] (11)

[0051] Calculate the minimum number of consecutive packet loss representing triggering unnecessary emergency braking of the train.

[0052] Preferably, step S3 specifically comprises:

[0053] By formula

[0054] (12)

[0055] (13)

[0056] Calculate the antenna gain of the antenna model; in formula (12): and respectively represent the horizontal angle and the vertical angle between the user equipment and the antenna, and respectively represent the horizontal gain and the vertical gain; in formula (13): and respectively represent the azimuth angle and the downtilt angle of the antenna, and respectively represent the horizontal beam width and the vertical beam width of the antenna;

[0057] Considering the free space propagation loss and the small-scale fading model subject to the Rice distribution, by formula

[0058] (14)

[0059] Calculate the received signal power of the train ; in the formula: is the transmission power of the base station , is the antenna gain of the base station, is the power loss caused by large-scale fading, is the power loss caused by small-scale fading;

[0060] By formula

[0061] (15)

[0062] Calculate the influence of the system internal and external interference and noise in the signal transmission process, which is represented by the signal-to-interference-noise ratio; in the formula: represents the received signal power of the train received from the serving cell , and the set of all base stations in the network is denoted as , the interference signal power from other base stations in the system, denotes the noise power, denotes the interference outside the system.

[0063] Preferably, in step S4:

[0064] Each base station in the urban rail train-ground wireless network is regarded as an agent, so that the respective base station agent can interact with the train-ground wireless network environment;

[0065] The state range observed by each base station agent includes the train state information in the cell covered by the base station agent and the radio frequency parameter information of the base station agent; the state space function of each base station agent is

[0066] (16)

[0067] In the formula: denotes the train state information of the service cell, including the position, reference signal received power, signal-to-interference-and-noise ratio, data packet transmission delay and packet loss rate of the train, is the radio frequency parameter information of the base station;

[0068] The action space of each base station agent includes the transmission power and antenna downtilt angle adjustment of the base station agent; the action space function of each base station agent is

[0069] (17)

[0070] In the formula: is the selected base station transmission power, is the selected antenna downtilt angle, both of which are in continuous space;

[0071] The reward function of each base station agent is used to guide the base station agent to learn a collaborative radio frequency parameter adjustment strategy that takes into account performance and interference control, and to evaluate the comprehensive influence of the radio frequency parameter configuration of each base station agent on the communication quality, the probability of unnecessary train braking triggering and the rationality of parameter adjustment; the reward function of each base station agent is

[0072] (18)

[0073] In the formula: denotes the improvement of the communication quality of the train, which is obtained by normalizing and weighting the RSRP and SINR of the train at the current position; denotes the probability of unnecessary train braking triggering, which is obtained by weighting the probabilities of unnecessary service braking triggering and emergency braking triggering obtained from the transmission delay and packet loss rate; denotes the radio frequency parameter configuration of the base station agent, , and respectively represent the weights of , and ; N represents the number of trains.

[0074] Preferably, step S5 specifically comprises:

[0075] S501 determining the interaction environment of the base station agent with the train-ground wireless network, the number of base station agents, the total number of rounds of training, and the maximum number of steps per round;

[0076] S502 initializing the policy network and the value network of each base station agent and their corresponding target networks, and setting an experience replay buffer B;

[0077] S503 for each round of training, resetting the interaction environment state of the base station agent with the train-ground wireless network, and obtaining an initial observation ;

[0078] S504 at each step in each round of training , each base station agent generates an action by combining the local observation state through the policy network , and adds exploration noise;

[0079] S505 taking the set of actions generated by all base station agents as a joint action , and executing the joint action in the interaction environment;

[0080] S506 after the base station agent and the train-ground wireless network execute the joint action, returning a new state set and a reward ;

[0081] S507 storing the tuple , , , ) composed of the state , the action , the reward , and the new state in the experience replay buffer B;

[0082] S508 randomly sampling a batch of training data from the experience replay buffer B, updating the value network parameters of each base station agent by minimizing the loss function, updating the policy network parameters by maximizing the expected Q value, and soft updating the target network;

[0083] S509 repeating the sub-steps S502 to S508 until the training of all rounds is completed;

[0084] S510 Based on the execution mode of sub-step S506, the base station agent collects local environment state information in real time, generates a radio frequency parameter adjustment strategy using the base station agent's own strategy network, and performs distributed transmission power adjustment and antenna tilt adjustment actions, and optimizes the urban rail train-ground wireless network coverage quality in real time according to the change of the environment state.

[0085] As can be seen from the technical solutions provided by the above embodiments of the present application, the present application provides a kind of urban rail train-ground wireless communication network coverage optimization method, comprising: based on train braking model, unnecessary braking of train triggered by communication delay and packet loss is considered, train minimum tracking interval distance and line capacity are calculated;With the triggering probability of train unnecessary common braking and emergency braking as evaluation index, a quantitative model of the influence of urban rail train-ground wireless network on train operation efficiency is constructed;For the topology structure of urban rail train-ground wireless network in the quantitative model, the train received power and signal-to-noise ratio are calculated in combination with the parameter adjustable antenna model;With each base station in the urban rail train-ground wireless network as an agent, the state space, action space and reward function of deep reinforcement learning are designed;Based on multi-agent deep deterministic gradient policy algorithm, the base station agent radio frequency parameter adjustment strategy is obtained by centralized learning and training strategy-value network;The base station agent radio frequency parameter adjustment strategy is distributedly executed by all base station agents, and the adaptive optimization of urban rail train-ground wireless network coverage quality is carried out.The beneficial effects of the present application are as follows:

[0086] One: unnecessary braking of train triggered by communication delay and packet loss is considered, train minimum tracking interval distance and line capacity are calculated, and a quantitative model of urban rail train-ground wireless network coverage quality is constructed.

[0087] Two: with the triggering probability of train unnecessary common braking and emergency braking as evaluation index, an urban rail train-ground wireless network coverage quality evaluation method is proposed, and a comprehensive reward function integrating communication index, train control index and base station radio frequency parameter is designed, which takes into account the demand for train operation efficiency and the performance of communication network.

[0088] Three: an urban rail train-ground wireless coverage adaptive optimization method based on MADDPG is proposed, which adjusts the radio frequency parameters of each base station cooperatively, dynamically adapts to complex and variable wireless environment, reduces the triggering probability of unnecessary train braking, and improves the network coverage quality.

[0089] Additional aspects and advantages of the application will be described in the following description, which will become apparent from the following description, or will be learned by practice of the application. BRIEF DESCRIPTION OF DRAWINGS

[0090] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments description. Obviously, the drawings in the following description only constitute some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from these drawings without any creative effort.

[0091] Figure 1 A processing flow chart of a city rail transit train-ground wireless communication network coverage optimization method provided by the present application;

[0092] Figure 2 A train tracking interval schematic diagram of a city rail transit train-ground wireless communication network coverage optimization method provided by the present application;

[0093] Figure 3 A train-ground communication schematic diagram of a city rail transit train-ground wireless communication network coverage optimization method provided by the present application;

[0094] Figure 4 A network topology schematic diagram of a base station cooperative adjustment of radio frequency parameters of a city rail transit train-ground wireless communication network coverage optimization method provided by the present application. DETAILED DESCRIPTION

[0095] The embodiments of the present application will be described in detail below, and examples of the embodiments are shown in the drawings, wherein the same or similar notations represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present application, and cannot be interpreted as a limitation of the present application.

[0096] Those skilled in the art can understand that, unless specifically stated, the singular forms "a", "an" and "the" used herein also include the plural forms. It should be further understood that the use of the phrase "comprises" in the specification of the present application means that the features, integers, steps, operations, elements and / or components exist, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to the other element, or there can be intermediate elements. In addition, "connected" or "coupled" used herein can include wireless connection or coupling. The phrase "and / or" used herein includes any one of the associated listed items and all combinations of the associated listed items.

[0097] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless defined as herein.

[0098] To facilitate understanding of the embodiments of the present invention, the following will provide further explanation and description with reference to the accompanying drawings and several specific embodiments. These embodiments do not constitute a limitation on the embodiments of the present invention.

[0099] See Figure 1 This invention provides a method for optimizing the coverage of urban rail transit vehicle-to-ground wireless communication networks, comprising the following steps:

[0100] S1 is based on the train braking model, considering communication delay and unnecessary train braking triggered by packet loss, and calculates the minimum train tracking interval and track throughput capacity.

[0101] S2 takes into account communication latency and packet loss, and uses the probability of unnecessary frequent braking and emergency braking of trains as evaluation indicators to construct a quantitative model of the impact of urban rail vehicle-to-ground wireless network on train operation efficiency.

[0102] S3 calculates the train's received power and signal-to-interference-plus-noise ratio for the urban rail vehicle-to-ground wireless network topology in the quantization model, combined with the parameter-adjustable antenna model.

[0103] S4 uses each base station in the urban rail vehicle-to-ground wireless network as an intelligent agent to design the state space, action space and reward function of deep reinforcement learning.

[0104] S5 is based on a multi-agent deep deterministic gradient policy algorithm. It obtains the base station agent radio frequency parameter adjustment policy through centralized learning and training policy—value network (Actor-Critic).

[0105] S6 executes base station agent radio frequency parameter adjustment strategies in a distributed manner through all base station agents, adjusting the base station radio frequency to adaptively optimize the coverage quality of the urban rail vehicle-to-ground wireless network.

[0106] like Figure 2 The diagram shows the train tracking interval obtained based on the train braking model decomposition. In step S1, from the perspective of train control, based on... Figure 2The interval model shown analyzes the impact of communication latency and packet loss on train operating efficiency. Considering the scenario where communication latency and packet loss trigger unnecessary braking, a model is established to quantify the relationship between the interval between two consecutive MA (Manual Access) receptions, the minimum tracking interval, and the track capacity. Specifically, this includes:

[0107] In the train control system, trains periodically receive movement authority (MA) information and, in conjunction with track parameters and their current position and speed, calculate a speed protection curve to ensure that trains track each other with a certain safety protection interval. Transmission delays and packet loss will increase the interval between two consecutive MA information receptions, potentially causing unnecessary braking by following trains.

[0108] The minimum tracking interval between trains 1 and 2 is expressed as follows:

[0109] (1)

[0110] In formula (1), It is the interval between two consecutive MA messages received by the train. The distance traveled by train 2 during this period, This is the safe braking distance of train 2. Among them, the safe braking distance... Depending on the emergency braking trigger speed of the train's current position And the train's trajectory after it triggers emergency braking at that point.

[0111] The time required for the train to reach the service braking curve at its current speed is expressed as:

[0112] (2)

[0113] In formula (2), The difference between the train's running speed and the normal braking trigger speed, This is the commonly used braking deceleration. If the interval between two consecutive MA messages received by the train... Exceed If so, the train will trigger the service brake; The longest transmission delay allowed by ATP, if Exceed If this occurs, the train control system will immediately apply emergency braking, which will continue until the train comes to a complete stop.

[0114] exist During this period, whether the train triggered the brakes had an impact. The value of can be divided into two cases:

[0115] Scenario 1: When If the service brake is not activated, the train will proceed as follows: The running distance of the train before receiving the next MA information is represented as:

[0116] (3)

[0117] If the train triggers the service brake, the running distance of the train is represented as:

[0118] (4)

[0119] Case 2: When , the train does not trigger the service brake or the emergency brake, the running distance of the train is represented as formula (3); when the train triggers the emergency brake, the minimum tracking interval distance of the train is represented as:

[0120] (5)

[0121] If the train runs at the minimum tracking interval distance, the time interval of the two trains continuously passing the same point is represented as:

[0122] (6)

[0123] In formula (6), is the protection margin and the ranging error, is the length of the train, and the line passing capacity is represented as:

[0124] (7)

[0125] S2 is based on the process of the MA information transmission of the urban rail train-ground wireless communication, considering the communication delay and packet loss, taking the unnecessary service brake and emergency brake triggering probability of the train as the evaluation index, a quantitative model of the influence of the urban rail train-ground wireless network on the train operation efficiency is constructed, which specifically includes:

[0126] The triggering probability of the unnecessary service brake and the emergency brake of the train in each communication cycle is calculated by the following method: the interval of the train receiving MA information for two consecutive times caused by continuous packet loss is greater than the time required for the train to run at the current speed to touch the service brake curve or the maximum transmission delay allowed by the ATP. The specific formula is as follows:

[0127] (8)

[0128] (9)

[0129] (10)

[0130] (11)

[0131] Formula (8) represents the probability of triggering unnecessary service braking of the train, wherein represents the th cycle, represents the number of consecutively lost data packets, is the packet loss rate of the th cycle;

[0132] Formula (9) represents the minimum number of consecutive lost packets for triggering unnecessary service braking of the train, wherein is the data packet transmission delay of the train in the th cycle, is the packet transmission cycle;

[0133] Formula (10) represents the probability of triggering unnecessary emergency braking of the train, wherein represents the number of consecutively lost data packets;

[0134] Formula (11) represents the minimum number of consecutive lost packets for triggering unnecessary emergency braking of the train.

[0135] Step S3 calculates the train receiving power and signal-to-noise ratio in combination with the parameter-adjustable antenna model for the urban rail train-ground wireless communication scenario as shown in FIG. 3, and specifically includes: Figure 3 (1) The urban rail train-ground wireless network is a linear topology structure,

[0136] The distance between the base stations is serially distributed along the track direction to form a linear cell with continuous coverage. Trains are distributed at a certain interval on the track line and pass through each coverage cell in turn. The received signal power presents the characteristics of periodic fluctuation. The number of train users simultaneously accessing each cell is usually no more than 6. At the cell boundary, the received signal power of the train triggers handover when certain conditions are met.

[0137] (2) In order to realize dynamic adjustment of the base station radio frequency parameters, an antenna model with adjustable azimuth angle, downtilt angle, vertical beam width and horizontal beam width is introduced. The gain of the antenna can be represented as:

[0138] (12)

[0139] (13)

[0140] In formula (12), and respectively represent the horizontal angle and the vertical angle between the specific train (i.e., the train access unit TAU in Figure 3 ) and the antenna; and respectively represent the horizontal gain and the vertical gain, which are represented by formula (13).

[0141] In equation (13), and are the azimuth and downtilt angles of the antenna, respectively, and are the horizontal and vertical beamwidths of the antenna, respectively.

[0142] (3) The quality of the coverage of the urban rail train-ground wireless network is affected by the large-scale and small-scale fading in the channel environment, the internal and external interference of the system, and the noise. Considering the free space propagation loss and the small-scale fading model subject to the Rice distribution, the received signal power of the train is expressed as:

[0143] (14)

[0144] In equation (14), is the transmission power of the base station, is the gain of the base station antenna, is the power loss of the free space propagation, is the power loss caused by the small-scale fading.

[0145] (4) The signal-to-interference-and-noise ratio is used to represent the influence of the internal and external interference of the system and the noise in the signal transmission process, and is expressed as:

[0146] (15)

[0147] In equation (15), represents the received signal power of the train from the serving cell , and the set of all base stations in the network is denoted as , is the interference signal power from other base stations in the system, represents the noise power. The external interference is denoted as Although the urban rail train-ground wireless communication network uses a dedicated frequency point for data transmission in the track line area, it does not exclusively use the frequency point, and therefore still faces interference from devices using the same frequency point outside the system.

[0148] It should be understood that the adjustable antenna model involved in this embodiment of the present application uses a known model architecture, for example, the model disclosed by Tekgul et al. in “Joint Uplink-Downlink Capacity and Coverage Optimization via Site-Specific Learning of Antenna Settings” can be used.

[0149] ​​In step S4, each base station in the urban rail train-ground wireless network is regarded as an agent, and a state space, an action space and a reward function of deep reinforcement learning are designed, specifically including:

[0150] Each base station in the urban rail train-ground wireless network is regarded as an agent, and each base station agent interacts with the train-ground wireless network environment to obtain experience and learns an optimal radio frequency parameter adjustment strategy from the experience. The state observable by the base station agent is a local environment state, including train state information in a cell covered by the base station and radio frequency parameter information of the base station, the action space is adjustment of transmit power and antenna downtilt angle of the base station agent, and the reward function guides the base station agent to learn a collaborative radio frequency parameter adjustment strategy considering performance and interference control, and evaluates comprehensive influence of radio frequency parameter configuration of each base station on communication quality, probability of unnecessary train braking triggering and rationality of parameter adjustment. The state space, the action space and the reward function are specifically represented as follows:

[0151] (16)

[0152] (17)

[0153] (18)

[0154] In formula (16), represents train state information of the service cell, including position, reference signal received power, signal-to-noise ratio, data packet transmission delay and packet loss rate of the train, is radio frequency parameter information of the base station, and the two are combined into a multi-dimensional vector as the state space;

[0155] In formula (17), is selected transmit power of the base station, is selected antenna downtilt angle, and the value range of both is a continuous space;

[0156] In formula (18), represents improvement of communication quality of the train, which is obtained by normalizing and weighting RSRP and SINR of the train at the current position; represents probability of unnecessary train braking triggering, which is obtained by weighting calculation after obtaining probabilities of unnecessary service braking and emergency braking triggering from the transmission delay and the packet loss rate; represents radio frequency parameter configuration of the base station agent, , , respectively represent weights of the three; and N represents the number of trains.

[0157] Step S5 is based on Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm, and the base station agent radio frequency parameter adjustment strategy is obtained by centralized learning and training of Critic and Actor networks, combined with Figure 4 The network topology schematic diagram is shown in FIG. 1, and the specific step process is as follows:

[0158] S503, for each round of training, resetting the interaction environment state of the base station agent and the train-ground wireless network, and obtaining the initial observation value ;

[0159] S504, at each step in each round of training , each base station agent generates an action by a policy network combined with the local observation state , and adds exploration noise;

[0160] S505, taking the action set generated by all base station agents as joint action , and executing the joint action in the interaction environment;

[0161] S506, the base station agent and the train-ground wireless network return a new state set and a reward after executing the joint action in the interaction environment;

[0162] S507, storing the tuple ( , , , ) composed of the state , the action , the reward , and the new state in the experience replay buffer B;

[0163] S508, randomly sampling a batch of training data from the experience replay buffer B, updating the value network parameters of each base station agent by minimizing the loss function, updating the policy network parameters by maximizing the expected Q value, and soft updating the target network;

[0164] S509, repeating the sub-steps S502 to S508 until the training of all rounds is completed;

[0165] S510 After the model training process described above is completed, in the manner of step S6, the local environment state information is collected in real time by each base station agent, the radio frequency parameter adjustment strategy is generated by the respective Actor network, the distributed transmission power adjustment and antenna tilt adjustment actions are performed, and the metro train-ground wireless network coverage quality is optimized in real time according to the environment state change.

[0166] In summary, the present application provides a metro train-ground wireless communication network coverage optimization method, which comprises: based on the train braking model, considering the unnecessary train braking triggered by communication delay and packet loss, calculating the minimum tracking interval distance of the train and the line passing capacity; considering the communication delay and packet loss, taking the unnecessary and emergency braking trigger probability of the train as the evaluation index, constructing a quantitative model of the influence of the metro train-ground wireless network on the train operation efficiency; for the metro train-ground wireless network topology in the quantitative model, combining the parameter adjustable antenna model, calculating the train received power and signal-to-interference-and-noise ratio; taking each base station in the metro train-ground wireless network as an agent, designing the state space, action space and reward function of deep reinforcement learning; based on the multi-agent deep deterministic gradient policy algorithm, obtaining the radio frequency parameter adjustment strategy of the base station agent through the centralized learning and training strategy-value network; through the distributed execution of the radio frequency parameter adjustment strategy of the base station agent by all base station agents, the adaptive optimization of the metro train-ground wireless network coverage quality is performed. The beneficial effects of the present application are:

[0167] One: considering the unnecessary train braking triggered by communication delay and packet loss, calculating the minimum tracking interval distance of the train and the line passing capacity, and constructing a quantitative model of the metro train-ground wireless network coverage quality.

[0168] Two: taking the unnecessary and emergency braking trigger probability of the train as the evaluation index, a metro train-ground wireless network coverage quality evaluation method is proposed, and a comprehensive reward function integrating communication indicators, train control indicators and base station radio frequency parameters is designed, which takes into account the demand for train efficiency and communication network performance.

[0169] Three: a metro train-ground wireless coverage adaptive optimization method based on MADDPG is proposed, which adjusts the radio frequency parameters of each base station to dynamically adapt to the complex and changeable wireless environment, reduces the unnecessary train braking trigger probability, and improves the network coverage quality.

[0170] Those skilled in the art can understand that the modules or processes in the drawings are not necessarily required to implement the present application.

[0171] Those skilled in the art can clearly understand the present application by the description of the above embodiments. Based on such an understanding, the technical solutions of the present application can be embodied in the form of a software product, and the computer software product can be stored in a storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, and the like, and includes a plurality of instructions to cause a computer device (which can be a personal computer, a server, or a network device, and the like) to execute the methods described in the various embodiments or some parts of the embodiments of the present application.

[0172] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each of the embodiments mainly describes the difference from other embodiments. In particular, for the device or system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the related parts can be referred to the part of the description of the method embodiments. The above-described device and system embodiments are merely illustrative, and the units described as separate components can be or can not be physically separated, and the components displayed as units can be or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiments according to the actual needs. Those skilled in the art can understand and implement it without creative labor.

[0173] The above describes only the preferred embodiments of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for optimizing the coverage of a vehicle-to-ground wireless communication network in urban rail transit, characterized in that, include: S1 is based on the train braking model, considering communication delay and unnecessary train braking triggered by packet loss, and calculates the minimum train tracking interval and track throughput capacity. S2 takes into account communication latency and packet loss, and uses the probability of unnecessary frequent braking and emergency braking of trains as evaluation indicators to construct a quantitative model of the impact of urban rail vehicle-to-ground wireless network on train operation efficiency. S3 calculates the train's received power and signal-to-interference-plus-noise ratio for the urban rail vehicle-to-ground wireless network topology in the quantization model, combined with the parameter-adjustable antenna model. S4 uses each base station in the urban rail vehicle-to-ground wireless network as an intelligent agent to design the state space, action space and reward function of deep reinforcement learning. S5 is based on a multi-agent deep deterministic gradient policy algorithm, which obtains the base station agent radio frequency parameter adjustment strategy through centralized learning and training of the policy-value network. S6 performs adaptive optimization of urban rail vehicle-to-ground wireless network coverage quality by distributing the base station agent radio frequency parameter adjustment strategy across all base station agents.

2. The method according to claim 1, characterized in that, Step S1 specifically includes: Based on the periodically received MA information, through the formula (1) Calculate the minimum tracking interval between trains; where: It is the interval between two consecutive MA messages received by the train. The distance traveled by subsequent trains during this period, It is the safe braking distance for following trains, where safe braking distance... Depending on the emergency braking trigger speed of the train's current position And the train's trajectory after it triggers emergency braking at that point; Through (2) Calculate the time required for the following train to reach the service braking curve at the current speed. In the formula: The difference between the train's running speed and the normal braking trigger speed, This is the commonly used braking deceleration. If the interval between two consecutive MA messages received by the train... Exceed If so, the train will trigger the service brake; set up The longest transmission delay allowed by ATP, if Exceed If this occurs, the train control system will immediately apply emergency braking until the train comes to a complete stop.

3. The method according to claim 2, characterized in that, The settings described The longest transmission delay allowed by ATP, if Exceed If the train comes to a complete stop, the train control system will immediately output an emergency brake, and the process will include the following: when If the train does not trigger the service brake, then the train will proceed via a bypass system. (3) Calculate trains Speed ​​is the distance traveled before receiving the next MA message. If the train triggers the service brake, then... (4) Calculate trains The speed is the distance traveled before receiving the next MA message; when If the train does not trigger service or emergency braking, then the train's braking force is calculated using equation (3). The speed is the distance traveled before receiving the next MA message; if the train triggers emergency braking, then... (5) Calculate the minimum train tracking interval; if the train is running at the minimum tracking interval, then use the formula... (6) Calculate the time interval between the preceding and following trains passing the same point consecutively. In equation (6), To protect margin and ranging error, The length of the train; Through (7) Calculate the line capacity .

4. The method according to claim 2, characterized in that, Step S2 specifically includes: Through (8) Calculate the probability of unnecessary frequent braking caused by the train: Where: Indicates the first One cycle, This indicates the number of consecutively lost data packets. For the first Packet loss rate over a period of time; Through (9) Calculate the minimum number of consecutive packet drops that would cause unnecessary frequent braking of the train; where: For the train at the Data packet transmission latency per cycle This refers to the contract issuance cycle; Through (10) Calculate the probability of unnecessary emergency braking of the train; where: Indicates the number of consecutively lost data packets; Through (11) The calculation represents the minimum number of consecutive packet drops required to trigger an unnecessary emergency braking of the train.

5. The method according to claim 4, characterized in that, Step S3 specifically includes: Through (12) (13) Calculate the antenna gain of the antenna model; in equation (12): and These represent the horizontal and vertical angles between the user equipment and the antenna, respectively. and Represent the horizontal gain and vertical gain respectively; in equation (13): and These are the antenna azimuth and downtilt angles, respectively. and These are the horizontal beamwidth and vertical beamwidth of the antenna, respectively. Considering free-space propagation loss and a small-scale fading model following a Rice distribution, the equation is used... (14) Calculate train Received signal power; where: For base stations The transmission power, For base station antenna gain, This is power loss caused by large-scale fading. This is power loss caused by small-scale fading; Through (15) The signal-to-interference-plus-noise ratio (SINR) is used to characterize the effects of internal and external interference and noise during signal transmission; where: Indicates train Received serving cell The received signal power, denoted as the set of all base stations in the network. , The power of interference signals from other base stations within the system. Indicates noise power. This indicates interference from outside the system.

6. The method according to claim 5, characterized in that, In step S4: Each base station in the urban rail vehicle-to-ground wireless network is treated as an intelligent agent, enabling each base station intelligent agent to interact with the vehicle-to-ground wireless network environment. The state range observed by each base station agent includes the train status information within the cell covered by that base station agent and the radio frequency parameter information of that base station agent; the state space function of each base station agent is: (16) In the formula: This indicates the status information of each train in the serving cell, including the train's location, reference signal received power, signal-to-interference-plus-noise ratio, data packet transmission delay, and packet loss rate. This refers to the radio frequency parameter information of the base station; The action space of each base station agent includes its own transmit power and antenna downtilt angle adjustment; the action space function of each base station agent is: (17) In the formula: For the selected base station transmit power, The selected antenna downtilt angle has a range of values ​​within continuous space. The reward function for each base station agent is used to guide the agent in learning a collaborative radio frequency (RF) parameter adjustment strategy that balances performance and interference control. It evaluates the comprehensive impact of each agent's RF parameter configuration on communication quality, the probability of unnecessary train braking, and the rationality of parameter adjustments. The reward function for each base station agent is as follows: (18) In the formula: This indicates an improvement in the train's communication quality, which is achieved by weighting the normalized RSRP and SINR values ​​of the train at its current location. This represents the probability of unnecessary braking of the train, which is calculated by weighting the probabilities of unnecessary regular braking and emergency braking obtained from transmission delay and packet loss rate. This indicates the radio frequency parameter configuration of the base station intelligent agent. , and They represent , and The weights; N represents the number of trains.

7. The method according to claim 6, characterized in that, Step S5 specifically includes: S501 determines the interaction environment between the base station agent and the vehicle-to-ground wireless network, the number of base station agents, the total number of training rounds, and the maximum number of steps per round; S502 initializes the policy network and value network of each base station agent and its corresponding target network, and sets the experience replay buffer B. For each training round, the S503 resets the interaction environment state between the base station agent and the vehicle-to-ground wireless network, and obtains initial observations. ; Every step of S504 in every round of training Each base station agent, through a policy network, combines local observation data... Generate Actions Add exploration noise; S505 treats the set of actions generated by all base station agents as a joint action. Perform joint actions in an interactive environment; After the S506 station's intelligent agent and the vehicle-to-ground wireless network interact, they perform a joint action and return. The new set of states for each step and rewards ; S507 will be in status ,action ,award New state The tuples formed , , , Store in experience replay buffer B; S508 randomly samples a batch of training data from the experience replay buffer B, updates the value network parameters of each base station agent by minimizing the loss function, updates the policy network parameters by maximizing the expected Q value, and softly updates the target network. S509 Repeat sub-steps S502 to S508 until all rounds of training are completed; S510, based on the execution method of sub-step S506, uses the base station agent to collect local environmental status information in real time, generates radio frequency parameter adjustment strategies using the base station agent's own strategy network, and performs transmit power adjustment and antenna tilt angle adjustment actions in a distributed manner, thereby optimizing the coverage quality of the urban rail vehicle-to-ground wireless network in real time according to changes in environmental status.

Citation Information

Patent Citations

  • D2D user resource allocation method based on deep reinforcement learning algorithm and storage medium

    CN116456493A

  • Unmanned aerial vehicle trajectory optimization method and system based on bionic algorithm and BP neural network

    CN116782269A

  • Deep Q learning-based virtual network function placement method with high and low orbit fusion

    CN118945669A

  • Intelligent agile networking method and device for large emergency wireless communication network

    CN120456039A

  • System and method for deep learning and wireless network optimization using deep learning

    WO2019007388A1