An urban rail transit emergency disposal recommendation method

By constructing a multi-dimensional equipment knowledge graph and combining reinforcement learning and deep learning algorithms, the configuration and execution of emergency response plans for urban rail transit are optimized. This solves the problems of rigid emergency response methods and cumbersome configuration, and realizes a close connection between emergency plans and equipment linkage, as well as flexible emergency response.

CN115936456BActive Publication Date: 2026-05-12NANJING RAIL TRANSIT SYST
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING RAIL TRANSIT SYST
Filing Date
2022-11-18
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

The existing emergency response methods for urban rail transit suffer from problems such as rigid emergency plans, cumbersome configuration, weak connection between emergency plans and emergency response equipment, and the inability of emergency response results or emergency drill evaluation results to flexibly provide effective references for subsequent handling.

Method used

A multi-dimensional equipment knowledge graph is constructed, and Markov decision process, Monte Carlo method and RNN recurrent neural network algorithm in reinforcement learning are combined to train the emergency response plan configuration and execution recommendation strategy, and optimize the emergency response plan configuration and execution process.

Benefits of technology

Simplify the configuration of emergency response plans, enhance the connection between emergency response plans and emergency response equipment, and improve the flexibility of emergency response results and the reference value of emergency drill evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115936456B_ABST
    Figure CN115936456B_ABST
Patent Text Reader

Abstract

The application discloses a kind of urban rail transit emergency disposal recommendation methods, steps are as follows: the device data relationship information is constructed;Device spatial relationship information is constructed;The multi-dimensional device knowledge graph including device static data relationship and dynamic spatial relationship is constructed;Emergency disposal preplan configuration recommendation strategy is trained;Emergency disposal preplan configuration recommendation strategy obtained by training is used to carry out configuration recommendation;Emergency disposal preplan execution recommendation strategy is trained;Emergency disposal preplan execution recommendation strategy obtained by training is used to carry out execution recommendation.The application is simplified by emergency configuration recommendation learning emergency preplan configuration, and the connection between emergency preplan and emergency linkage device is strengthened by emergency execution recommendation, and emergency disposal result or emergency exercise evaluation result can provide effective reference for subsequent emergency disposal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of emergency response technology in the context of big data in urban rail transit, and specifically relates to a recommended method for emergency response in urban rail transit. Background Technology

[0002] After the urban rail transit network is operational, there are complex interrelationships between the lines. The negative impact of an emergency on one line will directly or indirectly affect other lines in the network. Existing emergency response methods mostly employ digital emergency plans. When an emergency occurs, a pre-determined static emergency plan is selected based on the type of event. Emergency response is completed through processes such as reporting, event classification, plan activation, instruction issuance, and assessment. However, these existing emergency response methods suffer from problems such as rigid emergency plans, cumbersome plan configuration, weak integration between plan configuration and emergency response equipment, and the inability of emergency response results or drill evaluations to flexibly provide effective references for subsequent emergency responses.

[0003] A Markov chain is a stochastic process in probability theory and mathematical statistics that possesses the Markov property and exists within a discrete exponential set and state space. A Markov chain applicable to continuous exponential sets is called a Markov process. Markov chains can be defined using transition matrices and transition graphs. Besides the Markov property, Markov chains may exhibit irreducibility, recurrence, periodicity, and ergodicity. An irreducible and recurrence-normal Markov chain is a strictly stationary Markov chain with a unique stationary distribution; the limiting distribution of an ergodic Markov chain converges to its stationary distribution.

[0004] Markov decision processes (MDFs) are mathematical models of sequential decision-making used to simulate stochastic policies and rewards achievable by an agent in an environment where the system state exhibits Markov properties. A MDF is constructed based on a set of interacting objects: the agent and the environment. Its elements include state, action, policy, and reward. In the simulation of a MDF, the agent perceives the current system state, performs actions on the environment according to the policy, thereby changing the state of the environment and receiving a reward. The accumulation of this reward over time is called the payoff. Markov chains are applied to the Monte Carlo method, forming the Markov Chain Monte Carlo method.

[0005] Recurrent Neural Networks (RNNs) are a type of deep neural network that adds temporal relationships between sequences to a fully connected neural network, enabling them to better handle time-related problems. RNNs are highly capable of processing sequential data. By sharing weights across different parts of the network model, they can be extended to various sample types and generalize to long sequences to obtain the desired results. Summary of the Invention

[0006] In view of the shortcomings of the prior art, the purpose of this invention is to provide a recommended method for emergency response in urban rail transit, so as to solve the problems that existing emergency response methods have rigid emergency plans, cumbersome emergency plan configuration, weak connection between emergency plan configuration and emergency linkage equipment, and the inability of emergency response results or emergency drill evaluation results to flexibly provide effective reference for subsequent emergency response.

[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0008] The present invention provides a recommended method for emergency response in urban rail transit, comprising the following steps:

[0009] 1) Obtain basic information on urban rail transit emergency response equipment from the urban rail transit production system and construct equipment data relationship information;

[0010] 2) Obtain basic spatial information of emergency response equipment for urban rail transit from the urban rail transit BIM system, and construct spatial relationship information of the equipment;

[0011] 3) Based on the above-mentioned equipment data relationship information and equipment spatial relationship information, construct a multi-dimensional equipment knowledge graph that includes static equipment data relationships and dynamic spatial relationships;

[0012] 4) Construct an environment for recommending the configuration of emergency response plans for urban rail transit, and train the emergency response plan configuration recommendation strategy by combining Markov decision process, Monte Carlo method and RNN recurrent neural network algorithm in reinforcement learning.

[0013] 5) Use the emergency response plan configuration recommendation strategy obtained from step 4) to make configuration recommendations;

[0014] 6) Construct an intelligent agent atagt, an environment atenv, an execution strategy atply, and a reward / penalty function atrwd for the execution of urban rail transit emergency response plans. Train the recommended strategy for the execution of emergency response plans by combining Markov decision process, Monte Carlo method and RNN recurrent neural network algorithm in reinforcement learning.

[0015] 7) Utilize the recommended strategies for implementing the emergency response plan obtained from step 6) to make implementation recommendations.

[0016] Further, step 1) specifically includes:

[0017] 11) Determine the equipment type, protocol format, and parsing format of the integrated monitoring system and the integrated automation system of the substations for each station, depot, and parking lot of the urban rail transit line;

[0018] 12) Query the required electromechanical equipment data relationship information for environmental and equipment monitoring systems, power monitoring systems, communication management units, automatic train monitoring systems, fire alarm systems, and automatic fare collection systems by equipment type;

[0019] 13) Obtain and parse electromechanical equipment data relationship information according to the protocol format, including equipment tags, equipment types, analog point information, digital point information, limits, alarm information, error information, subordinate relationships, sequence relationships, and coloring relationships;

[0020] 14) Obtain and parse the basic spatial relationship information of the devices according to the protocol format;

[0021] 15) Construct the equipment model structure, add equipment data relationship information, and store it in the urban rail transit emergency response database.

[0022] Furthermore, step 2) specifically includes:

[0023] 21) Select basic equipment information from the urban rail transit emergency response database;

[0024] 22) Query the equipment spatial model from the urban rail transit BIM system based on the basic equipment information;

[0025] 23) Obtain the specific information of the equipment space model by parsing the BIM protocol of the equipment space model;

[0026] 24) Obtain the detailed spatial information of the required equipment from the specific information and proceed to step 25); If the detailed spatial information of the required equipment does not exist in the urban rail transit BIM system, manually enter the detailed spatial information of the equipment and proceed to step 25);

[0027] 25) Reconstruct the equipment model structure, add detailed spatial information of the equipment, form an equipment relationship information model and an equipment spatial information model, and store them in the urban rail transit emergency response database.

[0028] Furthermore, step 3) specifically includes:

[0029] 31) Query the equipment relationship information model and equipment spatial information model from the urban rail transit emergency response database;

[0030] 32) Based on the equipment relationship information model, construct a static knowledge graph of equipment with the reconstructed equipment model as the points and the relationships between equipment as the edges;

[0031] 33) Add the static data of the equipment, including but not limited to equipment subordination, equipment grouping, equipment sequence, equipment coloring, and equipment control relationships, to the equipment static knowledge graph.

[0032] 34) Based on the N-dimensional spatial information of equipment stored in the urban rail transit emergency response database, the distance between each piece of equipment in a local spatial area is calculated using the spatial region segmentation method and Minkowski distance, and the dynamic spatial relationship of the equipment is constructed.

[0033]

[0034] Where p and q are devices, ρ is a positive integer, G is the region segmented by the spatial region segmentation method, and h is a parameter variable bound to region G, which is optimized and updated through backpropagation based on the recommendation results;

[0035] 35) Based on static data relationships and dynamic spatial relationships, construct a multi-dimensional device knowledge graph that includes static data relationships and dynamic spatial relationships of devices;

[0036] 36) The constructed multi-dimensional equipment knowledge graph, which includes static data relationships and dynamic spatial relationships of equipment, will be stored in the urban rail transit emergency response knowledge graph database.

[0037] Furthermore, step 4) specifically includes:

[0038] 41) Construct an intelligent agent cfagt for configuring emergency response plans for urban rail transit, and perform emergency response plan configuration;

[0039] 42) Construct the environment cfenv and environment response function cfrpfn. Calculate based on the spatially segmented region and the dynamic spatial relationship D(p,q) between devices. When two devices do not belong to the same region, the environment response function cfrpfn returns a weakened response; when D(p,q)>h, it returns no response; when D(p,q)≤h, it returns an increased response.

[0040] 43) Construct a reward and punishment function cfrwd, and based on the historically accumulated emergency pre-configuration recommendation sequence cfrdrs{rs1,rs2,rs3…,rsn}, user selection results cfchrs, and environmental response function cfrpfn, provide value feedback on the ACT actions (including turning on devices, turning off devices, broadcasting, and real-time video) of the agent cfagt; reward and punishment function When cfchrs is not in cfrdrs, n = 1; otherwise, n = 2. result(cfchrs) represents the selection result, factor represents the fitting factor, which is a decimal greater than 0 and less than or equal to 1; E represents the expected value calculated from historical accumulated data.

[0041] 44) Utilize the RNN (Recurrent Neural Network) algorithm to configure the strategy cfply, where t represents time and x... t s represents the emergency response data at time t. tThis represents the output of the configuration strategy cfply at time t, y t This indicates the predicted emergency response output at time t. The input to the cfply configuration strategy has two sources: one is the current x. t The input is one of the two: the output of the previous state's hidden layer configuration policy cfply. t-1 W, U, and V are parameters; s t =tanh(Ux t +Ws t-1 );y t =softmax(Vs t Based on the forward and backward propagation algorithms of RNN, W, U, and V are obtained as parameters.

[0042] 45) The agent cfagt selects a historically accumulated emergency pre-configuration recommendation and selects a configuration ACT action to execute according to the configuration strategy cfply;

[0043] 46) The environment cfenv responds to the received configuration ACT action through the environment response function cfrpfn, and the reward and penalty value is calculated by the reward and penalty function cfrwd;

[0044] 47) Input the calculated reward and penalty values ​​into the configuration policy cfply, and update the parameters of the configuration policy cfply according to the calculation method of the configuration policy cfply;

[0045] 48) Using the Monte Carlo method, random sampling is used to select and repeat steps 45)-47) to train and obtain the emergency pre-configuration recommendation strategy.

[0046] Furthermore, step 5) specifically includes:

[0047] 51) Based on the user's selection criteria, configure the agent cfagt to select one ACT action from the user-configured action sequence;

[0048] 52) Configure the agent cfagt to configure the selected ACT action through the configuration policy cfply;

[0049] 53) Configure the reward / penalty function cfrwd to calculate the reward / penalty value of the selected ACT action and record it;

[0050] 54) Repeat steps 51)-53) to obtain the reward / penalty value for each ACT action in the configured action sequence, and sort them in descending order;

[0051] 55) Based on the set number, return the set number of ACT actions to the user in sequence;

[0052] 56) The user selects an action from the returned sequence of actions as the configuration result;

[0053] 57) Record the steps of the contingency plan, the recommended operation for each step, and the configuration results, and store all the plan configuration information in the urban rail transit emergency response database.

[0054] Furthermore, step 6) specifically includes:

[0055] 61) Construct an intelligent agent, atag, to execute the emergency response plan for urban rail transit;

[0056] 62) Construct the environment atenv and the environment response function atrpfn. Calculate the response based on the spatially segmented region and the dynamic spatial relationship D(p,q) between devices. When two devices do not belong to the same region, atrpfn returns a weakening response; when D(p,q)>h, it returns no response; when D(p,q)≤h, it returns an increasing response.

[0057] 63) Construct a reward and punishment function `atrwd`, and based on the historically accumulated emergency pre-configuration recommendation sequence `atrdrs{rs1,rs2,rs3…,rsn}`, user selection results `atchrs`, and the environmental response function `atrpfn`, provide value feedback to the ACT actions of the agent `atagt`; reward and punishment function If `atchrs` is not in `atrdrs`, then `n` = 1; otherwise, `n` = 2. `result(atchrs)` represents the selection result, `factor` represents the fitting factor, which is a decimal greater than 0 and less than or equal to 1, and `E` represents the expected value calculated from historical accumulated data.

[0058] 64) Execute the strategy atply using the RNN (Recurrent Neural Network) algorithm, where t represents time x t s represents the emergency response data at time t. t This represents the output of the configuration policy atply at time t, y t This indicates that the output of the emergency response is predicted at time t. The input to the execution strategy atply has two sources: one is the current x. t The input is the previous hidden layer execution policy atply, and the other is the output s. t-1 W, U, and V are parameters; s t =tanh(Ux t +Ws t-1 );y t =softmax(Vs t Based on the forward and backward propagation algorithms of RNN, W, U, and V are obtained as parameters.

[0059] 65) The agent atat selects a historically accumulated emergency pre-set execution recommendation and selects the ACT action to execute according to the execution strategy atply;

[0060] 66) The environment atenv responds to the received configuration ACT action through the environment response function atrpfn, and the reward and penalty value is calculated by the reward and penalty function atrwd;

[0061] 67) Input the calculated reward and penalty values ​​into the execution policy atply, and update the parameters of the execution policy atply according to the calculation method of the execution policy atply;

[0062] 68) Using the Monte Carlo method, random sampling is used to select repeated steps 65)-67) to train and obtain the emergency pre-set execution recommendation strategy.

[0063] Further, step 7) specifically includes:

[0064] 71) Execute the intelligent agent atat, and based on the configuration results and the type and location of the actual emergency, preliminarily screen the emergency response equipment sequence according to the dynamic spatial relationship;

[0065] 72) The execution agent atat selects one of the ACT actions from the configured sequence of execution actions; if the ACT action involves emergency linkage equipment, proceed to step 73); if it does not involve specific linkage equipment, calculate the reward or penalty value of the ACT action by executing the reward or penalty function atrwd.

[0066] 73) Select one device from the emergency response equipment sequence;

[0067] 74) The reward / penalty function atrwd calculates and records the reward / penalty values ​​for the selected ACT action and device;

[0068] 75) If the selected ACT action remains unchanged, select the next device in the emergency response device sequence, calculate the reward or penalty value by executing the reward or penalty function atrwd, and record it;

[0069] 76) Based on the set ACT action type, select to calculate the ACT action reward and penalty value by weighted average, expected value, or partial accumulation method.

[0070] 77) Repeat step 72) to obtain the reward / penalty value for each ACT action in the action sequence, and sort them in descending order;

[0071] 79) Based on the set number, return the recommended actions to the user in sequence according to the set number;

[0072] 710) Select one of the recommended actions to execute, and obtain user ratings after the plan is executed;

[0073] 711) Record the steps of the contingency plan, the recommended operation for each step, the execution results and scores, and store all execution information in the urban rail transit emergency response database.

[0074] The beneficial effects of this invention are:

[0075] This invention realizes a multi-dimensional device knowledge graph based on spatial region segmentation and Minkowski distance, which represents the static data relationship and dynamic spatial relationship of devices. It applies Markov decision process, Monte Carlo method and recurrent neural network algorithm in reinforcement learning, simplifies emergency plan configuration through emergency configuration recommendation learning, strengthens the connection between emergency plan and emergency linkage equipment through emergency execution recommendation, and provides effective reference for subsequent emergency response results or emergency drill evaluation results. Attached Figure Description

[0076] Figure 1 This is a flowchart illustrating the principle of the method of the present invention.

[0077] Figure 2 A training and learning flowchart recommended for the emergency response configuration of this invention.

[0078] Figure 3 This diagram illustrates the cfply learning and updating process for the configuration strategy of this invention. Detailed Implementation

[0079] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to embodiments and accompanying drawings. The content mentioned in the embodiments is not intended to limit the present invention.

[0080] Reference Figures 1-2 As shown, the recommended method for emergency response in urban rail transit according to the present invention comprises the following steps:

[0081] 1) Obtain basic information on urban rail transit emergency response equipment from the urban rail transit production system, and construct equipment data relationship information; specifically including:

[0082] 11) Determine the equipment type (analog, digital, custom, etc.), protocol format (MODBUS, IEC104, HTTP, etc.), and parsing format (JSON, XML, binary format, etc.) of the integrated monitoring system of each station, depot, parking lot, and substation of the urban rail transit line.

[0083] 12) Query the required electromechanical equipment data relationship information for the following systems by equipment type: Environmental and Equipment Monitoring System (BAS), Power Monitoring System (PSCADA), Communication System (TEL), Automatic Train Control System (ATS), Fire Alarm System (FAS), and Automatic Fare Collection System (AFC);

[0084] 13) Obtain and parse electromechanical equipment data relationship information according to the protocol format, including equipment tags, equipment types, analog point information, digital point information, limits, alarm information, error information, subordinate relationships, sequence relationships, and coloring relationships;

[0085] 14) Obtain and parse the basic spatial relationship information (basic spatial coordinates of the device) according to the protocol format;

[0086] 15) Construct the equipment model structure, add equipment data relationship information, and store it in the Urban Rail Transit Emergency Response Database (ULDB).

[0087] 2) Obtain basic spatial information of urban rail transit emergency response equipment from the urban rail transit BIM system, and construct spatial relationship information of the equipment; specifically including:

[0088] 21) Select basic equipment information from the urban rail transit emergency response database;

[0089] 22) Query the equipment spatial model from the urban rail transit BIM system based on the basic equipment information;

[0090] 23) Obtain the specific information of the equipment space model by parsing the BIM protocol of the equipment space model;

[0091] 24) Obtain the detailed spatial information of the required equipment from the specific information (detailed equipment spatial coordinates xyz, equipment outer enclosure geometric size lwh, etc.), and proceed to step 25); if the detailed spatial information of the required equipment does not exist in the urban rail transit BIM system, manually input the detailed spatial information of the equipment and proceed to step 25);

[0092] 25) Reconstruct the equipment model structure, add detailed spatial information of the equipment, form an equipment relationship information model (equipment label, equipment type, analog point information, digital point information, limit value, alarm information, error information, subordinate information, sequence information, coloring information, etc.) and an equipment spatial information model (equipment basic spatial coordinates, equipment spatial size, equipment shape, equipment outer size, etc.), and store them in the urban rail transit emergency response database.

[0093] 3) Based on the aforementioned equipment data relationship information and equipment spatial relationship information, construct a multi-dimensional equipment knowledge graph that includes static equipment data relationships and dynamic spatial relationships; specifically including:

[0094] 31) Query the equipment relationship information model and equipment spatial information model from the urban rail transit emergency response database;

[0095] 32) Based on the equipment relationship information model, construct a static knowledge graph of equipment with the reconstructed equipment model (including equipment relationship information and equipment spatial information) as points and the relationships between equipment as edges;

[0096] 33) Add the static data of the equipment, including but not limited to equipment subordination, equipment grouping, equipment sequence, equipment coloring, and equipment control relationships, to the equipment static knowledge graph.

[0097] 34) Based on the N-dimensional spatial information of equipment stored in the urban rail transit emergency response database, the distance between each piece of equipment in a local spatial area is calculated using the spatial region segmentation method and Minkowski distance, and the dynamic spatial relationship of the equipment is constructed.

[0098]

[0099] Where p and q are devices, ρ is a positive integer (such as 2, 3, 4, etc.), G is the region segmented by the spatial region segmentation method, and h is a parameter variable bound to the region G, which is optimized and updated through backpropagation based on the recommendation results;

[0100] 35) Based on static data relationships and dynamic spatial relationships, construct a multi-dimensional device knowledge graph that includes static data relationships and dynamic spatial relationships of devices;

[0101] 36) The constructed multi-dimensional equipment knowledge graph, which includes static data relationships and dynamic spatial relationships of equipment, will be stored in the Urban Rail Transit Emergency Response Knowledge Graph Database (ISKG).

[0102] 4) Construct an environment for recommending emergency response plans for urban rail transit, and train the emergency response plan recommendation strategy using a combination of Markov decision processes, Monte Carlo methods, and RNN recurrent neural network algorithms in reinforcement learning; specifically including:

[0103] 41) Construct an intelligent agent cfagt for configuring emergency response plans for urban rail transit, and perform emergency response plan configuration;

[0104] 42) Construct the environment cfenv and environment response function cfrpfn. Calculate based on the spatially segmented region and the dynamic spatial relationship D(p,q) between devices. When two devices do not belong to the same region, the environment response function cfrpfn returns a weakened response; when D(p,q)>h, it returns no response; when D(p,q)≤h, it returns an increased response.

[0105] 43) Construct a reward and punishment function cfrwd, and based on the historically accumulated emergency pre-configuration recommendation sequence cfrdrs{rs1,rs2,rs3…,rsn}, user selection results cfchrs, and environmental response function cfrpfn, provide value feedback on the ACT actions (including turning on devices, turning off devices, broadcasting, and real-time video) of the agent cfagt; reward and punishment function When cfchrs is not in cfrdrs, n = 1; otherwise, n = 2. result(cfchrs) represents the selection result, factor represents the fitting factor, which is a decimal greater than 0 and less than or equal to 1; E represents the expected value calculated from historical accumulated data.

[0106] 44) Utilize the RNN (Recurrent Neural Network) algorithm to configure the strategy cfply, where t represents time and x... t s represents the emergency response data at time t. t This represents the output of the configuration strategy cfply at time t, y t This indicates the predicted emergency response output at time t. The input to the cfply configuration strategy has two sources: one is the current x. t The input is one of the two: the output of the previous state's hidden layer configuration policy cfply. t-1 W, U, and V are parameters; s t =tanh(Ux t +Ws t-1 );y t =softmax(Vs t Based on the forward and backward propagation algorithms of RNNs, W, U, and V are obtained as parameters; refer to Figure 3 As shown;

[0107] 45) The agent cfagt selects a historically accumulated emergency pre-configuration recommendation and selects a configuration ACT action to execute according to the configuration strategy cfply;

[0108] 46) The environment cfenv responds to the received configuration ACT action through the environment response function cfrpfn, and the reward and penalty value is calculated by the reward and penalty function cfrwd;

[0109] 47) Input the calculated reward and penalty values ​​into the configuration policy cfply, and update the parameters of the configuration policy cfply according to the calculation method of the configuration policy cfply;

[0110] 48) Using the Monte Carlo method, random sampling is used to select and repeat steps 45)-47) to train and obtain the emergency pre-configuration recommendation strategy.

[0111] 5) Utilize the emergency response plan configuration recommendation strategy obtained from step 4) for configuration recommendations; specifically including:

[0112] 51) Based on the user's selected conditions (including emergency scenario selection, such as tunnel fire, tunnel flooding, train derailment, train malfunction, passenger stampede, etc.), configure the intelligent agent cfagt to select one ACT action from the user-configured action sequence;

[0113] 52) Configure the agent cfagt to configure the selected ACT action through the configuration policy cfply;

[0114] 53) Configure the reward / penalty function cfrwd to calculate the reward / penalty value of the selected ACT action and record it;

[0115] 54) Repeat steps 51)-53) to obtain the reward / penalty value for each ACT action in the configured action sequence, and sort them in descending order;

[0116] 55) Based on the set number, return the set number of ACT actions to the user in sequence;

[0117] 56) The user selects an action from the returned sequence of actions as the configuration result;

[0118] 57) Record the steps of the contingency plan, the recommended operation for each step, and the configuration results, and store all the plan configuration information in the urban rail transit emergency response database.

[0119] 6) Construct an intelligent agent (atagt), environment (atenv), execution policy (atply), and reward / penalty function (atrwd) for the execution of urban rail transit emergency response plans. Train the recommended strategy for emergency response plan execution using a combination of Markov decision processes, Monte Carlo methods, and recurrent neural network algorithms from reinforcement learning; specifically including:

[0120] 61) Construct an intelligent agent, atag, to execute the emergency response plan for urban rail transit;

[0121] 62) Construct the environment atenv and the environment response function atrpfn. Calculate the response based on the spatially segmented region and the dynamic spatial relationship D(p,q) between devices. When two devices do not belong to the same region, atrpfn returns a weakening response; when D(p,q)>h, it returns no response; when D(p,q)≤h, it returns an increasing response.

[0122] 63) Construct a reward and punishment function `atrwd`, and based on the historically accumulated emergency pre-configuration recommendation sequence `atrdrs{rs1,rs2,rs3…,rsn}`, user selection results `atchrs`, and the environmental response function `atrpfn`, provide value feedback to the ACT actions of the agent `atagt`; reward and punishment function If `atchrs` is not in `atrdrs`, then `n` = 1; otherwise, `n` = 2. `result(atchrs)` represents the selection result, `factor` represents the fitting factor, which is a decimal greater than 0 and less than or equal to 1, and `E` represents the expected value calculated from historical accumulated data.

[0123] 64) Execute the strategy atply using the RNN (Recurrent Neural Network) algorithm, where t represents time x t s represents the emergency response data at time t. t This represents the output of the configuration policy atply at time t, y t This indicates that the output of the emergency response is predicted at time t. The input to the execution strategy atply has two sources: one is the current x. t The input is the previous hidden layer execution policy atply, and the other is the output s. t-1 W, U, and V are parameters; s t =tanh(Ux t +Ws t-1 );y t =softmax(Vs t Based on the forward and backward propagation algorithms of RNN, W, U, and V are obtained as parameters.

[0124] 65) The agent atat selects a historically accumulated emergency pre-set execution recommendation and selects the ACT action to execute according to the execution strategy atply;

[0125] 66) The environment atenv responds to the received configuration ACT action through the environment response function atrpfn, and the reward and penalty value is calculated by the reward and penalty function atrwd;

[0126] 67) Input the calculated reward and penalty values ​​into the execution policy atply, and update the parameters of the execution policy atply according to the calculation method of the execution policy atply;

[0127] 68) Using the Monte Carlo method, random sampling is used to select repeated steps 65)-67) to train and obtain the emergency pre-set execution recommendation strategy.

[0128] 7) Utilize the recommended strategies for implementing the emergency response plan obtained from step 6) to make implementation recommendations; specifically including:

[0129] 71) Execute the intelligent agent atat, and based on the configuration results and the type and location of the actual emergency, preliminarily screen the emergency response equipment sequence according to the dynamic spatial relationship;

[0130] 72) The execution agent atat selects one of the ACT actions from the configured sequence of execution actions; if the ACT action involves emergency linkage equipment, proceed to step 73); if it does not involve specific linkage equipment, calculate the reward or penalty value of the ACT action by executing the reward or penalty function atrwd.

[0131] 73) Select one device from the emergency response equipment sequence;

[0132] 74) The reward / penalty function atrwd calculates and records the reward / penalty values ​​for the selected ACT action and device;

[0133] 75) If the selected ACT action remains unchanged, select the next device in the emergency response device sequence, calculate the reward or penalty value by executing the reward or penalty function atrwd, and record it;

[0134] 76) Based on the set ACT action type, select to calculate the ACT action reward and penalty value by weighted average, expected value, or partial accumulation method.

[0135] 77) Repeat step 72) to obtain the reward / penalty value for each ACT action in the action sequence, and sort them in descending order;

[0136] 79) Based on the set number, return the recommended actions to the user in sequence according to the set number;

[0137] 710) Select one of the recommended actions to execute, and obtain user ratings after the plan is executed;

[0138] 711) Record the steps of the contingency plan, the recommended operation for each step, the execution results and scores, and store all execution information in the urban rail transit emergency response database.

[0139] This invention has many specific applications. The above description is only a preferred embodiment of this invention. It should be noted that for those skilled in the art, several improvements can be made without departing from the principle of this invention, and these improvements should also be considered within the scope of protection of this invention.

Claims

1. A recommended method for emergency response in urban rail transit, characterized in that, The steps are as follows: 1) Obtain basic information on urban rail transit emergency response equipment from the urban rail transit production system and construct equipment data relationship information; 2) Obtain basic spatial information of emergency response equipment for urban rail transit from the urban rail transit BIM system, and construct spatial relationship information of the equipment; 3) Based on the above-mentioned equipment data relationship information and equipment spatial relationship information, construct a multi-dimensional equipment knowledge graph that includes static equipment data relationships and dynamic spatial relationships; 4) Construct an environment for recommending emergency response plans for urban rail transit, and train the emergency response plan recommendation strategy by combining Markov decision process, Monte Carlo method and RNN recurrent neural network algorithm in reinforcement learning. 5) Utilize the emergency response plan configuration recommendation strategy obtained from step 4) to make configuration recommendations; 6) Construct an intelligent agent atagt, an environment atenv, an execution strategy atply, and a reward / penalty function atrwd for the execution of urban rail transit emergency response plans. Train the recommended strategy for the execution of emergency response plans by combining Markov decision process, Monte Carlo method and RNN recurrent neural network algorithm in reinforcement learning. 7) Implement the recommended strategies for the emergency response plan obtained from step 6); Step 3) specifically includes: 31) Query the equipment relationship information model and equipment spatial information model from the urban rail transit emergency response database; 32) Based on the equipment relationship information model, construct a static knowledge graph of equipment with the reconstructed equipment model as the nodes and the relationships between equipment as the edges; 33) Add the static data of the equipment, including but not limited to equipment subordination, equipment grouping, equipment sequence, equipment coloring, and equipment control relationships, to the equipment static knowledge graph. 34) Based on the N-dimensional spatial information of equipment stored in the urban rail transit emergency response database, the spatial region segmentation method and Minkowski distance are used to calculate the distance between each piece of equipment in a local spatial region, and to construct the dynamic spatial relationship of the equipment; ; Where p and q are devices. Let G be a positive integer, where G is the region segmented by the spatial region segmentation method, and h is a parameter variable bound to region G. It is optimized and updated through backpropagation based on the recommendation results. 35) Based on static data relationships and dynamic spatial relationships, construct a multi-dimensional device knowledge graph that includes static data relationships and dynamic spatial relationships of devices; 36) The constructed multi-dimensional equipment knowledge graph, which includes static data relationships and dynamic spatial relationships of equipment, will be stored in the urban rail transit emergency response knowledge graph database.

2. The recommended method for emergency response in urban rail transit according to claim 1, characterized in that, Step 1) specifically includes: 11) Determine the equipment type, protocol format, and parsing format of the integrated monitoring system and the integrated automation system of the substations for each station, depot, and parking lot of the urban rail transit line; 12) Query the required electromechanical equipment data relationship information for environmental and equipment monitoring systems, power monitoring systems, communication systems, train automatic monitoring systems, fire alarm systems, and automatic fare collection systems by equipment type; 13) Obtain and parse electromechanical equipment data relationship information according to the protocol format, including equipment tags, equipment types, analog point information, digital point information, limits, alarm information, error information, subordinate relationships, sequence relationships, and coloring relationships; 14) Obtain and parse the basic spatial relationship information of the devices according to the protocol format; 15) Construct the equipment model structure, add equipment data relationship information, and store it in the urban rail transit emergency response database.

3. The recommended method for emergency response in urban rail transit according to claim 2, characterized in that, Step 2) specifically includes: 21) Select basic equipment information from the urban rail transit emergency response database; 22) Query the equipment spatial model from the urban rail transit BIM system based on the basic equipment information; 23) Obtain specific information about the equipment space model by parsing the BIM protocol of the equipment space model; 24) Obtain the detailed spatial information of the required equipment from the specific information and proceed to step 25); If the detailed spatial information of the required equipment does not exist in the urban rail transit BIM system, manually enter the detailed spatial information of the equipment and proceed to step 25); 25) Reconstruct the equipment model structure, add detailed spatial information of the equipment, form an equipment relationship information model and an equipment spatial information model, and store them in the urban rail transit emergency response database.

4. The recommended method for emergency response in urban rail transit according to claim 3, characterized in that, Step 4) specifically includes: 41) Construct an intelligent agent cfagt for configuring emergency response plans for urban rail transit, and perform emergency response plan configuration; 42) Construct the environment cfenv and environment response function cfrpfn, based on the spatially segmented regions and the dynamic spatial relationships of the devices. Calculations are performed; when the two devices do not belong to the same area, the environmental response function cfrpfn returns a weakened response; when... When, the return value remains unchanged; when When the time comes, return to increase; 43) Construct a reward / penalty function cfrwd, and based on the historically accumulated emergency pre-configuration recommendation sequence cfrdrs{rs1,rs2,rs3…,rsn}, user selection results cfchrs, and environmental response function cfrpfn, provide value feedback to the agent cfagt's ACT actions; reward / penalty function When cfchrs is not in cfrdrs, n=1; otherwise, n=2. result(cfchrs) represents the selection result, factor represents the fitting factor, which is a decimal greater than 0 and less than or equal to 1; E represents the expected value calculated from historical accumulated data. 44) Use the RNN (Recurrent Neural Network) algorithm to configure the strategy cfply, where t represents time. This represents the emergency response data at time t. This represents the output of the configuration policy cfply at time t. This indicates the predicted emergency response output at time t. The input to the cfply configuration strategy has two sources: one is the current... The other is the output of the previous state's hidden layer configuration strategy cfply. W, U, and V are parameters; ; Based on the forward and backward propagation algorithms of RNN, W, U, and V are obtained as parameters. 45) The agent cfagt selects a historically accumulated emergency pre-configuration recommendation and selects a configuration ACT action to execute according to the configuration strategy cfply; 46) The environment cfenv responds to the received configuration ACT action through the environment response function cfrpfn, and the reward and penalty value is calculated by the reward and penalty function cfrwd; 47) Input the calculated reward and penalty values ​​into the configuration policy cfply, and update the parameters of the configuration policy cfply according to the calculation method of the configuration policy cfply; 48) Using the Monte Carlo method, random sampling is used to select and repeat steps 45)-47) to train and obtain the emergency pre-configuration recommendation strategy.

5. The recommended method for emergency response in urban rail transit according to claim 4, characterized in that, Step 5) specifically includes: 51) Based on the user's selection criteria, configure the agent cfagt to select one ACT action from the user-configured action sequence; 52) Configure the agent cfagt to configure the selected ACT action through the configuration policy cfply; 53) Configure the reward / penalty function cfrwd to calculate the reward / penalty value of the selected ACT action and record it; 54) Repeat steps 51)-53) to obtain the reward / penalty value for each ACT action in the configured action sequence, and sort them in descending order; 55) Based on the set number, return the set number of ACT actions to the user in sequence; 56) The user selects an action from the returned sequence of actions as the configuration result; 57) Record the steps of the contingency plan, the recommended operation for each step, and the configuration results, and store all the plan configuration information in the urban rail transit emergency response database.

6. The recommended method for emergency response in urban rail transit according to claim 5, characterized in that, Step 6) specifically includes: 61) Construct an intelligent agent, atag, to execute the emergency response plan for urban rail transit; 62) Construct the environment atenv and environment response function atrpfn, based on the spatially segmented regions and the dynamic spatial relationships of the devices. When performing calculations, atrpfn returns a weakened value when the two devices do not belong to the same region; when... When, the return value remains unchanged; when When the time comes, return to increase; 63) Construct a reward / penalty function `atrwd`, and based on the historically accumulated emergency pre-configuration recommendation sequence `atrdrs{rs1,rs2,rs3…,rsn}`, user selection results `atchrs`, and the environmental response function `atrpfn`, provide value feedback to the ACT actions of the agent `atagt`; reward / penalty function When `atchrs` is not in `atrdrs`, `n` = 1; otherwise, `n` = 2. `result(atchrs)` represents the selection result, `factor` represents the fitting factor, which is a decimal greater than 0 and less than or equal to 1, and `E` represents the expected value calculated from historical accumulated data. 64) The strategy atply is executed using the RNN (Recurrent Neural Network) algorithm, where t represents time. This represents the emergency response data at time t. This represents the output of the configuration policy atply at time t. This indicates that the output of the emergency response is predicted at time t. The input to the execution strategy atply has two sources: one is the current... The other input is the output of the previous hidden layer execution policy atply. W, U, and V are parameters; ; Based on the forward and backward propagation algorithms of RNN, W, U, and V are obtained as parameters. 65) The agent atat selects a historically accumulated emergency pre-set execution recommendation and selects the ACT action to execute according to the execution strategy atply; 66) The environment atenv responds to the received configuration ACT action through the environment response function atrpfn, and the reward and penalty value is calculated by the reward and penalty function atrwd; 67) Input the calculated reward and penalty values ​​into the execution policy atply, and update the parameters of the execution policy atply according to the calculation method of the execution policy atply; 68) Using the Monte Carlo method, random sampling is used to select and repeat steps 65)-67) to train and obtain the emergency pre-set execution recommendation strategy.

7. The recommended method for emergency response in urban rail transit according to claim 6, characterized in that, Step 7) specifically includes: 71) Execute the intelligent agent atat, and based on the configuration results and the type and location of the actual emergency, preliminarily screen the emergency response equipment sequence according to the dynamic spatial relationship; 72) The execution agent atat selects one of the ACT actions from the configured sequence of execution actions; if the ACT action involves emergency linkage equipment, proceed to step 73); if it does not involve specific linkage equipment, calculate the reward or penalty value of the ACT action by executing the reward or penalty function atrwd; 73) Select one device from the emergency response equipment sequence; 74) The reward / penalty function atrwd calculates and records the reward / penalty values ​​for the selected ACT action and device; 75) If the selected ACT action remains unchanged, select the next device in the emergency response device sequence, calculate the reward or penalty value by executing the reward or penalty function atrwd, and record it; 76) Based on the set ACT action type, select to calculate the ACT action reward and penalty value by weighted average, expected value, or partial accumulation method. 77) Repeat step 72) to obtain the reward / penalty value for each ACT action in the action sequence, and sort them in descending order; 79) Based on the set number, return the recommended actions to the user in sequence according to the set number; 710) Select one of the recommended actions to execute, and obtain user ratings after the plan is executed; 711) Record the steps of the contingency plan, the recommended operations for each step, the execution results and scores, and store all execution information in the urban rail transit emergency response database.