Multi-agent dynamic nuclear emergency evacuation path planning method and system based on deep reinforcement learning
By employing deep reinforcement learning and multi-agent collaborative mechanisms, the problems of dynamic adaptability and multi-agent collaborative optimization in nuclear accident emergency evacuation were solved, enabling efficient and safe evacuation route planning under nuclear accident conditions, and improving the timeliness of emergency response and overall evacuation efficiency.
Patent Information
- Application Number
- CN202511402151.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-28
- Publication Date
- 2026-01-16
AI Technical Summary
Existing technologies are unable to achieve dynamic adaptability and multi-entity collaborative optimization in nuclear accident emergency evacuation, cannot effectively handle complex and ever-changing accident environments, and lack unified digital model support for pollution diffusion and road network status, resulting in untimely and insufficient optimization of evacuation routes.
A multi-agent dynamic kernel emergency evacuation path planning method based on deep reinforcement learning is adopted. By integrating multi-agent collaborative mechanism and path game modeling, deep neural network is used for real-time decision-making. Combined with pollution diffusion model and traffic flow simulation, efficient and safe evacuation path optimization is achieved.
It enables the automatic generation of optimal evacuation route plans in the dynamic environment of nuclear accidents, improves the intelligence level and practical effectiveness of emergency evacuation decision-making, ensures rapid response to changes in the route planning and environment, reduces congestion on peak road sections, and improves overall traffic efficiency and personnel safety.
Smart Images

Figure CN121352167A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of nuclear emergency response and intelligent route planning technology, and in particular to a multi-agent dynamic nuclear emergency evacuation route planning method and system based on deep reinforcement learning, which aims to achieve intelligent and collaborative dynamic evacuation route optimization for personnel and vehicles in nuclear accident scenarios. Background Technology
[0002] Following a nuclear accident, the safe evacuation of personnel from the affected area in the shortest possible time is one of the core issues of nuclear emergency response. Unlike general disasters, nuclear accidents are often accompanied by the release and spread of radioactive materials, and their hazards are characterized by their invisibility and dynamic changes: radiation contamination clouds spread and migrate over time, posing a continuous threat to personnel protection. Therefore, nuclear emergency evacuation route planning needs to simultaneously consider time sensitivity (rapid evacuation to reduce dose) and spatial dynamics (avoiding high-radiation areas), and organize the safe and orderly evacuation of large numbers of personnel and vehicles under large-scale real road network conditions. This problem involves many complex factors, including: real-time changes in the radiation field, multi-objective decision-making trade-offs (time vs. dose), multi-stakeholder interactions (competition for road resources), and uncertainties (road damage, weather changes, etc.). Existing technologies in this field have many shortcomings:
[0003] (1) Lack of dynamic radiation field coupling in path optimization: Traditional emergency evacuation path planning often adopts static or quasi-static methods, pre-calculating one or more alternative routes based on a fixed scenario. In the case of a nuclear accident, the range and intensity of radiation pollution change in real time with meteorological conditions, making it difficult for static path schemes to avoid newly emerging high-dose areas in a timely manner. For example, some studies have proposed to achieve dynamic optimization of evacuation paths by coupling high-resolution meteorological fields, radioactive diffusion models, and the A* algorithm for nuclear accident scenarios, but this is mainly used for post-accident risk assessment. The existing assessment system still has a significant lack of quantitative analysis of the dynamic interaction between emergency evacuation path optimization and the radiation field. This results in evacuation paths under traditional methods often failing to fully reflect the development of the accident and deviating from the true optimal path.
[0004] (2) Lack of multi-agent coordination mechanism: In large-scale evacuations, tens of thousands of people and vehicles often move simultaneously. If everyone chooses the shortest path to evacuate independently, traffic congestion is likely to occur on main roads, which will reduce the overall evacuation efficiency. In existing technologies, some urban emergency evacuation simulation systems have introduced multi-agent modeling to simulate the behavior of people and vehicles (e.g., abstracting people as agents for simulation exercises). However, these systems mainly focus on behavior simulation and exercise effects, lacking decision-making algorithms for path coordination optimization. As a result, the road occupancy competition problem during the evacuation process is not effectively solved, and there is a lack of necessary coordination among the various evacuation entities, which may lead to some routes being over-congested while other routes are underutilized.
[0005] (3) Lack of intelligent adaptation in path planning algorithms: Traditional path planning often employs heuristic or optimization algorithms, such as Dijkstra's algorithm, A* algorithm, and ant colony algorithm. For example, Zhou Huaifang et al. used a hybrid ant colony algorithm to optimize emergency vehicle routes under nuclear accidents, using cumulative radiation dose as an indicator to plan vehicle evacuation routes. These algorithms usually require manual setting of weights or offline calculation to determine the scheme, making it difficult to adapt to sudden changes in the environment and traffic during an accident. Moreover, heuristic algorithms struggle to handle multi-objective optimization needs simultaneously, often requiring simplification to a single weighted objective, which shows limitations when facing situations where both time and radiation dose need to be minimized.
[0006] (4) Lack of a unified state modeling and decision-making framework: Nuclear emergency evacuation involves spatially and temporally evolving contaminated fields and road networks, as well as the states of thousands of agents, making it difficult to uniformly describe and make real-time decisions using traditional methods. Current research employs reinforcement learning to guide evacuation in confined spaces such as building fires, or uses deep Q-networks with multi-agent collaboration to improve evacuation efficiency in fire emergencies. These studies demonstrate that deep reinforcement learning has superior adaptive optimization capabilities in evacuation decision-making. However, in scenarios involving large nuclear accident areas, real road networks, and multiple vehicles, a unified solution combining deep reinforcement learning, multi-agent game theory, and dynamic environment modeling is still lacking. Currently, there are no publicly reported systems that integrate radiation diffusion simulation, traffic flow simulation, and multi-agent reinforcement learning strategies to form a decision-making system for nuclear emergency evacuation. This leaves a gap in the theory and application of nuclear accident evacuation path optimization, urgently requiring an innovative technical solution.
[0007] In summary, existing technologies for nuclear emergency evacuation route planning struggle to simultaneously meet the requirements of dynamic adaptability and multi-stakeholder collaborative optimization. Specifically, they fail to fully leverage the automated decision-making capabilities of deep learning to handle complex and ever-changing accident environments; they lack path-playing strategies that balance safety and efficiency in scenarios involving multiple vehicles; and they lack a unified digital model to support pollution spread and road network conditions. These shortcomings can lead to untimely and insufficiently optimized nuclear accident emergency evacuation plans, impacting the effectiveness of emergency response and personnel safety. Summary of the Invention
[0008] Purpose of the Invention: The purpose of this invention is to overcome the shortcomings of the aforementioned background technology and provide a multi-agent dynamic nuclear emergency evacuation path planning method and system based on deep reinforcement learning. This method introduces deep reinforcement learning into the field of nuclear accident emergency evacuation, integrating multi-agent collaborative mechanisms and path game modeling to achieve efficient and safe evacuation path optimization in complex environments of dynamic pollution diffusion and real road networks. Through this invention, optimal evacuation route plans can be automatically generated for different accident scenarios, improving the intelligence level and practical effectiveness of emergency evacuation decision-making.
[0009] Technical Solution: To achieve the above objectives, this invention provides the following technical solution: A multi-agent dynamic kernel emergency evacuation path planning method based on deep reinforcement learning, comprising the following specific steps:
[0010] S1. Environmental Status Acquisition and Initialization: Establish an environmental model for a nuclear accident emergency scenario. First, collect multi-source data related to the nuclear accident, including: accident source information (e.g., release intensity, altitude), real-time radiation monitoring data (readings from environmental dose rate monitoring stations and air radiation monitoring equipment), meteorological data (wind direction, wind speed, stability, precipitation, etc.), and digital map and traffic network data (road topology, road length and speed limits, real-time traffic conditions, etc.). Input the above data into a nuclear accident atmospheric diffusion model (e.g., a Gaussian plume model or a regional atmospheric dynamics model) to predict the diffusion range and intensity distribution of the pollution cloud at different times. Simultaneously, combine this with a road traffic model to estimate the road network's capacity and congestion in the initial stage of the accident. Based on the simulation results and monitoring data, generate the environmental status of the accident-affected area at the current time: for example, representing the radiation dose rate distribution at each location in a grid format, and the traffic conditions of road network nodes / segments. Set the initial location and number of personnel or vehicles (agents) to be evacuated. The starting point set of agents can be determined based on population distribution or the location of important facilities, and the destination set of agents can be determined based on predetermined shelters or safe areas. This completes the initialization of the environmental state, providing a basis for subsequent decision-making;
[0011] S2. Multi-Agent System Establishment and Scheduling: The objects to be evacuated are modeled as a multi-agent system. Each agent represents an entity that needs to be evacuated (individual, family vehicle unit, etc.), and its state includes attributes such as current location, destination, travel speed, and cumulative dose. Action rules for all agents are defined: at discrete decision moments, agents can choose their next action (e.g., choosing the next road segment to travel on). To improve coordination efficiency, a central scheduling and control mechanism is introduced to macroscopically coordinate the evacuation actions of the agents. The scheduling mechanism includes: grouping agents into batches according to geographical regions, setting staggered start times for different groups; assigning higher priority to special groups (such as the injured, sick, elderly, and children) to ensure they have priority access to road passage or guidance resources; controlling the flow rate of agents entering certain key intersections based on road capacity to avoid congestion caused by sudden influx. The structure of the multi-agent system is attached. Figure 3 As shown, it includes a central command node and numerous intelligent agent nodes. The central command node communicates with each intelligent agent through information broadcasting to monitor and coordinate the overall evacuation process.
[0012] S3. Path Game Modeling and Road Network Competition Analysis: Based on the environmental model and multi-agent framework, a game theory model for nuclear emergency evacuation path planning is established. Each agent's path selection from the starting point to the destination is considered a strategy of an autonomous decision-making entity, and each road resource is a limited resource shared by all agents. To characterize the competition among agents, a road occupancy competition model is defined: when multiple agents choose to enter the same road segment at the same time, the effective travel speed of that road segment decreases, and the time per unit distance increases. The congestion cost can be represented by the following functional relationship:
[0013]
[0014] Where t0 is the reference time for passing through road segment i when unloaded, n is the number of agents on the current road segment, and N is the number of agents on the current road segment. max The saturation congestion threshold for a road segment is given by the formula, which shows that travel time increases linearly or non-linearly with increasing n. Meanwhile, the pollution diffusion model provides the radiation intensity ρ(x,t) of each road segment over time; the agent traversing this road segment will accumulate a dose D. i =ρ i ×time i , where time i This represents the actual travel time for each agent on that road segment. Considering the above factors, a cost function, or utility function, is defined for each agent:
[0015] J = α·T + β·D + γ·C
[0016] Where T is the total travel time, D is the cumulative radiation dose, C is the congestion cost (which can be reflected in T or measured independently), and α, β, γ are weighting coefficients used to balance the three objectives of risk avoidance, speed, and equilibrium. In the game model, each agent aims to minimize its own cost J, but since ρ and C are jointly determined by all agents, their strategies are coupled. This multi-agent game can be regarded as a non-cooperative game, and the ideal outcome corresponds to a Nash equilibrium state, that is, any single agent changing its route unilaterally will not reduce its cost. Due to the complex time-varying state and multi-objective properties of the game, traditional analytical methods are difficult to solve. Therefore, this invention introduces reinforcement learning to approximate the equilibrium strategy of the game.
[0017] S4. Deep Reinforcement Learning Strategy Solution: The above path game optimization problem is solved using Deep Reinforcement Learning (DRL) technology. First, a deep neural network policy architecture is designed, taking a multi-dimensional environment state tensor as input and outputting the action decisions of each agent. Specifically, the design of the environment state tensor is shown in the attached figure. Figure 2 As shown, it can include: a two-dimensional grid of the contaminated field and its time series (forming a 3D tensor input for ConvLSTM or 3D-CNN to extract spatiotemporal features), road network topology and congestion status (which can be extracted using a graph convolutional network, or mapped to the grid and the contaminated field superimposed input), and the current location distribution of the agents (which can be used as an additional channel). The neural network can adopt a parameter-sharing approach, where all agents share the same network structure (the input is the global state, and the output can represent the joint policy of all agents' actions, or the policy of a single agent but applied to each agent by replicating the network), thereby significantly reducing the number of parameters and naturally achieving experience sharing. Then, a reward function R is designed based on the optimization objective of nuclear emergency evacuation. To facilitate solving, the cost J defined in step S3 is converted into a reward R = -J (the smaller the cost, the higher the reward). At each decision time step, the immediate reward given to the agent consists of the negative cost experienced in that step, for example:
[0018] r t =-(α·δt+β·δd+γ·δc)
[0019] Where δt is the travel time for this step, δd is the additional dose for this step, and δc is the additional delay caused by congestion for this step. An additional reward R is given when the agent successfully reaches the safe zone. goal To encourage successful evacuation, penalties are imposed if agents remain in high-risk areas for extended periods or exceed the expected evacuation time, thus urging them to change their strategies. Next, an appropriate deep reinforcement learning algorithm is selected to train and optimize the multi-agent policy. Multi-agent dynamic deep Q-networks (MD-DQN) or policy gradient algorithms can be used. For example:
[0020] (1) Value-based algorithm: The above-mentioned shared neural network is used as an approximate Q function Q(s,a). θ Given a global state `s` and action `a` (which includes the joint action of all agents – a large space that can be practically decomposed), parameters `θ` are trained through multi-agent version experience replay and Q-learning updates. Due to the high dimensionality and complex combination of actions, a decomposed joint action-group output method can be used to reduce dimensionality, or a dual-deep Q-network can be used for stable training. Alternatively, a centralized training, distributed execution framework can be introduced. During centralized training, the Q-value of each agent is updated using the global state; during execution, each agent selects actions based on its local observations, approximating cooperation.
[0021] (2) Policy-based algorithm: adopts the Actor-Critic architecture, policy p i (a|s) θ And value function V w (s) is represented by two neural networks. During training, the value of the global policy is evaluated using a Critic network shared by all agents, and the advantage function 'a' for each agent's action is calculated. Then, the parameters θ of each agent's Actor network are updated using Policy Gradient. To promote multi-agent collaboration, a centralized Critic and independent Actor approach can be adopted, where the Critic network calculates the gradient using global information, and the Actors adjust their policies independently to achieve collaborative optimization. Alternatively, advanced algorithmic frameworks such as Multi-Agent Proximal Policy Optimization (MAPPO) or Multi-Agent Deep Deterministic Policy Gradient (MADDPG) can be used to accelerate convergence and enhance stability. The entire training process is repeated in a simulation environment: the simulation environment generates a large number of nuclear accident evacuation simulation samples (including random initial accident conditions, wind direction changes, road accident interference, etc.) based on the current policy. The agents continuously try and update their policies in the simulation until the policy converges and meets the performance requirements.
[0022] After training, the deep reinforcement learning model can approximate a solver for a multi-agent game. At this point, each agent's strategy fully considers the pollution spread trend and road congestion effects, and has the ability to adaptively adjust routes according to different situations.
[0023] S5. Dynamic Cooperative Path Planning: Deploying the trained strategy model for actual decision-making. Given the environment state tensor s at a certain moment. t(Obtained in real-time from step S1), this is input into the deep reinforcement learning policy model, which outputs action suggestions for each agent. For example, for decisions at discrete time steps, it can output the next road node each agent should go to; or output a complete path plan after a certain look-ahead step. Since the agent policies have shared information during training, the directly output set of routes for each agent already contains a certain degree of cooperation. However, to ensure the stability and effectiveness of the scheme in real-world execution, the initial scheme is further optimized and adjusted through a coordination module. Coordination optimization includes: simulating road load based on the current global scheme, identifying road segments that may become over-congested and the set of agents affected; for severely congested road segments, attempting to have some agents take suboptimal paths and evaluating the change in global cost; if the overall cost decreases, the adjustment is adopted (equivalent to performing a round of local game equilibrium adjustment); this process iterates until no agent can unilaterally change its route to further reduce its own cost, i.e., reaching an approximate Nash equilibrium state. This process is essentially a fine-grained search around the solution given by the reinforcement learning policy to ensure the global optimality and balance of the results. After optimization, the final multi-agent collaborative evacuation path planning scheme is obtained, including the specific evacuation route and estimated travel time for each agent (or each type of person / each vehicle);
[0024] S6. Implementation and Dynamic Adjustment of the Evacuation Plan: The final evacuation plan is submitted to the emergency command decision-making body and distributed to all implementing units and the public via the communication network. During implementation, emergency broadcasts, electronic displays, and navigation app notifications can be used to guide personnel to evacuate along designated routes. For the transportation system, traffic signal control can be integrated to manage recommended routes (e.g., extending green light durations on recommended routes) to ensure the smooth implementation of the plan. During the plan's execution, the method of this invention continues to monitor the evolution of the environment and the agent's state: the state update module obtains the latest radiation dose rate field, weather changes, and road accessibility information from the monitoring system, continuously updating the environmental state tensor s. t The decision-making module quickly recalculates the next action of each agent based on the latest state and existing strategies. In most cases, the trained strategy model can adaptively adjust to minor changes in the environment without completely replanning the route. However, in the event of a major sudden change (such as a sudden change in wind direction causing a pollution cloud to change its propagation direction, or a planned route being suddenly interrupted due to a traffic accident), the system will automatically enter a dynamic adjustment mode: inputting the new environmental state into the strategy model, obtaining the adjusted path decision, and using the coordination module to quickly correct the routes of the affected agents and guide them to change course. Throughout the execution process, the strategy management module monitors the progress and cumulative dose of each agent. If any agent deviates from the route or exhibits abnormal behavior, it can trigger manual intervention prompts. Finally, after the threat of the incident is eliminated, the system can store all action logs and environmental data in an experience base to provide a basis for post-event analysis and improvement.
[0025] Through the above steps, this invention achieves optimal evacuation path planning through multi-agent cooperation in a dynamic nuclear accident environment. The flowcharts for each step are attached. Figure 1 As shown: Environmental data continuously flows into the state model, reinforcement learning strategies provide decision suggestions, coordination modules optimize and adjust, and finally output an executable solution and receive feedback for closed-loop adjustment.
[0026] Preferably, the multi-source data acquisition in step S1 includes an automatic adaptation mechanism for data format and validity: an input adapter library for road and environmental data is established, supporting multiple sources such as static map files, real-time sensor streams, and traffic APIs. Data formats (such as GIS road data, JSON sensor streams, CSV population distribution, etc.) are automatically determined through pattern recognition and content indexing, and environmental data is adaptively loaded using batch or streaming processing modes. To address potential data gaps or anomalies in emergency situations (such as sensor disconnections), a hybrid strategy of rule-based and reinforcement learning is introduced during the data acquisition phase for data imputation or estimation to ensure the integrity of the environmental model. The environmental state information generated in this phase is standardized, encoded, and then passed to subsequent steps as the basis for evacuation decisions.
[0027] Preferably, in step S2, corresponding evacuation strategy templates are defined for different accident scenarios: In the case of a nuclear power plant accident, priority is given to long-distance, large-scale evacuation of personnel. The system will delineate key evacuation zones (such as a range of several kilometers centered on the accident source) and determine the evacuation order by distance and wind direction. In the case of a radioactive source transportation accident, the impact may be smaller, and the system will implement local traffic control and crowd evacuation in the neighborhood near the accident site. Under severe weather conditions (such as heavy rain affecting traffic), the scheduling strategy will increase its sensitivity to changes in road capacity and avoid easily flooded or accident-prone road sections in advance. In the case of an urban nuclear terrorist attack, the core area will be quickly sealed off and multi-directional, multi-exit diversion routes will be planned. The scheduling process utilizes a reinforcement learning model trained offline based on historical exercise data. In real-time applications, the system adaptively selects the most suitable evacuation strategy parameter set according to the accident type and continuously fine-tunes it through online feedback to meet the needs of specific scenarios.
[0028] Preferably, the deep reinforcement learning path optimization in step S3 adopts a multi-agent reinforcement learning architecture of centralized training and distributed execution (CTDE). During training, offline learning is performed on various evacuation scenarios using a large number of simulation environments. Agents optimize their policies with the reward objective of reducing total evacuation time and radiation exposure; at the same time, a penalty term is introduced to constrain individual behavioral biases (e.g., if a single agent excessively delays its own evacuation to make way for others, it is appropriately restricted to balance fairness and overall efficiency). In the execution phase, each agent independently decides on a specific route based on the trained policy, supplemented by real-time local optimization. A high-level deep Q-network or policy gradient model is responsible for global game decision-making, while a low-level bandit or local policy network adjusts individual behavior according to the real-time state, thereby achieving a unity of global optimum and individual adaptation. Through this hierarchical reinforcement learning strategy, the system can effectively explore a vast path combination space and produce an effective evacuation plan within a reasonable time.
[0029] Preferably, the multi-agent collaborative mechanism in step S4 includes two aspects: information sharing and game equilibrium. Agents share key state information (such as regional traffic density and path risk index) through a cloud platform, reducing unnecessary communication overhead while ensuring real-time performance. A distributed strategy game is used to solve for an approximate Nash equilibrium of the current traffic allocation, ensuring that no single agent is motivated to unilaterally change routes for higher gains (i.e., the overall solution achieves Pareto optimization). In the event of drastic environmental changes (e.g., a sudden road closure), the system triggers a rapid replanning module. The local agent set re-optimizes sub-problems (e.g., local evacuation plans bypassing closed roads) through game theory, while the global agent strategy is adjusted according to the principle of minimum perturbation, minimizing new conflicts in already evacuated areas. Through this mechanism, step S4 ensures the robustness of the evacuation process and significantly reduces the risk of the global solution collapsing due to the unexpected failure of a single path.
[0030] Preferably, the standardized data structure defined in step S5 can be expanded as needed to include environmental and process audit information. For example, it can record the environmental snapshot ID, key constraints (such as the maximum permissible dose threshold), and estimated uncertainty range used when each route plan is generated, for post-event analysis and decision tracing. The output supports visualization and manual verification: for example, the route is drawn on a GIS map based on the route coordinate set, and a heat map is generated based on the dose estimate and overlaid on the evacuation area map for command personnel to intuitively assess route safety. The standardized output also considers cross-platform compatibility, allowing direct access to mainstream navigation software or traffic signal control systems to achieve automated emergency coordination.
[0031] Preferably, the interface adaptation in step S6 supports multiple communication methods and data formats, including XML / JSON Web services, mobile push notifications, and traffic broadcast protocols, ensuring that different emergency participants can obtain evacuation instructions in real time. The system provides permission management and encryption verification mechanisms to ensure the security and reliability of the evacuation route plan during its release. Depending on actual deployment needs, a centralized evacuation decision server can be deployed in the cloud to send information to the local command terminal via 5G / satellite communication; alternatively, a simplified model can be deployed on edge computing nodes to enable local evacuation guidance in the event of a network outage. The interface adaptation layer issues reminders or adopts backup plans (such as using loudspeakers to guide personnel) based on feedback (e.g., some users do not respond in time) to improve the actual execution effect of the plan.
[0032] Compared with the prior art, the present invention has the following beneficial technical effects:
[0033] (1) Real-time dynamic decision-making capability: This invention utilizes deep reinforcement learning strategies to perceive dynamically changing radiation fields and traffic conditions in real time, and autonomously adjust evacuation routes, achieving rapid response of route planning to environmental changes. Compared with traditional methods of pre-planning static routes, this invention can continuously optimize decisions during the development of an accident, ensuring that the evacuation plan always matches the latest accident situation, greatly improving the timeliness and effectiveness of emergency response.
[0034] (2) Multi-agent cooperative optimization: This invention models the evacuation problem as a multi-agent game, achieving cooperative optimization among the evacuation entities through mechanisms such as shared strategies and centralized evaluation. Different vehicles and personnel no longer seek their own shortest paths independently, but learn to consider the behavior of others during reinforcement learning, achieving overall balanced traffic distribution—avoiding excessive congestion on individual main routes while ensuring no waste of evacuation resources. Experiments show that the method of this invention can significantly reduce congestion on peak road sections and improve overall traffic efficiency.
[0035] (3) Multi-objective integrated balance: This invention effectively integrates objectives such as risk avoidance (radiation dose), accessibility (time / distance), and balance (road load) by customizing a reward function, overcoming the difficulty of multi-objective coordination in traditional methods. During training, the algorithm automatically explores a risk-efficiency trade-off strategy, minimizing evacuation time while ensuring personnel dose does not exceed limits, or ensuring the optimal overall risk-reward ratio when slightly increasing the distance to avoid high-risk areas. The generated route plan thus better meets the actual needs of nuclear emergency response, balancing safety and efficiency.
[0036] (4) Unified Intelligent Decision-Making Framework: This invention proposes a multi-dimensional tensor modeling method for pollution diffusion and road network status, and inputs it into a deep neural network for decision-making, realizing end-to-end intelligent optimization from sensor data to evacuation action decisions. Compared with the traditional approach of manual coordination between radiation protection experts and traffic management personnel, this solution directly maps complex environmental information into machine-understandable states, and uses artificial intelligence to complete complex reasoning in one go, reducing human judgment errors and delays, and improving the automation and reliability of the decision-making chain.
[0037] (5) Strong engineering feasibility and applicability: The method and system of this invention fully consider the engineering conditions of nuclear accident emergency scenarios. Through the strategy of offline pre-training + online fine-tuning, it ensures that the model can be quickly deployed and calculate solutions in real time when an accident occurs, and can run inference in real time on ordinary computing servers, meeting the emergency timeliness requirements. At the same time, the system supports docking with existing nuclear emergency monitoring and command platforms, with clear input and output interfaces, making it easy to integrate into actual emergency command processes. In addition, the method of this invention has a certain degree of versatility. After parameter adjustment, it can be applied to different nuclear power plant site layouts and evacuation planning for different population sizes, and can also be extended to emergency evacuation decision support for other large-scale disasters (such as chemical spill accidents).
[0038] In summary, this invention innovatively integrates deep reinforcement learning and multi-agent game theory to achieve intelligent optimization of personnel evacuation routes in nuclear emergency situations. It has clear technological advancements and practical value, and can significantly improve the efficiency and safety of personnel evacuation in nuclear accident emergency responses. Attached Figure Description
[0039] Figure 1 This is a flowchart of the method of the present invention.
[0040] Figure 2 Design diagram for the environmental state tensor.
[0041] Figure 3 This is a diagram of the structure of a multi-agent system.
[0042] Figure 4 This is a diagram illustrating a multi-agent evacuation path planning example.
[0043] Figure 5 This is a schematic diagram of the environment tensor.
[0044] Figure 6 This is a schematic diagram illustrating the effect of an example. Detailed Implementation
[0045] The method and system of the present invention will be further described in detail below with reference to specific embodiments. It should be understood that the embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0046] Example 1
[0047] Assume a nuclear terrorist attack (such as a small nuclear device explosion) occurs in the core area of a city center, releasing high-intensity radioactive material. The epicenter is located in Area A, a densely populated area, requiring a large-scale evacuation under the dynamic and hazardous environment of nuclear radiation spread. This embodiment uses street network data from OpenStreetMap of a certain city for simulation, and combines it with a nuclear accident consequence assessment model to generate a radiation dose distribution heatmap after the accident, verifying the effectiveness of the multi-agent path planning method of this invention. An example diagram of multi-agent evacuation path planning is shown below. Figure 4 As shown.
[0048] S1 Environment Status Acquisition and Status Initialization:
[0049] S11. Environmental Data Acquisition: First, acquire digital maps and road network data for Area A of a certain city, including road nodes and connectivity, road speed limits, etc. Simultaneously, collect population distribution data and the locations of major shelters in the area. Assuming a nuclear explosion occurs, obtain the initial radioactive source term (release intensity, nuclide type) at the epicenter and meteorological conditions such as wind direction and speed at the time of the explosion through accident models and monitoring. Based on this information, use a Gaussian plume diffusion model to calculate the diffusion range of the radioactive contamination cloud and the dose rate at various locations within one hour after the accident. Obtain the radiation dose rate field covering area A (e.g.) Figure 4 (The blue-green radiation heat zone in the image). In addition, the traffic monitoring system obtains the current road conditions, such as when some roads are unavailable due to accident damage or traffic control.
[0050] S12. Environmental State Modeling: Input the above data into the environmental modeling module to construct the state tensor E(x,y) of the simulation environment. Using region A as the computational domain, discretize it into grid cells, for example, 101×101. Add the following channel information to each grid: (1) Dose rate channel, storing the radiation dose rate value (μSv / h) at the grid location calculated by the nuclear diffusion model; (2) Road access channel, indicating whether there is a passable road in the grid (1 indicates a road, and a basic passage time is assigned according to the road speed limit, 0 indicates no road or obstruction); (3) Population density channel, storing the population number or density of the area where the grid is located; (4) Shelter channel, marking the location of the main shelters and their maximum capacity, etc. The above multi-layer data are fused to form a three-dimensional tensor environment E. For example, Figure 5 The image shows an example environment tensor, where the red channel represents radiation intensity, the black and white channels represent road distribution, and the blue channel represents the spatial variation of population density.
[0051] S2 Multi-Agent System Establishment and Scheduling:
[0052] In the context of an emergency evacuation following a nuclear terrorist attack in a certain city, we model all evacuees as a multi-agent system (in this case, 8 agents) where multiple agents work collaboratively. Their initial positions are distributed across eight points within and around the radiation contamination area. Each agent is assigned action attributes, including current location, destination, speed, and cumulative dose. To improve collaboration efficiency, we introduce two agent roles: "Leader" and "Follower," with two Leader agents, two Follower agents, and four other agents. Specifically:
[0053] S21. Leader Agent: Acting as the leader of the group, each leader is responsible for a group of nearby evacuees. Leader agents possess broader environmental awareness and communication capabilities, obtaining global information (such as citywide road conditions and distribution) from the central command system and communicating with their followers within the group. Their main responsibilities include formulating evacuation routes and strategies for the group, coordinating the action pace of agents within the group, and negotiating with other agents to avoid conflict. The leader agent is equivalent to a "navigator" for evacuating convoys or crowds, exploring the route and guiding the group forward during the evacuation process. To improve decision-making quality, the leader agent can use methods such as model predictive control to anticipate future road conditions, thereby leading followers to reach a consensus on the path more quickly.
[0054] S22. Follower Agents: Representing specific individuals or vehicles involved in evacuation, these agents primarily act according to the instructions of the leader agent and their own established rules of conduct. Followers typically possess only localized perception capabilities (e.g., local obstacles, nearby vehicles) and limited decision-making authority. They follow routes planned by the leader or adhere to established rules of conduct, autonomously avoiding collisions or hazards within their local area (e.g., navigating around obstacles). Followers also report their position, speed, and any anomalies encountered to the leader, allowing the leader to adjust strategies promptly. Through this mechanism, follower agents form a coordinated and obedient team, reducing route conflicts that could result from individual agents acting independently.
[0055] S23. Other intelligent agents: Intelligent agents with action rules.
[0056] Leader-Follower Coordination Mechanism: In terms of scheduling, the leader agent is responsible for task allocation and formation maintenance. The leader breaks down the overall evacuation task into smaller, manageable parts for each follower, such as assigning a refuge location or sub-route to each follower. Each follower proceeds along their assigned path, and the leader monitors the direction, formation, and spacing of the evacuation in real time, issuing acceleration, deceleration, or detour commands as needed to maintain the safety and efficiency of the evacuation. To prevent conflicts between multiple evacuation groups, the leader agents share their location and route intentions through vehicle networks or a command system, implementing a yield negotiation mechanism: when two leader-led evacuation groups are about to meet on a road segment, they negotiate priority through communication (e.g., allowing the group closer to the danger zone to pass first), thereby avoiding congestion or collisions.
[0057] Dynamic Role Switching Strategy: The system employs a flexible leadership election and switching strategy. When a leading agent malfunctions (e.g., due to vehicle breakdown or communication interruption), its most senior or capable follower will be promoted to the new leader to ensure the team's continued evacuation. Conversely, if two evacuation teams merge en route, one leader can temporarily assume joint command, while the other steps down to a deputy or follower role, reducing multiple command structures. Role switching is based on predetermined rules (such as priority or task requirements) and is automatically triggered by the scheduling system when necessary. This dynamic hierarchical mechanism enhances the system's robustness, ensuring the overall evacuation mission continues even if some agents fail.
[0058] Global Scheduling and Priority Control: The scheduling mechanism coordinates the evacuation actions of each agent in both time and space. Considering the limited capacity of a city's road network and the varying risk levels in different areas, the system groups agents by region and population type, assigning different evacuation priorities. For example, area A can be divided into several evacuation zones based on geographical location, with a designated leader agent responsible for each zone, and the start times for evacuation teams in each zone are staggered. Agent groups in the accident epicenter (high-risk area) initiate evacuation first, while those in safer outer areas depart later, evacuating in batches to avoid peak congestion. This scheduling method allocates different batches in time and plans different routes in space, preventing large flows of people and vehicles from simultaneously converging on the same path and causing congestion. By limiting the departure time window of each group through the scheduling control module, peak load reduction is achieved, ensuring the orderly utilization of the road network.
[0059] The aforementioned multi-agent system architecture utilizes a leader-follower hierarchical collaboration to achieve hierarchical transmission of information from the central to the local level and decision-making from the global to the individual level. On the one hand, the leader agent coordinates the overall situation and resolves conflicts; on the other hand, the follower agents autonomously execute micro-level actions to ensure safety and obstacle avoidance. This architecture leverages both the global optimization capabilities of centralized command and the flexibility of distributed autonomy, providing an engineering-feasible and efficient organizational mechanism for large-scale urban emergency evacuation.
[0060] S3 Path Game Modeling and Road Network Competition Analysis:
[0061] To address the route competition problem caused by the simultaneous evacuation of multiple agents, we incorporate game theory mechanisms into our reinforcement learning model, treating route planning as a strategy game with limited road resources. Each agent (or its leader agent) chooses a route from its starting point to the refuge as its strategy. All agents share the city's road network resources, thus route choices influence each other. If too many agents concentrate on the same road, congestion, delays, and even increased radiation exposure risks will occur—this manifests as road occupancy competition in the game model: when multiple agents travel on the same road segment, the unit cost of travel on that segment increases with the number of agents, resulting in additional penalties for all participants. Through this modeling, we quantitatively characterize the interaction between agents under dynamic pollution environments and limited road network resources; that is, each agent's decision affects the outcomes for others.
[0062] Payoff / Cost Function Modeling: To formalize the above game, a utility function or cost function for each agent is defined in the path planning, incorporating multiple factors. Let the path chosen by agent i be P. i Then its total cost can be expressed as:
[0063]
[0064] Where t(e) is the baseline travel time for road segment e in the path, and N e D(P) represents the number of agents simultaneously using road segment e, where α is the congestion impact coefficient; i ) indicates that the agent traverses path P i The accumulated radiation dose or risk value, β is the corresponding risk weight. This formula indicates that if other agents choose the same travel route (N... e If the travel time cost of i is greater than 1), then the congestion factor will increase, leading to a rise in the total cost; simultaneously, the path passing through a high-radiation area will also increase the risk cost. This is equivalent to a non-cooperative game model: each agent hopes to minimize its own cost J. i However, the limited road resources create a competitive relationship of give-and-take in their decision-making.
[0065] Local Broadcasting and State Sharing: To address potential conflicts arising from independent agent decision-making, we introduce a local communication mechanism. Agents can broadcast their state (e.g., current location, speed) and intentions (e.g., planned next road / destination) within a certain range via vehicle-to-everything (V2X) or mobile communication. For example, a leader agent can broadcast its convoy's planned route or the time it will approach an intersection, allowing other leaders in the vicinity to be informed in advance. Follower agents can also broadcast abnormal situations (e.g., traffic jams ahead) to their group or nearby vehicles. Through this state sharing, each agent gains a wider perception range than its own sensors, approximating the strategy information of other players in the local game. This provides an informational basis for subsequent negotiation, helping to prevent multiple agents from simultaneously choosing the same path due to information asymmetry.
[0066] Negotiation Mechanism and Game Theory Solution: When a potential conflict is detected (e.g., multiple agents broadcast their intention to occupy the same road segment), the system triggers a local negotiation algorithm to resolve the conflict. Negotiation can take place between leading agents or distributedly with the assistance of a central scheduler. The basic idea is to have the relevant agents engage in a small-scale game to find the equilibrium for the conflicting road segment: each party can choose a strategy such as "staying on the straight route," "changing the route and taking a detour," or "slowing down and yielding," each strategy corresponding to a different cost (delay time, increased risk, etc.). By exchanging their estimated costs for their respective strategies, the agents jointly find a strategy combination that minimizes the overall risk of occupancy. For example, in the case of two groups of vehicles vying for the same exit ramp, if both sides rush through, it will cause severe congestion (a lose-lose situation), while allowing one side to go first and the other to wait minimizes the total delay. The negotiation algorithm can calculate the waiting cost and rushing cost for each party, finding a solution that minimizes the total cost and eliminates the incentive for either party to change their strategy; this is a Nash equilibrium solution. Studies have shown that, when the cost function is designed appropriately, the globally optimal solution in a evacuation path game is often also the Nash equilibrium of the game, meaning that no single agent can further reduce its cost by unilaterally changing its route. Based on this principle, we use methods such as iterative optimal response to solve local games: relevant agents take turns adjusting their strategies to reduce their own costs until the strategy combination converges, at which point any further change by any party will worsen the situation, and equilibrium is reached.
[0067] By sharing information through localized broadcasting and resolving issues through negotiation and game theory, agents can achieve a degree of cooperation while making decentralized decisions, avoiding collective suboptimal outcomes caused by simplistic greed. The ultimate result is that agents rationally allocate road resources in an equilibrium manner: no one concentrates on a particular path, thus approximately minimizing the total risk of road occupancy and congestion losses. This game-theory-enhanced path selection model ensures the stability and efficiency of the evacuation strategy under multi-agent interaction—once each agent follows the equilibrium strategy, there is no incentive for unilateral deviation, which is equivalent to reaching a tacit understanding that benefits everyone, significantly alleviating competitive conflicts under limited road conditions.
[0068] Solving the S4 deep reinforcement learning strategy:
[0069] Agent and Reinforcement Learning Setup: In this disaster scenario, we consider multiple evacuation guidance vehicles as agents, responsible for guiding people to evacuate along safe routes. We set initial positions for several evacuation vehicles (e.g., 8 in this case), dispersed near the accident-affected area. The state of the reinforcement learning agent is defined as a fragment of the environmental tensor within a certain range around its current position, as well as information such as the destination (shelter) location (here, moving to a point outside the evaluation area or certain uncontaminated locations). The agent's action is defined as selecting the direction of the next adjacent intersection (or grid cell) to proceed on the road network. Within each step time δt, the agent moves a distance of one road network node. To encourage the agent to approach the evacuation target or destination as quickly and safely as possible, the reward function is designed as follows: When the agent moves from position x to the neighboring position x':
[0070] (1) If the final refuge destination is reached, a one-time positive reward R will be given. g (Indicates a reward for successful evacuation), and terminates the agent's task;
[0071] (2) Otherwise, a step reward will be given:
[0072] r t =-(α·δt+β·δd+γ·δc)
[0073] Where δt is the travel time for this step, δd is the additional dose for this step (which can be the average of the dose rates at the current position and the next position multiplied by δt), and δc is the additional delay caused by congestion for this step. This reward function is consistent with the aforementioned formula and reflects the trade-off between travel time and radiation risk.
[0074] Before training begins, the agent's policy network parameters are randomly initialized, and hyperparameters such as the discount factor γ and learning rate α are set. The DQN algorithm is used for centralized training: the global value network inputs the joint state and actions of all agents to evaluate the team value, thereby guiding each agent's policy update. During training, by continuously interacting with the simulation environment (allowing the agents to attempt evacuation routes in a simulated accident environment in area A), the agents receive rewards and update their policies. After tens of thousands of iterations, the agents gradually learn to avoid high-dose areas and seek low-risk and fast path coordination solutions.
[0075] After training convergence, the obtained agent policies are applied to evacuation route planning in a city nuclear accident scenario. Assume that within 30 minutes of the accident, people need to be guided to evacuate to several shelters. Each agent makes action decisions based on the current environmental state tensor E. Due to the multi-agent collaboration, the agents implicitly coordinate through a shared value function: for example, they automatically avoid simultaneously choosing the same path, which could cause congestion, or dynamically allocate different shelters to balance the load. In the simulation, the agents advance one step every δt time steps, gradually forming a complete route from the starting point to the shelter. Path generation process details:
[0076] S41. Initial Phase: When the agent is at its starting position, it senses the level of radiation dose in the surrounding area and the accessibility of the road. If an agent detects that the radiation dose rate in the block directly in front of it is extremely high, it will tend to choose a detour route rather than heading straight towards the target under the influence of the policy network.
[0077] S42. Intermediate Stage: As the agent progresses, it continuously perceives new local environmental tensors and adjusts its strategy. For example, if a previously safe road suddenly becomes congested (traffic data update) or its radiation spread changes, the agent in this algorithm can perceive this in real time and replan the route. Because multi-objective optimization has been considered during training, the agent can flexibly respond to environmental changes, ensuring the rationality and safety of the path.
[0078] S43. Arrival Phase: When an agent successfully arrives at the target shelter or evacuation zone boundary, its route, time, and accumulated risks are recorded. If there are still agents that have not arrived, guidance continues until all evacuation tasks are completed.
[0079] S5 Dynamic Collaborative Path Planning:
[0080] After the aforementioned game theory modeling and reinforcement learning training (S4 has obtained the agent cooperation strategy), this step applies it to a real-world urban nuclear attack evacuation scenario to generate the final globally optimal evacuation plan. We designed a cooperative path planning module that comprehensively considers multiple objective requirements and ensures that the routes of each agent do not overlap as much as possible and avoid mutual interference.
[0081] S51. Multi-objective Weighted Optimization: Evacuation route planning needs to simultaneously satisfy objectives such as minimum time, minimum safety risk, and balanced refuge resources. Therefore, a multi-objective weighted optimization strategy is introduced during planning, unifying factors such as time, risk, congestion, and refuge capacity into a single evaluation function. For example, the comprehensive cost of each candidate route P is defined as follows:
[0082] F(P)=ω t ·T(P)+ω d ·D(P)+ω c ·C(P)
[0083] Where T(P) is the travel time, D(P) is the cumulative radiation dose along the path, C(P) can represent the degree of overlap with other agent paths or the cost of causing congestion, and ω t ,ω d ,ω c These are the corresponding weight coefficients. By appropriately setting the weights, we transform the original multi-objective optimization into a single-objective optimization with scalar cost values. The reward function in the reinforcement learning phase employs a similar weighted summation (e.g., assigning negative weights to time, dosage, and congestion level), enabling the agent to learn to balance risk and efficiency during training. Therefore, the trained policy naturally tends to lead each agent to choose the optimal or suboptimal path under the aforementioned comprehensive cost.
[0084] S52. Generating Cooperative Path Solutions: In a real-time environment (a city after an accident), the current state is input into a pre-trained multi-agent policy model, which outputs preliminary path decisions for all agents in batches. These decisions provide suggestions on which direction or road each agent should travel at this moment. After receiving these suggestions, the cooperative planning module combines the game theory model from S3 to perform conflict detection and coordination optimization: if multiple agents' preliminary decisions show overlapping or competing paths, the game solver is synchronously invoked to adjust the route choices of some agents until there are no significant conflicts. This process may require multiple iterations: continuously adjusting the agent path solutions so that no one can significantly improve their own situation by unilaterally changing their path, i.e., reaching an approximate Nash equilibrium. The solution obtained after iterative convergence is both stable (equilibrium) and globally optimized, balancing overall traffic efficiency and road load balance while satisfying safety and risk avoidance requirements.
[0085] S53. Shelter Allocation and Capacity Constraints: The collaborative planning module also considers the constraint of shelter capacity limitations. Since each shelter in a city has a limited capacity, the system must ensure that too many people are not concentrated at the same shelter during planning. This is achieved by introducing shelter capacity costs during agent path selection and target allocation: if the number of agents allocated to a shelter is close to its limit, the cost of continuing to allocate to that target increases significantly, thus prompting some agents to choose other shelters. This approach is equivalent to adding a penalty to multi-objective optimization to prevent the solution from violating capacity constraints. Alternatively, the planning process can adopt a strategy of allocation followed by adjustment: initially allocating shelters based on distance or minimum risk, then checking the number of people in each shelter; if the limit is exceeded, those exceeding the limit are reassigned to the next best shelter, iterating until all capacity constraints are met. Through these measures, the final solution ensures that individual time / risk is minimized while globally dispersing the flow of people seeking shelter, preventing individual shelters from becoming overloaded.
[0086] S54. Output of Global Optimal Evacuation Plan: Once collaborative optimization is complete, the system generates the final evacuation plan for each agent, including:
[0087] (1) Route: A series of road nodes or navigation guides from the starting point to the designated shelter;
[0088] (2) Estimated time: Based on the current traffic conditions and possible congestion, estimate the time it will take for the agent to arrive from its departure point;
[0089] (3) Risk assessment value: such as cumulative radiation dose or exposure time in high-risk areas along the route, a safety risk score is given for each route.
[0090] The above information can be organized into lists or tables for commanders to view, or it can be directly distributed to the navigation terminals of each agent. For example, agent A (located in Midtown A) is assigned the route: "North along Avenue C - turn onto Bridge D - proceed to Vault X," with an estimated travel time of approximately 25 minutes and an estimated radiation dose of 2.5 mSv. Agent B (located in Downtown) is assigned the route: "Enter Vault Y via Bridge B," with a travel time of 20 minutes and a dose of 1.8 mSv. The system ensures that the routes of A and B only meet at necessary intersections and are staggered, thus satisfying their respective rapid evacuation needs while avoiding congestion at bridgeheads.
[0091] S55. The final generated global evacuation plan achieves overall cooperative optimality: no single agent can change its route without increasing its own time or risk costs; in other words, the plan reaches an approximate Nash equilibrium in path game theory. Simultaneously, various optimization objectives (time, risk, and shelter load) achieve a comprehensive balance—the overall evacuation time is close to the shortest, the total radiation exposure is the lowest, and the load on critical roads and shelter facilities is within acceptable limits. This provides command with an executable, intelligently optimized evacuation route plan, significantly improving the scientific rigor and effectiveness of evacuation plans in complex scenarios such as nuclear terrorist attacks.
[0092] S6 Scheme Implementation and Dynamic Adjustment:
[0093] After the optimized evacuation plan is finalized, the actual implementation phase begins. The system disseminates the plan to relevant departments and intelligent agent terminals (such as navigation devices and smartphones) via the emergency command communication network. Each leadership intelligent agent then directs followers to begin evacuation along designated routes, while the entire system enters real-time monitoring and feedback mode.
[0094] S61. Environmental Status Feedback: The emergency monitoring network continuously collects the latest data on the accident site and the urban environment, including radiation diffusion (dose rate changes detected by monitoring stations and drones) and traffic flow dynamics (road speed and congestion levels monitored by traffic cameras and vehicle location data). This data updates the environmental status model in real time, and the system will trigger corresponding adjustment procedures once changes that do not conform to previous predictions are detected. For example, when a sudden increase in radioactive dose or expansion of the contaminated area is detected in a certain direction, the risk weight of that area in the model will be increased; or if a traffic accident is detected causing congestion on a main road, the traffic capacity of that road segment will be reduced or marked as congested.
[0095] S62. Project Execution Monitoring: The system tracks the progress of each agent, utilizing mobile communication or IoT devices to obtain vehicle GPS locations or personnel movement information. This allows for comparison of actual progress with the original plan: if an agent is found to be significantly behind schedule or deviating from the planned path, the system will record the anomaly. The leader agent will also report its team's status, indicating whether the team is progressing smoothly and whether all followers are keeping up. If everything is proceeding as planned, the system only needs to maintain monitoring and information dissemination without intervention.
[0096] S63. Dynamic Replanning Mechanism: When the accident scenario undergoes sudden changes or unforeseen circumstances occur during execution, the system will initiate local or global route replanning to ensure the effective execution of evacuation operations. Possible triggering events and corresponding strategies include:
[0097] (1) Expanding contamination area: For example, due to a change in wind direction, radioactive clouds spread in a new direction, causing previously safe paths to suddenly fall into high-dose areas. Upon detecting this change, the system immediately marks the affected paths as "high-risk" and identifies the set of agents currently using these paths. For these agents, a reinforcement learning policy model is activated to replan their actions: using their current location information and the updated environmental state as input, their new optimal actions are calculated. If there is an alternative route to avoid the contaminated area, they are guided to switch to the new route; if there is no alternative route and the risk is too high, they may be instructed to temporarily stay put or go to a temporary refuge point, awaiting further instructions.
[0098] (2) Road Congestion or Traffic Jams: If monitoring detects an unexpected blockage (such as a building collapse or vehicle accident) or severe traffic congestion causing extremely slow traffic on a critical road segment, this information will be broadcast to nearby agents. Agents on the affected road segment will enter a local task replanning state: their leader agent will attempt to lead the team to change routes, such as choosing pre-selected parallel roads or detours to bypass the congested area. At this time, a game-theoretic negotiation mechanism needs to be used again to ensure that the multiple teams that have changed routes do not converge on another road and cause secondary congestion—this can be guided by global information such as knowledge graphs to achieve real-time optimization and adjustment of route selection under the new environmental conditions. The collaborative path planning module re-runs a round of local optimization in the background, reallocates routes and priorities to the agents that are congested, and then sends the adjustment results down for execution.
[0099] (3) Shelter Changes: For example, if a shelter suddenly becomes unavailable due to capacity or security reasons (exceeding capacity or being subjected to secondary threats), agents originally planning to go there need to be reassigned to other shelters. The system treats these agents as a task subset, re-executes the goal allocation and path planning (similar to step S5) to find alternative safe havens for them as quickly as possible, and notifies the relevant leadership agents of the adjusted destinations for implementation.
[0100] The aforementioned replanning process is handled as locally as possible, meaning that only the paths of affected agents (or groups) are adjusted, while other unaffected areas continue according to the original plan, in order to reduce global disturbances. However, if the accident escalates significantly (e.g., a secondary explosion occurs, requiring a city-wide change in evacuation direction), the system also has a contingency plan for global re-optimization: a new evacuation plan can be quickly recalculated in the background using previously trained models and current data, and all agents can be notified to switch plans via broadcast.
[0101] S64. Disturbance Tolerance and Flexible Scheduling: To maintain order amidst dynamic changes, we designed a flexible queue mechanism and fault-tolerant strategy for task scheduling. During scheduling, a certain time margin is reserved for each evacuation batch, allowing for slight fluctuations in actual departure or arrival times relative to the plan without triggering a chain reaction of conflicts. For example, a buffer time is provided between groups departing at different times; if the previous batch is slightly delayed, the subsequent batch can be appropriately delayed to avoid rear-end collisions along the way. The system treats evacuation tasks in each area as a priority queue, dynamically adjusting their execution order: if a more urgent situation occurs in a certain area (such as a sudden increase in radiation), the priority of that task is increased, and it is immediately executed; conversely, if progress in a certain area is temporarily hindered, tasks in other areas can be executed ahead of schedule to fully utilize road resources. This flexible adjustment of the task queue ensures that the overall evacuation rhythm remains stable under disturbances, somewhat similar to a sliding window in traffic control, avoiding both excessive idleness and congestion.
[0102] At the individual level, each leading agent typically prepares backup routes or contingency plans in addition to the original plan. If minor changes in road conditions occur during execution (e.g., slight congestion ahead but not serious impact), the leader can make minor adjustments within their authority without requiring central replanning—such as changing the team's speed or briefly using the emergency lane. Only when the deviation exceeds a threshold (e.g., complete blockage or emergency) is a new route requested. Following agents are also endowed with a certain degree of fault tolerance: if they temporarily fall behind or deviate from the team, they can catch up and rejoin based on the memorized route, while the leader will sense this through communication and wait or proceed slowly to reorganize the formation. This entire mechanism ensures that even in the event of disturbances, evacuation is not interrupted: either the disturbance is absorbed through local adjustments, or the impact is minimized by quickly switching tasks / routes.
[0103] In summary, this solution fully demonstrates adaptability and robustness during the execution phase. Real-time monitoring and multi-agent collaborative replanning make the solution a "living" plan, continuously modified and optimized as the environment changes. The leader-follower architecture plays a crucial role on-site—the leader responds promptly to changes and adjusts the team, while followers flexibly cooperate to ensure safety. Even in a highly dynamic scenario like this city with its complex road network and radioactive spread, this multi-agent system can consistently evolve towards overall evacuation optimization through pre-trained intelligent strategies and runtime collaborative adjustments, ensuring a safe and efficient evacuation process.
[0104] Example 2
[0105] Assuming a specific location is selected as the evacuation starting point in District A of a city, and multiple refuge locations are set around the city, this embodiment compares the path planning using the MD-DQN (Multi-Agent Dynamic Deep Reinforcement Learning) algorithm, the A* algorithm, and the Dijkstra algorithm under the same accident scenario, based on the map and road network data of District A. To ensure comparability, the evacuation starting point and refuge target points (edge of the evaluation area) are fixed, and the three algorithms' respective evacuation path schemes are generated under the same road network environment and accident conditions. Then, the path advantages and disadvantages of different algorithms are compared and evaluated from three aspects: total evacuation path length, total computation time, and cumulative radiation dose. The results are summarized in Table 1. The data in the table show that the path planned by the MD-DQN algorithm is on par with traditional route generation algorithms in terms of computation time and total evacuation path length, which can meet the requirements of timely evacuation, but its cumulative radiation dose is lower than that planned by the A* and Dijkstra algorithms. This indicates that while meeting efficiency requirements, the MD-DQN method effectively reduces the risk of radiation exposure during evacuation.
[0106] Table 1 Comparison of evacuation route schemes generated by three path planning algorithms
[0107] algorithm Total length of evacuation route (meters) Calculate the total time (seconds). Cumulative radiation dose (mSv) The algorithm of this invention (MD-DQN) 1500 1.8 620.73 A* Algorithm 1500 1.6 627.81 Dijkstra's algorithm 1500 2.0 653.92
[0108] The effect diagram of the embodiment is as follows: Figure 6As shown (the green route represents the route generated by the algorithm of this invention and its dynamic changes), the evacuation routes generated by the three algorithms exhibit significant differences in spatial distribution. The route planned by the MD-DQN algorithm avoids areas of high radiation pollution after the accident, prioritizing routes with lower radiation doses and relatively smooth traffic flow; effectively avoiding prolonged stays in high-radiation areas and ensuring the safety of personnel during evacuation. Simultaneously, the MD-DQN path appropriately avoids main roads that may become congested after the accident, alleviating traffic pressure by flexibly adjusting the driving route, thus balancing safety with a certain level of traffic efficiency. In contrast, the traditional A* algorithm and Dijkstra's algorithm, because they only plan based on the static shortest path criterion and do not fully consider dynamic radiation diffusion and traffic congestion factors, generate routes that partially traverse areas with high radiation intensity and continue to use conventional main roads in the early stages of the accident. This results in evacuation vehicles receiving higher cumulative radiation doses in these areas and may increase travel time due to entering congested sections. As the above comparison demonstrates, the MD-DQN algorithm exhibits superior performance and reliability in dynamic evacuation route planning for nuclear accidents. On one hand, the multi-agent collaborative decision-making based on deep reinforcement learning gives its route planning greater flexibility and adaptability, enabling dynamic route adjustments based on real-time changes in pollution distribution and road conditions. On the other hand, the MD-DQN method introduces a safety redundancy mechanism in its planning, tending to avoid potentially high-risk points and bottleneck sections on a single path, thereby improving the robustness and safety margin of the overall evacuation plan. Conversely, traditional algorithms such as A* and Dijkstra lack responsiveness to dynamic environmental changes, have fixed route selections, and limited safety margins, making it difficult to provide optimal evacuation routes in complex scenarios such as sudden nuclear accidents. These results fully demonstrate the significant advantages of the MD-DQN route planning method described in this invention compared to traditional algorithms in emergency evacuation during sudden nuclear accidents.
[0109] As can be seen from the above embodiments, the method of the present invention can autonomously formulate feasible personnel evacuation route plans based on the dynamic radiation threat and traffic conditions at the nuclear accident site. For example, in Embodiment 2, the strategy of the present invention automatically calculates the proportion of diversions along different routes, making full use of available road resources while avoiding high-dose risks, achieving better auxiliary decision-making than traditional algorithms. In addition, when environmental changes exceed expectations (such as the occurrence of dynamic changes and adjustments in the environment, routes, actors, and targets), the system can also dynamically adjust its strategy through real-time monitoring and feedback, demonstrating good robustness and adaptability.
Claims
1. A multi-agent dynamic core emergency evacuation path planning method based on deep reinforcement learning, characterized in that, Comprise the following specific steps: S1, environmental data acquisition and state initialization: collect multi-source data of nuclear accident emergency scene, establish the initial environmental state model of the accident affected area, and set the initial distribution and evacuation target of personnel or vehicles; S2, multi-agent system establishment and scheduling: abstract the objects to be evacuated into multiple agents, and construct a multi-agent cooperative evacuation model; S3, path game modeling and road network competition analysis: based on the environmental state model, a multi-agent path game model is established, and the evacuation route selection of each agent is modeled as a game decision on shared road resources; accordingly, the interaction between agents in a dynamic polluted environment and limited road network resources is described; S4, deep reinforcement learning strategy solving: a deep reinforcement learning algorithm is used to solve the multi-agent path game model; S5, dynamic cooperative path planning: the trained reinforcement learning strategy is applied to the actual nuclear accident emergency scene, and a cooperative evacuation route scheme is generated for all agents according to the current real-time state; S6, scheme execution and dynamic adjustment: the optimized evacuation route scheme is published to the emergency command system or navigation terminal to guide each agent to evacuate according to the allocated path; the change of environmental state is continuously monitored and the state model is updated in real time during the evacuation process.
2. The multi-agent dynamic core emergency evacuation path planning method based on deep reinforcement learning according to claim 1, characterized in that: The multi-source data collected in step S1 includes accident initial information, real-time nuclear pollution diffusion monitoring data, meteorological data and regional road traffic network data; The multi-source data includes nuclear accident source information and accident type information, real-time dose rate data of radiation monitoring stations, meteorological forecast data, and regional road network and traffic flow data; by inputting the above data into the pre-constructed nuclear pollution diffusion model and traffic model, the pollution intensity distribution and road capacity of the accident affected area at each time step are predicted, and are part of the initial environmental state.
3. The multi-agent dynamic core emergency evacuation path planning method based on deep reinforcement learning according to claim 1, characterized in that: In step S2, the agents are grouped and priority scheduled, including dividing the evacuees into different regional agent groups according to geographical location, or setting different priorities according to personnel categories; through the scheduling control module, the starting evacuation time window of each group of agents is allocated, and the evacuation is started staggeredly to reduce the peak load of the road in the same time period.
4. The multi-agent dynamic core emergency evacuation path planning method based on deep reinforcement learning according to claim 1, characterized in that: In step S3, the multi-agent path planning is modeled as a non-cooperative game model containing road occupancy competition: the benefit / cost function of each agent is defined, including travel time, radiation dose on the way, road congestion cost factors, and the road occupancy competition model is introduced to make the number of agents on the same road section affect the travel cost.
5. The multi-agent dynamic core emergency evacuation path planning method based on deep reinforcement learning according to claim 1, characterized in that: In step S4, a deep neural network strategy model is constructed with multi-dimensional environmental state as input and agent action decision as output, and a reward function for evacuation tasks is designed; the training process adopts a combination of centralized training and decentralized decision making, so that each agent learns the optimal or suboptimal path selection strategy for collaboration; The deep reinforcement learning algorithm uses a multi-agent reinforcement learning framework to solve the approximate equilibrium solution of the above path game model, including: The decision-making strategy of each agent is represented as a deep neural network with shared parameters, the input of which is a multi-dimensional tensor describing the entire environment state, and the output of which is the action decision that each agent can take in the current state; A reward function is designed for the evacuation task, which includes risk avoidance indicators, accessibility indicators, and balance indicators. The multi-objective optimization is converted into a scalar reward signal by weighted summation. Positive rewards are given when an agent successfully reaches a safe area, and penalties are given if it stays in a high-risk area for a long time or causes severe congestion.
6. The multi-agent dynamic core emergency evacuation path planning method based on deep reinforcement learning according to claim 5, characterized in that: The deep reinforcement learning algorithm uses a policy gradient algorithm or a value function algorithm to optimize multi-agent strategies. The training and execution architecture is centralized: in the training phase, a centralized evaluator is introduced to obtain the global state and reward information of all agents, and the joint update of each agent's policy gradient is performed; in the execution phase, each agent only makes independent decisions based on its own local observations and the trained strategy, thereby ensuring agent cooperation while improving the feasibility and robustness of deployment; during training, the strategies of other agents are considered as part of the environment, and an adaptive iterative method is used to approximate the Nash equilibrium solution of the game, ensuring the stability of the resulting strategy under multi-agent interaction.
7. The multi-agent dynamic core emergency evacuation path planning method based on deep reinforcement learning according to claim 5, characterized in that: The environment state is input into a deep neural network in the form of a multi-dimensional tensor, which includes at least the following levels or channels: Pollution diffusion state layer: represents the radiation dose rate or pollutant concentration at each geographic location in the accident area in a grid form, as well as the dynamic change information over time, which can be composed of a three-dimensional spatio-temporal tensor by stacking pollution distribution at multiple time points; Road network traffic state layer: represents the traffic state of road network nodes and sections in a manner aligned with the above-mentioned geographic grid, including whether the road is unobstructed, the current road section vehicle occupancy rate or speed information; Agent distribution layer: represents the current position distribution of each agent and the target position distribution.
8. The multi-agent dynamic core emergency evacuation path planning method based on deep reinforcement learning according to claim 1, characterized in that: During the training process of the deep reinforcement learning strategy, rule-based strategies and multi-strategy coordination mechanisms are integrated, including: Introducing expert rules and safety constraints: setting rule masks for infeasible or high-risk actions; Multi-strategy fusion training: using heuristic strategies to guide agent behavior at the beginning of reinforcement learning training as a strategy initialization or auxiliary reference; gradually increasing the proportion of autonomous exploration of agents as training progresses, and using strategy entropy regulation or mixed strategy gradient methods to smoothly transition between rule-based strategies and learned strategies; Supporting offline pre-training and online fine-tuning: first pre-train the strategy using historical accident data and simulation environments to shorten the model convergence time; when actually deployed, make small online updates combined with real-time monitoring data, use low-exploration-rate strategies to fine-tune to adapt to the current accident scenario, and ensure the reliability of the decision-making result by using safety threshold control to ensure that strategy updates do not cause drastic behavior changes.
9. A deep reinforcement learning based multi-agent dynamic core emergency evacuation path planning system, characterized in that: The system includes a modular unit for executing the method of any one of claims 1 to 8, and at least includes: The environmental data acquisition module is configured to receive and collect multi-source environmental data, including nuclear accident monitoring and forecasting data, road traffic state data and personnel position data, input the data into an environmental state model for preprocessing, and generate a multi-dimensional data set describing the current state of the accident-affected area; The state modeling and updating module is configured to generate and maintain a multi-dimensional state tensor model of the nuclear emergency scene according to the environmental data, update the state information such as pollution diffusion, road traffic condition and agent position in real time, and provide the state information for the decision-making module; The strategy decision-making module is configured to internally embed a deep reinforcement learning strategy model, calculate the path planning decision scheme of each agent based on the current state tensor, give the next action or complete path suggestion to each agent by using the trained multi-agent collaborative strategy, and form a globally coordinated evacuation scheme by comprehensively considering the decisions of all agents; The coordination and scheduling module is configured to detect and optimize the multi-agent path scheme output by the strategy decision-making module, including identifying possible road resource conflicts or overload conditions, and adjusting the path selection or start time of part of the agents to relieve congestion, so as to adjust the scheme to a final route scheme meeting the balance condition; The output and communication module is configured to publish the finally determined evacuation path scheme to emergency management personnel, rescue teams and public evacuation navigation terminals through a visual interface or a communication network, including marking the recommended route on an electronic map, sending evacuation instructions in batches, and receiving feedback information in the execution process and feeding back the feedback information to the state updating module for closed-loop adjustment; The system further comprises auxiliary modules related to deep reinforcement learning: The strategy training and management module is configured to manage the training process and version update of the deep reinforcement learning strategy, save the pre-trained strategy model under different accident scenarios, trigger an online training fine-tuning mechanism when the environment exceeds the training data distribution, and be responsible for the permission management and security audit of the strategy parameters; The experience data storage module is configured to store the state-decision-feedback data generated by the multi-agent in the simulation training and real execution process, support the strategy training module to sample the data for further training and optimization, and provide data support for post-accident analysis; The rule engine module is configured to store and execute predefined expert rules and safety constraints, filter and constrain the candidate actions output by the strategy decision-making module in the decision-making process, modify or prohibit the decisions when safety rules are violated, and ensure the safety compliance of the system decision results.
10. The deep reinforcement learning based multi-agent dynamic core emergency evacuation path planning system according to claim 9, characterized in that: The system is deployed in the information platform of the nuclear emergency command center as an independent decision support unit and integrated with other emergency response modules; The interaction process among the modules is as follows: the environmental data acquisition module sends real-time monitoring and prediction data to the state modeling and updating module, generates a current environmental state tensor and transmits the state tensor to the strategy decision-making module, the strategy decision-making module calculates an initial multi-agent path scheme based on the state tensor, the path scheme is adjusted and optimized by the coordination and scheduling module, and the execution is carried out by the output and communication module. During the execution, the environmental data acquisition module continuously collects new data and updates the state model, the experience data storage module records the action trajectory and environmental feedback of each agent, and the strategy training and management module is used for accident process analysis and periodic correction and upgrading of the strategy model, so as to realize the closed-loop optimization of nuclear emergency evacuation decision.
Citation Information
Cited By
Multi-agent emergency evacuation method in dynamic radiation scene of nuclear accident and storage medium
CN122114322A
Multi-agent nuclear accident dynamic emergency evacuation method and storage medium
CN122155060A