Evacuation path optimization method based on coupling numerical simulation and hierarchical reinforcement learning

By constructing a high-precision coupled numerical simulation framework and a hierarchical deep reinforcement learning framework, combined with inverse distance weighted interpolation and graph neural network model, the problems of high computing resource demand, low simulation efficiency and insufficient disaster avoidance path planning in numerical simulation of storm surges and emergency evacuation are solved, and efficient and refined risk division and dynamic evacuation path optimization are achieved.

CN120471247APending Publication Date: 2025-08-12SHENZHEN UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510546749.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-12

Smart Images

  • Figure CN120471247A_ABST
    Figure CN120471247A_ABST
Patent Text Reader

Abstract

The invention provides an evacuation path optimization method based on coupling numerical simulation and hierarchical reinforcement learning, and relates to the field of storm surge, and the method comprises the steps: constructing a coupling numerical simulation framework of high-precision storm surge; performing storm surge numerical simulation through a coupling numerical simulation framework to generate a submerged depth data set; performing high-resolution submerged depth interpolation through the submerged depth data set in combination with an inverse distance weighted interpolation algorithm to generate a high-resolution submerged depth interpolation data set; performing risk grading on the high-resolution submerged depth interpolation data set to generate a risk zoning map; generating an evacuation road network through a path planning neural network model based on a graph neural network and the risk zoning map; an evacuation strategy of an evacuation road network is optimized by adopting a hierarchical deep reinforcement learning framework, a combined optimal dynamic evacuation strategy is generated, and the generalization performance and the calculation efficiency of evacuation path planning are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of storm surges, and in particular to an evacuation path optimization method based on coupled numerical simulation and hierarchical reinforcement learning. Background Art

[0002] Storm surge is one of the major marine disasters faced by my country's coastal areas. Its powerful destructive power and wide impact range have caused huge casualties and economic losses to coastal areas.

[0003] In order to achieve refined prevention and control of storm surge disasters and emergency management, it is crucial to build a high-precision storm surge numerical simulation and forecasting system. At present, the coupling method of the WRF (Weather Research and Forecasting) atmospheric model and the ADCIRC (Advanced Circulation) and SWAN (Simulating Waves Nearshore) ocean models is widely used at home and abroad to numerically simulate storm surges. Through the accurate description of the interaction between the sea and the air and the collaborative solution of multi-scale processes, high-fidelity reconstruction of multi-physical fields such as storm surge water level increase, nearshore waves, and storm surge inundation can be achieved. On this basis, by performing high-resolution inundation depth interpolation on the storm surge numerical simulation results and combining them with GIS spatial analysis technology, a series of products required for emergency management can be quickly generated, such as risk area demarcation maps and evacuation route planning maps.

[0004] However, the existing storm surge numerical simulation and emergency evacuation technologies still have several urgent problems to be solved in practical applications. First, storm surge disasters involve multi-scale physical processes from atmospheric circulation to ocean eddies. Carrying out high-resolution, large-area coupled numerical simulations requires extremely high computing resources. The existing simulation efficiency is difficult to meet the needs of rapid simulation of ultra-large-scale, high-resolution scenarios, affecting the timeliness of forecast results. Secondly, due to the limitations of model resolution and observational data assimilation capabilities, the temporal and spatial resolution of existing storm surge inundation simulations is not fine enough, and it is difficult to fully meet the requirements of disaster prevention and mitigation departments for the timeliness and spatial refinement of risk information. Finally, in terms of emergency evacuation of storm surge disasters, how to achieve integrated and dynamic optimization decisions from regional personnel evacuation plans to specific disaster avoidance route planning based on the population distribution characteristics, important facility layout, and road network conditions in the risk area remains a challenging problem.

[0005] Traditional disaster avoidance route planning methods based on the shortest path are often unable to effectively cope with complex situations such as ultra-large-scale road networks and dynamic traffic flow evolution, and lack the necessary adaptive and real-time adjustment capabilities.

[0006] To address these challenges, significant breakthroughs are urgently needed in improving the efficiency of storm surge numerical simulations, refining storm surge risk analysis, and innovating adaptive disaster avoidance path planning. This is crucial for enhancing the scientific, refined, and intelligent level of storm surge emergency management in my country's coastal areas, minimizing disaster losses, and protecting the lives and property of the people. Summary of the Invention

[0007] The purpose of the present invention is to provide an evacuation path optimization method based on coupled numerical simulation and hierarchical reinforcement learning in order to solve the problems that the existing emergency evacuation planning has low spatial resolution and disaster avoidance path planning is difficult to apply to dynamic traffic flows and large-scale complex road network environments.

[0008] The above-mentioned purpose of this application is achieved through the following technical solutions:

[0009] S1: Build a high-precision coupled numerical simulation framework for storm surges, which includes the WRF atmospheric model, the ADCIRC model, and the SWAN ocean model.

[0010] S2: Conduct storm surge numerical simulations through a coupled numerical simulation framework to generate an inundation depth dataset;

[0011] S3: Perform high-resolution inundation depth interpolation using the inundation depth dataset and the inverse distance weighted interpolation algorithm to generate a high-resolution inundation depth interpolation dataset; perform risk classification on the high-resolution inundation depth interpolation dataset to generate a risk zoning map;

[0012] S4: Generate an evacuation road network through a graph neural network-based path planning neural network model and a risk zoning map;

[0013] S5: Adopt a hierarchical deep reinforcement learning framework to optimize the evacuation strategy of the evacuation road network and generate a joint optimal dynamic evacuation strategy.

[0014] An electronic device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device performs an evacuation path optimization method based on coupled numerical simulation and hierarchical reinforcement learning.

[0015] A computer-readable storage medium stores instructions. When the instructions are executed, an evacuation path optimization method based on coupled numerical simulation and hierarchical reinforcement learning is performed.

[0016] The beneficial effects of the technical solution provided by this application are:

[0017] 1. Construct a high-precision coupled numerical simulation framework for storm surges. This framework employs a parallel strategy that combines domain decomposition with task decomposition. It uses METIS to divide non-overlapping subdomains and introduces ghost cells to facilitate data exchange. WRF, ADCIRC, and SWAN each employ different parallel strategies, transferring multi-physics variables via couplers. Furthermore, multi-threaded parallel simulation is employed within non-overlapping subdomains, with optimizations such as non-blocking communication and I / O overlap implemented to achieve efficient numerical simulation of storm surge scenarios.

[0018] 2. A new method for high-resolution interpolation and refined risk grading of storm surge inundation has been developed. Using an inverse distance weighted interpolation algorithm, inundation depths are interpolated from an unstructured source grid to a 10-meter resolution structured target grid. Furthermore, taking into account the characteristics of the hazard-prone areas and historical disaster experience, a four-level risk zoning system (red, orange, yellow, and blue) has been designed. Through grid data threshold segmentation and a local correction principle based on "higher rather than lower" risk, this system automatically delineates risk zones, providing refined risk information support for coastal disaster prevention and mitigation.

[0019] 3. Propose a method for intelligently representing road networks and optimizing evacuation strategies. A graph neural network architecture consisting of three modules, node embedding, edge embedding, and knowledge fusion, is designed. Through graph attention mechanisms and knowledge graph embedding, it effectively learns the spatiotemporal dependencies of road networks. Based on this, a hierarchical deep reinforcement learning framework is constructed, employing a centralized multi-agent reinforcement learning model. Using a multi-agent deep deterministic policy gradient algorithm with an actor-critic architecture, this method achieves decoupled optimization of regional evacuation plans and segment-level path planning. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The present application will be further described below with reference to the accompanying drawings and embodiments, in which:

[0021] Figure 1 It is a step diagram in the embodiment of the present application;

[0022] Figure 2 It is a schematic diagram of the structure of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION

[0023] In order to have a clearer understanding of the technical features, purposes and effects of this application, the specific implementation methods of this application are now described in detail with reference to the accompanying drawings.

[0024] The embodiments of the present application provide an evacuation path optimization method based on coupled numerical simulation and hierarchical reinforcement learning.

[0025] Please refer to Figure 1 , Figure 1 This is a step diagram of an evacuation path optimization method based on coupled numerical simulation and hierarchical reinforcement learning in an embodiment of the present application, including:

[0026] S1: Build a high-precision coupled numerical simulation framework for storm surges, which includes the WRF atmospheric model, the ADCIRC model, and the SWAN ocean model.

[0027] As an embodiment, a two-way coupling method of the WRF atmospheric model, ADCIRC and SWAN ocean models is adopted to achieve high-precision and high-resolution numerical simulation of storm surges. The WRF model provides atmospheric driving conditions, ADCIRC simulates storm surge water increase, and SWAN simulates nearshore winds and waves. In terms of spatial discretization, a 30-50m high-resolution unstructured triangular grid is used in the nearshore area, and a kilometer-level low-resolution grid is used in the farshore area. In terms of time integration, a combination of semi-implicit, explicit and adaptive step sizes is adopted, with an average time step of 30-60s. Based on historical typhoon best path data, 1,000 sets of storm surge numerical simulation conditions covering typical to extreme intensities are generated.

[0028] S2: Conduct storm surge numerical simulations through a coupled numerical simulation framework to generate an inundation depth dataset;

[0029] As an example, based on the coupled simulation framework constructed in step 1, a parallel strategy combining domain decomposition and task decomposition is adopted to distribute computing tasks to multiple processor nodes. METIS is used to divide non-overlapping sub-areas, and ghost units are introduced to achieve data exchange. WRF, ADCIRC, and SWAN each adopt different parallel strategies to transfer multi-physics field variables through couplers. Multi-threaded parallel simulation is adopted within non-overlapping sub-areas, and optimization measures such as non-blocking communication and IO overlap are introduced.

[0030] S3: Perform high-resolution inundation depth interpolation using the inundation depth dataset and the inverse distance weighted interpolation algorithm to generate a high-resolution inundation depth interpolation dataset; perform risk classification on the high-resolution inundation depth interpolation dataset to generate a risk zoning map;

[0031] As an example, an inverse distance weighted interpolation algorithm is used to interpolate flood depths onto a 10-meter resolution structured grid. A four-level risk grading system ("red-orange-yellow-blue") is designed. Through threshold segmentation and local correction of grid data, risk zones are automatically delineated. Thematic maps of risk zones are generated, and risk statistics are calculated to produce a comprehensive assessment report.

[0032] S4: Generate an evacuation road network through a graph neural network-based path planning neural network model and a risk zoning map;

[0033] As an example, the evacuation road network is abstracted into a multi-attribute weighted directed graph, and a neural network architecture is designed, comprising three modules: node embedding, edge embedding, and knowledge fusion. A graph attention network is used to learn low-dimensional, dense representations of network nodes, and a self-attention mechanism is employed to implement edge embedding. Knowledge graph embedding technology is introduced to incorporate empirical knowledge from historical evacuation cases. A neural network architecture search is used to optimize the model structure, improving inference efficiency in ultra-large-scale road network scenarios. This method provides an intelligent network representation foundation for subsequent evacuation strategy optimization.

[0034] S5: Adopt a hierarchical deep reinforcement learning framework to optimize the evacuation strategy of the evacuation road network and generate a joint optimal dynamic evacuation strategy.

[0035] As an example, a centralized multi-agent reinforcement learning model is constructed to model the evacuation problem as a multi-agent decision-making process. The state space, action space, and reward function are designed, taking into account regional population distribution, shelter layout, and real-time road conditions. A multi-agent deep deterministic policy gradient algorithm with an actor-critic architecture is used for policy learning to generate a dynamic evacuation plan that considers global coordination. This method adaptively optimizes regional evacuation plans and segment-level path planning, minimizing total evacuation time and maximizing shelter coverage.

[0036] Step S1 includes:

[0037] As an embodiment, a high-precision, high-resolution numerical simulation of storm surges is achieved by using a two-way coupling method of the WRF (Weather Research and Forecasting) atmospheric model, ADCIRC (Advanced Circulation) and SWAN (Simulating Waves Nearshore) ocean model.

[0038] The OASIS3-MCT coupler is used to control the WRF atmospheric model, ADCIRC model and SWAN ocean model, and the coupling variables are exchanged within a specified time interval to achieve

[0039]

[0040]

[0041] (ζ a ):ADCIRC→OASIS→(ζ w ):WRF

[0042] Among them, OASIS is a general coupling framework responsible for controlling the exchange and synchronization of coupling variables between different modes; and P wrepresent the wind field and pressure field provided by the WRF model respectively; P a , P s represent the upper boundary wind field and pressure field driving of ADCIRC model and SWAN model respectively; and represents the wave field transferred from the SWAN mode to the ADCIRC mode; ζ a and ζ w Indicates the water level for transferring from ADCIRC mode to SWAN mode.

[0043] As an example, WRF provides wind and pressure fields to ADCIRC and SWAN, SWAN provides radiation stress to ADCIRC, and ADCIRC feeds back tidal changes to WRF. OASIS3-MCT is responsible for interpolation mapping and block transmission between different grids, enabling data exchange and synchronization between models.

[0044] As an example, in the time step setting of numerical solution and mode coupling, ADCIRC adopts the finite element method for spatial discretization and semi-implicit Crank-Nicolson time integration with a time step of 100-200s; SWAN adopts the finite volume method for spatial discretization and explicit Euler forward time integration with a time step of 1-10s; WRF adopts Arakawa-C grid finite difference and explicit Runge-Kutta time integration with a time step of 30-60s.

[0045] As an embodiment, the finite element method is based on the variational principle. By selecting a suitable trial function, the continuous problem is converted into a discrete algebraic equation system, which has good numerical stability and computational accuracy. The Crank-Nicolson time integral is a semi-implicit algorithm that couples the decomposition of the current moment and the previous moment by linear interpolation, which can improve computational efficiency while ensuring stability. In terms of spatial grid division, a 30-50m high-resolution unstructured triangular grid is used in the nearshore area of interest, and a thousand-meter low-resolution unstructured triangular grid is used in the offshore deep-sea area, achieving a balance between computational accuracy and efficiency. The unstructured triangular grid has geometric flexibility and can better adapt to complex coastlines and terrains, while facilitating the implementation of local encryption and adaptive grids. The grid generation process usually adopts the Delaunay triangulation algorithm to improve the accuracy and stability of the numerical solution by optimizing the grid quality (such as maximum angle, minimum angle, area, etc.).

[0046] Step S2 includes:

[0047] Step S2 includes: S21: Extract the key parameters of tropical cyclones that land on China from the typhoon best track database. The key parameters include: time of landfall, latitude and longitude of the landing point, central pressure, maximum wind speed, moving speed, and moving direction. As an example, in order to generate storm surge scenarios (flooding depth datasets) covering typical to extreme intensity levels, this patent extracts key element parameters of tropical cyclones that have landed on Chinese land since 1949 from the National Typhoon Optimal Path Database released by the Tropical Cyclone Data Center of the China Meteorological Administration, including: landing time; , longitude and latitude of landing point , central air pressure , maximum wind speed , Movement Speed , moving direction etc., as shown in Table 1.

[0048] Table 1 Optimal track parameters of tropical cyclones making landfall in China

[0049] S22: Based on the statistical analysis of key typhoon parameters, the probability distribution characteristics are obtained;

[0050] Using the Monte Carlo method, we sampled several typhoon path and intensity combinations from the probability distribution characteristics to obtain the storm surge numerical simulation conditions.

[0051] As an example, a Monte Carlo method was used to randomly sample a large number of typhoon path and intensity combinations from these probability distributions. Physically unreasonable or duplicate combinations were eliminated, ultimately resulting in 1,000 sets of storm surge numerical simulation conditions, covering a range of typical to extreme intensity levels. These conditions encompass variations in multiple dimensions, including typhoon landfall location, intensity level, and movement speed, ensuring a certain degree of representativeness and completeness.

[0052] S23: For each set of storm surge numerical simulation conditions, the Holland wind field model is used to generate the spatiotemporal distribution characteristics of the typhoon wind field and pressure field according to the typhoon optimal path parameters;

[0053] Tangential wind speed V of the Holland wind field model t And the pressure field P satisfies:

[0054]

[0055] Where r is the radial distance from the typhoon center; A and B are dimensionless parameters related to the typhoon's intensity and structure; ρ a is the air density; p n Indicates the typhoon center pressure; p eIndicates the ambient air pressure value outside the typhoon; r m is the maximum wind speed radius;

[0056] The optimal typhoon path parameters include: typhoon center location, typhoon center pressure, and typhoon maximum wind speed radius;

[0057] As an example, commonly used empirical formulas include the Rankine vortex model, the Schloemer model, the Holland model, etc. They are based on the symmetry assumption of typhoons. Through a few parameters (such as maximum wind speed radius, central pressure difference, etc.), the analytical expressions of the wind field and pressure field can be determined to generate wind fields and pressure fields that are consistent with observations.

[0058] S24: Interpolate the spatiotemporal distribution characteristics of typhoon wind and pressure fields onto the grid points of the WRF, ADCIRC, and SWAN marine meteorological models;

[0059] The interpolated grid points are fused with the background field to form a submerged depth dataset;

[0060] The background field uses reanalysis data or climate mean field;

[0061] Fusion methods include spectrum matching and optimal interpolation.

[0062] As an example, for marine meteorological models (WRF, ADCIRC, SWAN), it is necessary to discretize the study area through their respective specific grid systems to simulate relevant physical processes. The original resolution and grid division of typhoon wind field and pressure field data are inconsistent with the WRF, ADCIRC, and SWAN grid points. Therefore, interpolation methods are needed to interpolate the wind field and pressure field to the grid points of these marine meteorological models.

[0063] As an example, the background field uses reanalysis data or climate mean fields to ensure that the simulated conditions are embedded in the large-scale environmental field. Fusion methods include spectral matching and optimal interpolation, aiming to smoothly transition to the background field while maintaining the typhoon's intensity and structural characteristics.

[0064] Step S2 further includes:

[0065] A parallel strategy combining domain decomposition and task decomposition is adopted to carry out numerical simulation of storm surge.

[0066] As an example, by coupling the numerical simulation framework, in order to achieve efficient calculation of 1,000 typical to extreme intensity storm surge scenarios, a parallel strategy combining domain decomposition and task decomposition is adopted to distribute the computational tasks of storm surge numerical simulation to multiple processor nodes, thereby realizing the rapid generation of massive scenarios.

[0067] As an example, the three-dimensional computational domain Ω is spatially discretized to generate an unstructured hybrid grid. In the horizontal direction, a triangular grid is used in the nearshore area to meet the requirements for resolving terrain and features; a quadrilateral grid is used in the offshore area to improve computational efficiency. In the vertical direction, the σ coordinate transformation is used:

[0068]

[0069] Where z is the vertical physical coordinate, ζ is the free surface elevation, and H is the still water depth. The σ coordinate can better adapt to irregular terrain and liquid level changes.

[0070] Assuming the number of horizontal grid cells is N and the number of vertical σ layers is M, the total number of cells is M·N, and the number of nodes is about 6 times that. Considering the high requirements of storm surge numerical simulation on terrain resolution, a 50m high-resolution unstructured triangular grid is used to discretize the nearshore area, a 100m resolution unstructured quadrilateral grid is used to discretize the offshore area, and 100 vertical layers of σ coordinates are used. The total degree of freedom scale reaches O(10 9 ) magnitude.

[0071] Based on the grid topology, METIS software package is used to divide Ω into L non-overlapping sub-regions Ω. i , i=1,2,…,L, the division principle is:

[0072] Subregion Ω i The number of grid cells should be as equal as possible to ensure load balancing;

[0073] Subregion Ω i The boundary Γ i Be as regular as possible to reduce communication overhead;

[0074] The number of sub-regions L should match the total number of parallel computing cores P, generally L = O(P).

[0075] The number of sub-regions L is estimated by the following empirical formula:

[0076] Where P is the total number of parallel cores, and dof is the degrees of freedom of a single storm surge scenario. For the parallel computation of 1000 scenarios, this patent selects a total number of cores P = 512, and the number of sub-regions L = 32.

[0077] Then, in each subregion Ω i The boundary Γ i On the top, a layer of ghost units is introduced to connect with the adjacent sub-region Ω j The boundary cells of (j≠i) overlap, which is used for data exchange between sub-regions. The physical quantities of ghost cells are explicitly provided by the adjacent sub-regions.

[0078] For the i-th subregion Ω i, the semi-discrete finite volume format of the storm surge governing equation is:

[0079]

[0080] Where k is Ω i Local number of internal unit; V ik is the volume of unit k; F ijk is the interface flux between k and its adjacent unit j; S ik is the source-sink term on unit k; N(k) is the set of units adjacent to k, including ghost units.

[0081] Boundary flux F ijk Using Riemann approximation: F ijk =F(V ik ,V jk ,n ij )

[0082] Among them, n ij is the interface normal vector. The Riemann flux is solved using the HLLC format, which can effectively capture discontinuous structures such as storm surge fronts and tide ranges. ik Including influences such as seabed topography, Coriolis force, sea surface wind stress, and wave radiation stress.

[0083] Considering that WRF, ADCIRC, and SWAN use different spatiotemporal discretization methods in different regions, this patent adopts a parallel framework of non-overlapping grid coupling to achieve parallelization in the grid subdivision of each mode. The number of processes is distributed as shown in Table 2.

[0084] The WRF mode adopts the parallel strategy of longitude and latitude blocks, setting N W MPI processes;

[0085] The ADCIRC mode adopts the parallel strategy of finite element sub-domain decomposition, setting N A MPI processes;

[0086] The SWAN mode adopts the parallel strategy of finite element sub-domain decomposition, setting N S MPI processes.

[0087] Table 2. Coupling model process number distribution table

[0088] Mode Name Number of parallel processes Allocate cores effect WRF 4 32 Atmospheric driving ADCIRC 32 320 Storm surge SWAN 16 160 nearshore waves

[0089] On this basis, the key steps to achieve WRF-ADCIRC-SWAN coupled parallel computing are as follows:

[0090] Step (1): Distribute the meshing results and initial boundary conditions of WRF, ADCIRC, and SWAN to the corresponding N W 、N A 、N S MPI processes;

[0091] Step (2): Each process calculates the sub-region it is responsible for, controlled by the local time step, first calculating the internal unit and then the ghost unit;

[0092] Step (3): Through parallel I / O in PnetCDF format, the interaction of wind field, pressure field and tide feedback between WRF and ADCIRC, as well as the radiation stress transfer between SWAN and ADCIRC are realized;

[0093] Step (4): Each process packages the physical quantity data on the boundary and transmits it through MPI Isend and MPI Irecv Wait for non-blocking communication functions to send to the ghost unit of the adjacent process, and receive updates to its own ghost unit;

[0094] Step (5): Repeat steps (2)-(4) until all time steps are calculated;

[0095] Step (6): Each process aggregates the sub-region solutions, splices them into a global solution and outputs it through parallel I / O.

[0096] Step S3 includes:

[0097] As an example, after completing the high-resolution parallel simulation of 1,000 storm surge scenarios in step 2, a massive inundation depth dataset was obtained. To refine storm surge warnings, it is necessary to use interpolation technology based on the original inundation depth results generated by numerical simulation to generate a high-resolution inundation depth interpolation dataset that meets the temporal and spatial resolution requirements of the warning. Based on this, risk grading can be carried out to form an intuitive risk zoning map, providing more refined and dynamic risk information support for coastal disaster prevention and mitigation decision-making.

[0098] As an example, spatial interpolation of flooding depth refers to the process of interpolating numerical simulation results from an unstructured source grid (typically a triangular or quadrilateral grid) to a high-resolution structured target grid (typically a regular rectangular grid). This algorithm uses an inverse distance weighted interpolation algorithm, which assumes that the value of the estimated point can be represented by a weighted average of the known points in its neighborhood, with the weight being inversely proportional to the distance.

[0099] The specific steps of spatial interpolation of high-resolution inundation depth include: interpolating numerical simulation results from an unstructured source grid to a high-resolution structured target grid; extracting numerical simulation results from an inundation depth dataset; the unstructured source grid is a triangular grid or a quadrilateral grid; the high-resolution structured target grid is a regular rectangular grid;

[0100] Assume D s (i,j) and D t (p,q) are the flooding depths of the source grid cell (i,j) and the target grid cell (p,q), and their spatial coordinates are (x i ,y j ) and (X p ,Y q );

[0101] Let w k is the weight of the interpolation of the kth source grid cell to the target grid cell (p,q), satisfying the normalization condition ∑w k =1, then the inverse distance weighted interpolation formula is:

[0102]

[0103] Among them, K is the total number of source grid cells involved in interpolation; weight w k Defined using the inverse distance function; represents the nth power of the Euclidean distance between the center of the kth source grid cell and the center of the target grid cell; d k is the Euclidean distance between the center of the kth source grid cell and the center of the target grid cell, and n is the inverse distance weighted index; It represents the nth power of the Euclidean distance between the center of the lth source grid cell and the center of the target grid cell;

[0104] As an embodiment, after obtaining high-resolution inundation depth data using the above-mentioned interpolation method, it is necessary to further carry out risk grading to intuitively reveal the geographical spatial risk pattern of storm surge inundation. This requires a full understanding of the disaster intensity of different inundation depths and the reasonable setting of risk level division thresholds. This patent comprehensively considers the characteristics of the hazard-bearing bodies in my country's coastal areas, historical storm surge disaster experience, and the national standard "Storm Surge Disaster Classification", takes inundation depth as the key indicator for risk grading, and constructs a "red-orange-yellow-blue" four-level risk zone division system.

[0105] The risk zoning map includes: blue zone, yellow zone, orange zone and red zone;

[0106] Blue zone: 0.2m≤submergence depth<0.5m, slight impact on personnel and farmland, low risk level;

[0107] Yellow zone: 0.5m≤submergence depth<1.0m, travel is obstructed, crops are flooded on a large scale, and the risk level is medium;

[0108] Orange zone: 1.0m≤submergence depth<1.5m, building structures begin to be damaged, personnel safety is seriously threatened, and the risk level is high;

[0109] Red zone: Flooding depth ≥ 1.5m, large-scale building collapse, human life is in danger, and the risk level is extremely high;

[0110] As an example, the above thresholds are primarily determined based on the degree of damage to different hazard-bearing structures caused by storm surge inundation. For example, when the inundation depth exceeds 0.2m, large areas of farmland begin to become waterlogged, and some roads begin to accumulate water. When it exceeds 0.5m, crop yields drop by more than half, and urban and rural transportation is disrupted. When it exceeds 1.0m, water enters houses, causing structural damage and posing a serious threat to human safety. When it exceeds 1.5m, widespread building collapse occurs, making it difficult to protect lives. These thresholds effectively quantify the intensification of inundation disasters, reflecting the differences in risk levels and the gradual increase in disaster intensity.

[0111] As an example, after clarifying the classification criteria, it is necessary to further match the classification threshold to the high-resolution flooding depth interpolation data to achieve rapid identification and automatic division of risk areas. To this end, this patent designs a dynamic risk area division algorithm based on raster data, the key steps of which are as follows:

[0112] (1) Define the risk classification threshold vector. Based on the aforementioned four-level risk classification standard, define the risk classification threshold vector H = [0.2, 0.5, 1.0, 1.5], where H[k] represents the lower limit of the flooding depth of the k-th risk zone. For example, H[0] = 0.2 means that the lower limit of the flooding depth of the blue risk zone is 0.2m, and H[3] = 1.5 means that the lower limit of the flooding depth of the red risk zone is 1.5m.

[0113] (2) Threshold segmentation of flooding depth. Using the high-resolution flooding depth grid data D generated by interpolation as input, traverse each grid cell (i, j) within the study area and extract its flooding depth value D ij , judge D ij Which interval the threshold vector H falls into is used to determine the initial risk level R of the grid cell. ij For example, if 0.5≤D ij <1.0, then the initial risk level R of grid (i, j) ij =2, indicating a yellow risk area. This step achieves the threshold segmentation of the flooding depth grid into the graded risk area grid, which is the key to risk area delineation.

[0114] (3) Local correction of risk zones. Considering that risk zones of adjacent levels may be交错分布locally, it is necessary to locally correct the initial risk zone division results to simplify the risk zone boundaries and improve the mapping aesthetics. Specifically, the principle of "choosing the higher level instead of the lower level" is adopted. That is, if a grid cell (i, j) meets the inundation depth threshold conditions of multiple risk levels simultaneously, the highest level among them is taken as the final risk level R of this cell. ij , Mathematically expressed as: R ij = max{k | H[k] ≤ D ij < H[k + 1]}, k = 0, 1, 2, 3. For example, if there are individual orange grid cells distributed in the red area, their risk levels are raised to red to ensure the integrity of the high-risk red area. Conversely, if there are individual yellow grid cells distributed in the blue area, their risk levels are lowered to blue to ensure the continuity of the low-risk blue area. Although this local correction will slightly change the area of the risk zone, it can effectively eliminate the "noise" in the risk zone distribution and make the risk zone division more regular and reasonable.

[0115] Import the division results into the GIS platform, overlay thematic elements such as administrative divisions, key facilities, and evacuation routes to generate intuitive inundation risk distribution maps and risk zone thematic maps. In GIS, a red-orange-yellow-blue color scheme is usually adopted to represent the four risk levels, and the spatial distribution of risk zones of each level is presented in the form of a color patch map.

[0116] Step S4 includes: The path planning neural network model includes: a node embedding module, an edge embedding module, and a knowledge fusion module; the point embedding module, the edge embedding module, and the knowledge fusion module are connected in sequence.

[0117] As an embodiment, after obtaining the refined risk zoning of storm surge disasters, it is necessary to further plan scientific and efficient emergency evacuation routes for high-risk areas to minimize casualties and property losses. Traditional evacuation route planning methods are mainly based on shortest path algorithms such as the Dijkstra algorithm and the A* algorithm. However, such methods often have difficulty fully considering the spatio-temporal complexities such as road network topology and dynamic traffic flow evolution, and the obtained evacuation routes lack overall coordination and dynamic adaptability. To overcome the above challenges, this patent introduces the Graph Neural Network (GNN) technology. By abstracting the evacuation road network into a complex network with multiple attributes, it learns the spatial dependence relationships of different paths to form an intelligent evacuation route planning method.

[0118] It should be noted that the phrase "交错分布" in the original text seems to be a misspelling or an incomplete expression. I translated it as "interleaved distribution" according to the context, but it might need to be further confirmed based on the accurate meaning.The evacuation road network is abstracted as an undirected weighted graph G = (V, E, X, Z), where V is the set of road network nodes, including intersection nodes and refuge nodes; E is the set of road segment edges, indicating the existence of a connected road segment between node i and node j; X is the set of node attributes, including the population distribution, refuge capacity, and traffic status of each node; Z is the set of edge attributes, including the length, width, and traffic capacity of the road segment.

[0119] (1) Node Embedding Module

[0120] The node embedding module aims to learn the low-dimensional dense representation of road network nodes and aggregate the node's own attributes and neighborhood structure information. Graph Attention Network (GAT) is used to implement node embedding. GAT introduces an attention mechanism to adaptively assign importance weights to different neighboring nodes, thereby highlighting the influence of key nodes. Suppose the input of the l-th layer GAT is The output is Where N is the number of nodes, d l is the embedding dimension of the l-th layer node; for node i, the self-attention mechanism is used to calculate the attention weight of its neighbor node j∈Ni:

[0121]

[0122] in, is a learnable linear transformation matrix, is the attention vector, "||" represents vector concatenation, and is the LeakyReLU activation function;

[0123] The attention weight of node j reflects its contribution to the aggregated features of node i.

[0124] After calculating the weights of all neighboring nodes, the embedding representation of node i is updated using a weighted summation method:

[0125]

[0126] Where σ(·) is a nonlinear activation function, such as the ReLU function. To improve the feature expression capability, multiple attention heads can be stacked in the GAT, that is, multiple sets of attention weights are calculated in parallel, and then the embedded features of different heads are spliced or summed to obtain the final node embedding. After stacking L layers of GAT, the embedding vector of each node can be obtained. It integrates the node's own attributes and the structural information of the L-hop neighborhood.

[0127] (2) Edge Embedding Module

[0128] The edge embedding module aims to learn a low-dimensional, dense representation of road segment edges, revealing the differences in traffic characteristics between different road segments. Similar to node embedding, edge embedding is achieved using a self-attention mechanism.

[0129] Assume that the input of the l-th layer edge embedding module is The output is Where M is the number of road edges, p l is the embedding dimension of the l-th layer edge. ij , first concatenate the embedding vectors of nodes i and j at both ends, and then map them to the edge embedding space through a multi-layer perceptron (MLP):

[0130]

[0131] in, is the road edge e ij The original attribute features of the model. MLPl consists of multiple fully connected layers and nonlinear activation functions, which can flexibly adjust the feature extraction capability of the model.

[0132] After calculating the initial embeddings of all road segment edges, a similar attention mechanism is used to aggregate the feature information of neighboring road segments and refine the edge embedding representation. Taking the intersection node i as the center and assuming its associated edge set as Ei, the attention weight of the road segment edge is:

[0133]

[0134] Among them, q i is the query vector of intersection node i, W E is the attention weight matrix, and tanh(·) is the hyperbolic tangent activation function. The optimized embedding of the road segment edge is:

[0135]

[0136] After stacking L layers of edge embedding modules, the embedding vector of each road segment can be obtained It contains information about the local topological structure of the road network and dynamic changes in traffic.

[0137] (3) Knowledge fusion module

[0138] While learning embedded representations of road network nodes and edges, this patent introduces knowledge graph embedding technology to further improve the effectiveness of evacuation route planning, incorporating empirical knowledge from historical successful evacuation cases into route representation optimization. This not only fully utilizes prior knowledge but also helps alleviate the data sparsity problem of the GNN model.

[0139] First, based on ontology engineering, a knowledge graph covering the field of evacuation route planning was constructed, denoted as K = (C, R, F). C is a set of concept nodes, including key concepts such as "intersection," "road section," "population density," "refuge point," and "capacity"; R is a set of relationship edges, including semantic relationships such as "connected," "located in," and "belongs to"; and F is a set of links between the knowledge graph and the road network, reflecting the correspondence between concept nodes and specific road network nodes, edges, and attributes. During the knowledge graph construction process, we fully absorb the successful experience of historical evacuation events and incorporate relevant knowledge on optimizing route selection, such as avoiding highly populated areas, staying away from high-risk areas, and fully utilizing main roads, to form a high-quality prior knowledge base.

[0140] Then, the TransE model is applied to the constructed knowledge graph K to learn the low-dimensional embedding representation of the knowledge graph. The TransE model is based on the translation invariance assumption, that is, for a triple (h, r, t), h+r=t should be satisfied in the embedding space. Therefore, the concept node c i and the relationship edge r k The embedding vector f i and g k The following loss function should be optimized:

[0141]

[0142] Where d(·) is the distance metric between embedding vectors, usually using the L1 norm or L2 norm; is a set of positive triples, and negative triples are obtained by randomly replacing the head entity or the tail entity from the knowledge graph; γ is an interval hyperparameter used to distinguish positive and negative triples. By minimizing the loss function L, the embedding vectors f of concept nodes and relationship edges in the knowledge graph can be learned. i and g k , which reflects the semantic structure of knowledge related to evacuation path planning.

[0143] Finally, the knowledge graph embedding is fused with the road network embedding through the link set F. Specifically, for the road network node i, find the corresponding concept node c in the knowledge graph i , get c i The embedding vector f i Then, a gated attention mechanism is used to adaptively adjust the original embedding of the node and knowledge embedding f i Importance:

[0144]

[0145] Among them, W Kis the attention weight matrix, and sigmoid is the gating function. The fused node embedding incorporates both the road network structure characteristics and evacuation path planning knowledge, forming a more comprehensive and effective node representation. The knowledge fusion process for road segment edges is similar.

[0146] To further improve the inference and training efficiency of the aforementioned GNN model, especially for ultra-large-scale road network scenarios, this patent employs Neural Architecture Search (NAS) technology to optimize the model's network structure. NAS aims to automatically discover the optimal network architecture, including hyperparameters such as the number of neural network layers, the number of neurons per layer, and the choice of activation function. Here, a differentiable NAS method using gradient descent is employed to unify the network architecture search process and the model training process into a single optimization framework.

[0147] Define an over-parameterized basic network structure, including N node embedding layers, M edge embedding layers and sufficient skip connections, denoted as S (sup) Then, a set of architecture parameters α=[α1,…,α k ,…,α K ], where α_k∈[0,1] represents the probability of retaining the kth network layer or skip connection. Basic network structure S (sup) After sampling the architecture parameter α, a sub-network structure S can be obtained (sup) The loss function of embedding learning is recorded as, and the entire NAS optimization process can be described as: min α E S(sub)~α [L emb (w S(sub) )]

[0148] in, For subnetwork S (sup) The model parameters are shared. Using the concept of weight sharing, all sub-networks share the same set of parameters, avoiding repeated training. At the beginning of optimization, the architecture parameters α are fixed, and the model parameters w are updated via stochastic gradient descent. Then, the model parameters w are fixed, and the architecture parameters α are updated using feedback rewards from the validation set (e.g., improvement in evaluation metrics). Sub-networks are sampled based on the updated α. This iterative process ultimately leads to the search for the optimal network architecture.

[0149] To control model complexity and avoid overfitting during the NAS search process, this patent adopts the following strategies: (1) network decoupling, searching node embeddings and edge embeddings separately to reduce the search space; (2) layering, dividing the network into multiple layers and searching each layer separately to achieve inter-layer decoupling; (3) error correction mechanism, introducing random perturbations to enhance architectural robustness. These strategies can minimize the model's parameter size while ensuring key topological information, demonstrating excellent computational efficiency on ultra-large road networks.

[0150] The GNN embedding model and NAS optimization described above ultimately yield a concise and efficient road network representation model that fully exploits the spatiotemporal dependencies of evacuation routes and integrates multi-source heterogeneous knowledge, providing more intelligent and accurate decision support for subsequent evacuation route planning. Based on this intelligent road network representation, we will now construct a hierarchical reinforcement learning model to adaptively generate dynamic evacuation strategies.

[0151] Step S5 includes:

[0152] As an embodiment, after obtaining the intelligent embedded representation of the road network, it is necessary to further adaptively generate dynamic evacuation strategies based on multi-source information such as population distribution in risk areas, spatial layout of refuge points, and real-time road conditions, including regional crowd evacuation plans and specific evacuation route planning, to minimize the total evacuation time and maximize the coverage of refuge points. Traditional evacuation optimization models usually use heuristic rules or linear programming methods, which have insufficient generalization capabilities and are difficult to handle ultra-large-scale, dynamically evolving evacuation scenarios. To address this problem, this patent is based on the Hierarchical Deep Reinforcement Learning (HDRL) framework, which realizes the decoupling optimization of regional-level evacuation plans and section-level path planning through multi-agent collaboration and macro-micro decision-making stratification, greatly improving the real-time and effectiveness of evacuation strategies.

[0153] Using the CMARL centralized multi-agent reinforcement learning method, the problem of allocating evacuees from different areas to shelters is modeled as a multi-agent system decision-making process. The specific steps are as follows:

[0154] The N high-risk areas in the evacuation road network are divided into M evacuation cells. Each evacuation cell corresponds to an agent, and its state space, action space, and reward function are designed as follows:

[0155] Extract the state space S from the embedded representation of the evacuation road network i : The state space S of the evacuation cell i i Including: population density of the area ρ i , the road network distance d from the nearest refuge point i , the proportion of main roads in the region r i , the proportion of evacuated population i ;

[0156] The state space features of the evacuation road network are spliced together to obtain the global state vector S=S1,S2,…,S M ;

[0157] Action space A i : Action space A for evacuating cell i iis a continuous real number space, representing the proportion of evacuees assigned to each refuge point; assuming there are K refuge points, the action vector of evacuation cell i at decision step t is satisfy

[0158] The proportion of people evacuated can be converted into the actual number of people, i.e. where N i is the total population of evacuation area i;

[0159] The action vectors of all evacuation cells are spliced together to obtain the joint action vector of the multi-agent system at step t

[0160] Reward function R i : Reward function R for evacuating cell i i It consists of two parts, namely individual rewards and collaboration rewards Right now λ is the balance coefficient, which controls the weight of individual rewards and collaborative rewards;

[0161] Individual rewards Indicates the evacuation effect of the evacuation area;

[0162] Design individual rewards based on indicators such as evacuation completion time and evacuation path length, Where T i is the total time to complete the evacuation of cell i, T max is the global maximum evacuation time;

[0163] Collaboration Rewards Indicates the degree of evacuation coordination between the evacuation cell and other cells;

[0164] Design collaborative rewards based on indicators such as shelter capacity balance and path similarity, such as Where N(i) is the neighborhood set of evacuation cell i, P i and P j are the evacuation path sets of two cells respectively, sim(·) is the path similarity function;

[0165] A multi-agent deep deterministic policy gradient algorithm based on the Actor-Critic framework is used to perform joint strategy learning on the multi-agent decision-making system to obtain the final joint optimal dynamic evacuation strategy.

[0166] As an embodiment, the hierarchical deep reinforcement learning framework proposed in this patent deconstructs the evacuation path planning problem into two subtasks: regional route selection and section-level path optimization, through macro-micro decision-making hierarchies, thereby reducing the complexity of solving the problem. The macro layer focuses on the overall situation and is responsible for selecting high-level regional evacuation plans; the micro layer focuses on the local situation and is responsible for adjusting specific evacuation routes in real time. The two levels complement each other and optimize collaboratively. Based on the intelligent road network representation, this hierarchical architecture can significantly improve the generalization performance and computational efficiency of evacuation path planning, and better cope with ultra-large-scale, dynamic and changeable evacuation scenarios.

[0167] This application also discloses an electronic device. Figure 2 , Figure 2 Schematic diagram of the structure of an electronic device disclosed in an embodiment of the present application. The electronic device 500 may include: at least one processor 501, at least one network interface 504, a user interface 503, a memory 505, and at least one communication bus 502.

[0168] The communication bus 502 is used to implement the connection and communication between these components.

[0169] The user interface 503 may include a display screen, and the optional user interface 503 may also include a standard wired interface or a wireless interface.

[0170] The network interface 504 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface).

[0171] The present application also discloses a computer-readable storage medium storing a plurality of instructions suitable for loading by a processor to execute the above-mentioned evacuation path optimization method based on coupled numerical simulation and hierarchical reinforcement learning.

[0172] The above are merely exemplary embodiments of the present disclosure and are not intended to limit the scope of the present disclosure. In other words, any equivalent changes and modifications made according to the teachings of the present disclosure are still within the scope of the present disclosure.

[0173] This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not described herein. The description and examples are to be considered as exemplary only, and the scope and spirit of the present disclosure are to be defined by the claims.

Claims

1. An evacuation path optimization method based on coupled numerical simulation and hierarchical reinforcement learning, characterized in that: The method comprises the following steps: S1: Build a high-precision coupled numerical simulation framework for storm surges, which includes the WRF atmospheric model, the ADCIRC model, and the SWAN ocean model. S2: Conduct storm surge numerical simulations through a coupled numerical simulation framework to generate an inundation depth dataset; S3: Perform high-resolution inundation depth interpolation using the inundation depth dataset and the inverse distance weighted interpolation algorithm to generate a high-resolution inundation depth interpolation dataset; perform risk classification on the high-resolution inundation depth interpolation dataset to generate a risk zoning map; S4: Generate an evacuation road network through a graph neural network-based path planning neural network model and a risk zoning map; S5: Adopt a hierarchical deep reinforcement learning framework to optimize the evacuation strategy of the evacuation road network and generate a joint optimal dynamic evacuation strategy.

2. The evacuation path optimization method based on coupled numerical simulation and hierarchical reinforcement learning according to claim 1, characterized in that: Step S1 includes: The OASIS3-MCT coupler is used to control the WRF atmospheric model, ADCIRC model, and SWAN ocean model. The coupling variables are exchanged within a specified time interval to achieve bidirectional coupling between WRF, ADCIRC, and SWAN. The basic coupling relationship is as follows: (g) a ):ADCIRC→OASIS→(ζ w ):WRF Among them, OASIS is a general coupling framework responsible for controlling the exchange and synchronization of coupling variables between different modes; and P w represent the wind field and pressure field provided by the WRF model respectively; represent the upper boundary wind field and pressure field driving of ADCIRC model and SWAN model respectively; and represents the wave field transferred from the SWAN mode to the ADCIRC mode; ζ a and ζ w Indicates the water level for transferring from ADCIRC mode to SWAN mode.

3. The evacuation path optimization method based on coupled numerical simulation and hierarchical reinforcement learning according to claim 1, characterized in that: Step S2 includes: S21: Extract the key parameters of tropical cyclones that land on China from the typhoon best track database. The key parameters include: time of landfall, latitude and longitude of the landing point, central pressure, maximum wind speed, moving speed, and moving direction. S22: Based on the statistical analysis of key typhoon parameters, the probability distribution characteristics are obtained; Using the Monte Carlo method, we sampled several typhoon path and intensity combinations from the probability distribution characteristics to obtain the storm surge numerical simulation conditions. S23: For each set of storm surge numerical simulation conditions, the Holland wind field model is used to generate the spatiotemporal distribution characteristics of the typhoon wind field and pressure field according to the typhoon optimal path parameters; Tangential wind speed of the Holland wind model and pressure field satisfy: in, is the radial distance from the typhoon center; A and B are dimensionless parameters related to the typhoon intensity and structure; is the air density; Indicates the central pressure of the typhoon; Indicates the ambient air pressure value outside the typhoon; is the maximum wind speed radius; The optimal typhoon path parameters include: typhoon center location, typhoon center pressure, and typhoon maximum wind speed radius; S24: Interpolate the spatiotemporal distribution characteristics of typhoon wind and pressure fields onto the grid points of the WRF, ADCIRC, and SWAN marine meteorological models; The interpolated grid points are fused with the background field to form a submerged depth dataset; The background field uses reanalysis data or climate mean field; Fusion methods include spectrum matching and optimal interpolation.

4. The evacuation path optimization method based on coupled numerical simulation and hierarchical reinforcement learning according to claim 1, characterized in that: Step S2 further includes: A parallel strategy combining domain decomposition and task decomposition is adopted to carry out numerical simulation of storm surge.

5. The evacuation path optimization method based on coupled numerical simulation and hierarchical reinforcement learning according to claim 1, characterized in that: Step S3 includes: The specific steps of spatial interpolation of high-resolution inundation depth include: interpolating numerical simulation results from an unstructured source grid to a high-resolution structured target grid; extracting numerical simulation results from an inundation depth dataset; the unstructured source grid is a triangular grid or a quadrilateral grid; the high-resolution structured target grid is a regular rectangular grid; Assume D s (i,j) and D t (p,q) are the flooding depths of the source grid cell (i,j) and the target grid cell (p,q), and their spatial coordinates are (x i ,y j ) and (X p ,Y q ); Let w k is the weight of the interpolation of the kth source grid cell to the target grid cell (p,q), satisfying the normalization condition ∑w k =1, then the inverse distance weighted interpolation formula is: Among them, K is the total number of source grid cells involved in interpolation; weight w k Defined using the inverse distance function; represents the nth power of the Euclidean distance between the center of the kth source grid cell and the center of the target grid cell; d k is the Euclidean distance between the center of the kth source grid cell and the center of the target grid cell, and n is the inverse distance weighted index; It represents the nth power of the Euclidean distance between the center of the lth source grid cell and the center of the target grid cell; The risk zoning map includes: blue zone, yellow zone, orange zone and red zone; Blue zone: 0.2m≤submergence depth<0.5m, slight impact on personnel and farmland, low risk level; Yellow zone: 0.5m≤submergence depth<1.0m, travel is obstructed, crops are flooded on a large scale, and the risk level is medium; Orange zone: 1.0m≤submergence depth<1.5m, building structures begin to be damaged, personnel safety is seriously threatened, and the risk level is high; Red zone: Flooding depth ≥ 1.5m, large-scale building collapse, people's lives are difficult to protect, and the risk level is extremely high.

6. The evacuation path optimization method based on coupled numerical simulation and hierarchical reinforcement learning according to claim 1, characterized in that: Step S4 includes: the path planning neural network model includes: a node embedding module, an edge embedding module and a knowledge fusion module; the node embedding module, the edge embedding module and the knowledge fusion module are connected in sequence.

7. The evacuation path optimization method based on coupled numerical simulation and hierarchical reinforcement learning according to claim 1, characterized in that: Step S5 includes: Using the CMARL centralized multi-agent reinforcement learning method, the problem of allocating evacuees from different areas to shelters is modeled as a multi-agent system decision-making process. The specific steps are as follows: The N high-risk areas in the evacuation road network are divided into M evacuation cells. Each evacuation cell corresponds to an agent, and its state space, action space, and reward function are designed as follows: Extract the state space S from the embedded representation of the evacuation road network i : The state space S of evacuated cell i i Including: population density of the area ρ i , the road network distance d from the nearest refuge point i , the proportion of main roads in the region r i , the proportion of evacuated population i ; The state space features of the evacuation road network are spliced together to obtain the global state vector S=S1,S2,…,S M ; Action space A i : Action space A for evacuating cell i i is a continuous real number space, representing the proportion of evacuees assigned to each refuge point; assuming there are K refuge points, the action vector of evacuation cell i at decision step t is satisfy The proportion of people evacuated can be converted into the actual number of people, i.e. where N i is the total population of evacuation area i; The action vectors of all evacuation cells are spliced together to obtain the joint action vector of the multi-agent system at step t Reward function R i : Reward function R for evacuating cell i i It consists of two parts, namely individual rewards and collaboration rewards Right now λ is the balance coefficient, which controls the weight of individual rewards and collaborative rewards; Individual rewards Indicates the evacuation effect of the evacuation area; Design individual rewards based on indicators such as evacuation completion time and evacuation path length, Where T i is the total time to complete the evacuation of cell i, T max is the global maximum evacuation time; Collaboration Rewards Indicates the degree of evacuation coordination between the evacuation cell and other cells; Design collaborative rewards based on indicators such as shelter capacity balance and path similarity, such as Where N(i) is the neighborhood set of evacuation cell i, P i and P j are the evacuation path sets of two cells respectively, sim(·) is the path similarity function; A multi-agent deep deterministic policy gradient algorithm based on the Actor-Critic framework is used to perform joint strategy learning on the multi-agent decision-making system to obtain the final joint optimal dynamic evacuation strategy.

8. An electronic device, characterized in that: The electronic device comprises a processor, a memory, a user interface and a network interface, wherein the memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions, and when the instructions are executed by a computer, the method according to any one of claims 1 to 7 is executed.

Citation Information

Cited By

  • Regional collaborative storm surge water increase forecasting method of coupling graph attention and gated cycle network

    CN121118702A

  • A Regional Cooperative Storm Surge Prediction Method Using Coupled Graph Attention and Gated Recurrent Networks

    CN121118702B