A fire emergency dispatch optimization method based on deep reinforcement learning

By employing a fire emergency dispatching method based on deep reinforcement learning, and utilizing the spatiotemporal hypergraph Transformer and the hierarchical multi-agent graph Transformer, combined with the world model and CVaR risk reinforcement learning, the problem of insufficient expression of complex water supply topology and weak risk handling in fire emergency dispatching is solved, and continuous water supply guarantee and dispatching efficiency are achieved under uncertain conditions.

CN122114456APending Publication Date: 2026-05-29NANJING CHENGLANG INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING CHENGLANG INFORMATION TECHNOLOGY CO LTD
Filing Date
2026-01-23
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing emergency dispatch methods for fire sites lack sufficient expression of complex water supply topology, have weak handling of risks and uncertainties, and lack calculable guarantees for safety constraints. This leads to deviations in continuous water supply capacity and bottleneck assessment, insufficient handling of risks and uncertainties, and lack of calculable guarantees for safety constraints, resulting in unstable strategy executability and compliance rates.

Method used

A deep reinforcement learning-based approach is adopted, which uses a spatiotemporal hypergraph Transformer and a hierarchical multi-agent graph Transformer, combined with a world model, to predict the distribution of consumable consumption, arrival time, water availability, and inventory balance. Fire safety red lines and risk gating are introduced to select high-rise tasks and allocate time windows. Candidate actions are projected to the feasible region through secondary planning safety filtering, and CVaR risk reinforcement learning is used for policy updates.

Benefits of technology

Ensuring continuous water supply under uncertain conditions reduces the risk of supply interruptions and delays, improves scheduling efficiency and compliance, and ensures the feasibility and compliance of the strategy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122114456A_ABST
    Figure CN122114456A_ABST
Patent Text Reader

Abstract

The application discloses a fire emergency dispatching optimization method based on deep reinforcement learning, and aims at solving the space-time coordination and safety compliance problems of fire supply link and consumable dispatching. The application realizes the technical effects of continuous water supply guarantee, reduction of tail risk of delay and supply interruption, improvement of compliance rate of access prohibition and time window constraint, and overall scheduling efficiency by constructing a space-time hypergraph, adopting a world model for distribution prediction, implementing gating in a constraint-aware hypergraph Transformer, and performing hierarchical multi-agent strategy reasoning, combining quadratic programming safety filtering and CVaR risk reinforcement learning closed-loop updating.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of emergency dispatch, and in particular to an optimization method for fire emergency dispatch based on deep reinforcement learning. Background Technology

[0002] With the advancement of urbanization and the increasing complexity of fire scene environments, fire emergency support involves the coordination of multiple links such as water intake, laying of water hoses, temporary connection, foam agent replenishment, and replacement of breathing apparatus cylinders, exhibiting significant spatiotemporal coupling and resource sharing constraints.

[0003] In terms of existing technologies, command and dispatch typically rely on GIS and vehicle positioning, combined with road traffic and equipment status, and use deterministic optimization methods such as shortest path, network flow, and vehicle routing problems, or heuristic methods such as genetic algorithms and tabu search for task and vehicle allocation. Some studies have begun to introduce graph neural networks and reinforcement learning for path and task optimization, but most of them are centralized or single-layer architectures, which have limited expression of complex water supply topologies and operational dependencies. Safety and compliance are mostly handled through rule bases and manual checks, lacking endogenous integration with optimization and learning algorithms.

[0004] However, existing technologies still have shortcomings:

[0005] 1. Insufficient expressive ability: Simple diagrams or task lists are insufficient to depict the hyper-edge relationships and operational sequence dependencies of multiple water sources, multiple node convergence, parallel laying and connection, resulting in deviations in continuous water supply capacity and bottleneck assessment.

[0006] 2. Insufficient handling of risk and uncertainty: Most methods are based on point estimation or expectation optimization, lacking tail risk measurement and optimization for supply interruption duration and arrival delay, which can easily lead to extreme events in high uncertainty scenarios;

[0007] 3. Lack of computable security constraints: Learning or heuristic strategies lack a unified feasible domain projection and online correction for hard constraints such as restricted areas, minimum inventory, time windows, and continuous water supply, resulting in unstable strategy executability and compliance rates.

[0008] Therefore, an optimization method for emergency dispatching in fire situations that can overcome the shortcomings of the existing technologies is a problem that needs to be solved by those skilled in the art. Summary of the Invention

[0009] One objective of this invention is to propose a fire emergency dispatch optimization method based on deep reinforcement learning. Addressing the shortcomings of existing technologies in representing complex water supply topologies, handling risks and uncertainties, and lacking computationally achievable guarantees for safety constraints, this invention proposes a method that uses a spatiotemporal hypergraph as a carrier, combined with a world model to predict the distribution of consumable consumption, arrival time, water availability, and inventory reserves. It introduces fire safety red lines and risk gating (hard shielding + soft scaling) into a constraint-aware spatiotemporal hypergraph Transformer. A hierarchical multi-agent graph Transformer is used for high-level task selection and time window allocation, and low-level path and connection execution. A secondary planning safety filter projects candidate actions to a feasible region including continuous water supply, minimum inventory, restricted areas, and sequence dependencies. Finally, a CVaR risk reinforcement learning-based closed-loop update strategy based on on-site feedback is employed. This invention achieves the technical effects of ensuring continuous water supply under uncertain conditions, reducing the risk of supply interruptions and late arrivals, and improving constraint compliance and overall dispatch efficiency.

[0010] A fire emergency dispatch optimization method based on deep reinforcement learning according to an embodiment of the present invention is characterized by comprising the following steps:

[0011] S1. Collect vehicle positioning, road traffic status, building topology, water source information, water pressure sensor data, consumable inventory data and combat mission progress data. Based on the data, construct an initial spatiotemporal hypergraph. Generate a fire protection red line set and a constraint parameter set based on regulations and operating procedures. Initialize the strategy parameters of the hierarchical multi-agent graph Transformer. Output the initial spatiotemporal hypergraph, fire protection red line set, constraint parameter set and strategy parameters.

[0012] S2. Using the initial spatiotemporal hypergraph as input, the world model is used to predict the distribution of consumable consumption rate, arrival time, water availability and inventory balance, and to obtain the predicted distribution set, quantile set and risk score set.

[0013] S3. The initial spatiotemporal hypergraph, fire red line set, constraint parameter set, strategy parameter, prediction distribution set, quantile set and risk score set are taken as inputs. Constraint gating is performed in the attention calculation of the constraint-aware spatiotemporal hypergraph Transformer to obtain the gated hypergraph representation.

[0014] S4. Using the gated hypergraph representation and policy parameters as input, a hierarchical multi-agent graph Transformer is used for policy reasoning. The higher layer outputs the supply task selection and time window allocation decision, and the lower layer outputs the path planning and connection execution actions. The combination yields a set of candidate actions.

[0015] S5. Taking the candidate action set and constraint parameter set as input, the project solution is performed in the quadratic programming safety filtering module based on the continuous water supply constraint parameters, minimum inventory threshold parameters, time window parameters, restricted area parameters and sequence dependency parameters to obtain safe and feasible actions.

[0016] S6. Execute safe and feasible actions and strategy parameters on-site and collect execution feedback data;

[0017] S7. Using the execution feedback data and policy parameters as input, CVaR risk reinforcement learning is used to calculate the conditional risk loss of the supply interruption duration and late timeout, and to update the policy parameters of the high-level and low-level layers to obtain the updated policy parameters. The updated policy parameters are then output for use in policy reasoning in S4 at subsequent times.

[0018] Optionally, step S1 specifically includes:

[0019] The system aligns vehicle positioning, road traffic status, building topology, water source information, water pressure sensor data, consumable inventory data, and combat mission progress data with timestamps and coordinates, fills in missing data and removes anomalies to obtain aligned multi-source situational data and outputs the aligned multi-source situational data.

[0020] Using aligned multi-source situational awareness data as input, an initial spatiotemporal hypergraph is constructed and output. The initial spatiotemporal hypergraph includes a set of nodes and a set of hyperedges. The set of nodes includes water source nodes, temporary connection nodes, hose confluence nodes, pump truck nodes, combat unit nodes, and supply station nodes. The set of hyperedges is used to represent the supply link from the water source node to the combat unit node. Hyperedge features include water pressure, flow limit, altitude difference, pipe diameter, hose laying length, number of joints, connection time, laying time, and current occupancy status. The water pressure and flow limit are calculated or estimated based on water pressure sensor data and pipe diameter. The altitude difference is calculated based on terrain elevation or building elevation. The hose laying length is estimated based on the length of the passable path between nodes. The connection time and laying time are estimated based on distance and operating conditions. The current occupancy status is determined based on the on-site equipment and mission progress.

[0021] The initial spatiotemporal hypergraph is used as input to generate a set of fire protection red lines based on regulations, operating procedures, and on-site control information, and the set of fire protection red lines is output.

[0022] Simultaneously, based on the fire protection red line set and the aligned multi-source situational data, a set of constraint parameters is derived and output. The set of constraint parameters includes continuous water supply constraint parameters, minimum inventory threshold parameters, time window parameters, restricted area parameters, and sequence dependency parameters. Among them, the continuous water supply constraint parameters are set according to the relationship between water pressure and flow rate limits and consumable consumption; the minimum inventory threshold parameters are set according to consumable inventory data; the time window parameters are set according to road traffic status and combat mission progress; the restricted area parameters are set according to the spatial range of the fire protection red line set; and the sequence dependency parameters are set according to the operation sequence of connection and laying.

[0023] The initial spatiotemporal hypergraph, the fire protection red line set, and the constraint parameter set are used as inputs to initialize the policy parameters of the hierarchical multi-agent graph Transformer and output the policy parameters. The initialization includes assigning initial values ​​to the network weights, the gating threshold corresponding to the fire protection red line set, and the mask corresponding to the constraint parameter set in the policy parameters, so that the policy parameters are consistent with the initial spatiotemporal hypergraph, the fire protection red line set, and the constraint parameter set.

[0024] Optionally, step S2 specifically includes:

[0025] The initial spatiotemporal hypergraph is used as input. The node features and hyperedge features in the initial spatiotemporal hypergraph are encoded to form the model input sequence for the world model and output the model input sequence.

[0026] The model input sequence is used as input, and the encoder and state-space dynamics in the world model (DreamerV3) are used to perform time-series inference on the model input sequence to obtain the latent state sequence and output the latent state sequence.

[0027] Using the potential state sequence as input, parameterized probability modeling of consumable consumption rate, arrival time, water availability and inventory balance is performed through the corresponding decoder to generate a set of predicted distributions covering the variables and output the set of predicted distributions.

[0028] The predicted distribution set is used as input, and the quantile set is calculated based on the target quantile set. Combined with the fire protection red line set and the constraint parameter set, the risk scores of continuous water supply default and time window default are calculated based on the quantile exceedance probability or the conditional tail expectation, forming a risk score set and outputting the quantile set and risk score set.

[0029] The predicted distribution set, quantile set, and risk score set are output together with the initial spatiotemporal hypermap, fire protection red line set, constraint parameter set, and strategy parameters.

[0030] Optionally, step S3 specifically includes:

[0031] The initial spatiotemporal hypergraph, fire protection red line set, constraint parameter set, strategy parameter, prediction distribution set, quantile set, and risk score set are taken as input. Based on the fire protection red line set and constraint parameter set, the constraints of each hyperedge in the initial spatiotemporal hypergraph are judged to determine whether they violate the continuous water supply constraint, minimum inventory threshold, time window constraint, restricted area constraint, and sequence dependency constraint, forming a constraint mask set and outputting the constraint mask set.

[0032] Simultaneously, risk scaling coefficients are calculated for each hyperedge based on the quantile set and risk score set, forming a risk scaling coefficient set, which is then output together with the initial spatiotemporal hypergraph, constraint parameter set, and strategy parameters.

[0033] The initial spatiotemporal hypergraph, constraint parameter set, policy parameter, constraint mask set, and risk scaling factor set are taken as input. Attention is calculated on the initial spatiotemporal hypergraph in the constraint-aware spatiotemporal hypergraph Transformer to generate an ungated attention weight set. The ungated attention weight set, constraint mask set, risk scaling factor set, initial spatiotemporal hypergraph, constraint parameter set, and policy parameter set are output together.

[0034] The ungated attention weight set, constraint mask set, risk scaling factor set, initial spatiotemporal hypergraph, constraint parameter set and policy parameters are taken as input. The attention weights corresponding to hyperedges that violate the constraint mask set are hard masked and the corresponding weights are reset to zero.

[0035] The attention weights that are not hard-masked are soft-scaled according to the risk scaling factor set to obtain the gated attention weight set. The gated attention weight set is then output together with the initial spatiotemporal hypergraph, the constraint parameter set, and the policy parameters.

[0036] The gated attention weight set, the initial spatiotemporal hypergraph, and the policy parameters are taken as input and forward computation is performed in the constraint-aware spatiotemporal hypergraph Transformer to form the gated hypergraph representation. The gated hypergraph representation, the constraint parameter set, and the policy parameters are then output.

[0037] Optionally, step S4 specifically includes:

[0038] Using the gated hypergraph representation, constraint parameter set, and policy parameters as input, input embeddings are constructed for each agent in the hierarchical multi-agent graph Transformer based on the gated hypergraph representation. These are used to form high-level input embeddings for high-level reasoning and low-level input embeddings for low-level reasoning, and the outputs are high-level input embeddings, low-level input embeddings, constraint parameter set, and policy parameters.

[0039] The high-level input embedding, constraint parameter set, and policy parameters are used as inputs. Policy reasoning is performed at the high level of the hierarchical multi-agent graph Transformer to generate supply task selection and time window allocation decisions, forming a high-level decision set and outputting the high-level decision set, constraint parameter set, and policy parameters.

[0040] The high-level decision set, as well as the low-level input embedding, constraint parameter set, and policy parameter set, are used as inputs. Policy reasoning is performed at the low level of the hierarchical multi-agent graph Transformer to generate path planning and connection execution actions, forming a low-level action set and outputting the low-level action set, constraint parameter set, and policy parameter set.

[0041] The set of low-level actions, the set of high-level decisions, the set of constraint parameters, and the set of policy parameters are taken as input, combined and aligned to obtain a set of candidate actions, and the set of candidate actions, the set of constraint parameters, and the set of policy parameters are output.

[0042] Optionally, step S5 specifically includes:

[0043] The candidate action set, constraint parameter set, and strategy parameters are taken as input. In the quadratic programming safety filtering module, the candidate action set is numerically encoded to establish a quadratic programming model and output the quadratic programming model and strategy parameters. The decision variables of the quadratic programming model are used to represent the action parameters of the candidate action set. The objective function is used to minimize the deviation of the decision variables from the candidate action set and simultaneously minimize the weighted value of the time window delay. The constraints are set according to the constraint parameter set, including continuous water supply constraint parameters, minimum inventory threshold parameters, time window parameters, restricted area parameters, and sequence dependency parameters.

[0044] The quadratic programming model and policy parameters are used as inputs. The projection is solved in the quadratic programming security filtering module to obtain the preliminary projection action vector, and the preliminary projection action vector and policy parameters are output.

[0045] The initial projection action vector, the set of constraint parameters, and the strategy parameters are taken as input. The feasibility of the initial projection action vector is checked based on the set of constraint parameters. If it is found that the continuous water supply constraint parameters, minimum inventory threshold parameters, time window parameters, restricted area parameters, or sequence dependency parameters are violated, the nearest projection correction to the feasible domain boundary is performed in the secondary planning safety filtering module to obtain the feasible action vector and output the feasible action vector and strategy parameters.

[0046] The action vector and policy parameters are taken as input, and the action parameters are restored and aligned according to the action parameter definition of the candidate action set to form safe actions. The safe actions and policy parameters are then output.

[0047] Optionally, step S6 specifically includes:

[0048] Taking safe and feasible actions and policy parameters as input, the safe and feasible actions are parsed into instructions to generate a set of execution instructions containing target location, target node, time window, resource quantity and job type. A unified timestamp is generated for each instruction, and the set of execution instructions and policy parameters are output.

[0049] The execution instruction set and strategy parameters are taken as input. The execution instruction set is sent to the site and status acquisition is started. The site status data related to the execution is collected to form raw execution data. The raw execution data includes at least execution instruction identifier, timestamp, location record, water supply status record and inventory sampling record. The raw execution data and strategy parameters are output.

[0050] The original execution data and strategy parameters are used as input. Timestamp alignment and anomaly removal are performed. The execution result indicators are calculated based on the execution instruction identifier and timestamp in the original execution data to obtain inventory status, water supply status, arrival time and compliance status. Among them, inventory status represents the remaining quantity of key consumables and the remaining operation time; water supply status represents the continuous water supply indication, water supply interruption duration and water supply capacity assessment value; arrival time represents the time interval from the issuance of the instruction to the arrival at the target location; and compliance status represents whether the execution violates the gating threshold and the spatial and time window constraints corresponding to the mask encoded in the strategy parameters. The output is a set of indicators and strategy parameters.

[0051] The system takes a set of indicators and strategy parameters as input, aggregates and formats the indicator set to form execution feedback data, which includes at least inventory status, water supply status, arrival time and compliance status. The system outputs execution feedback data and strategy parameters.

[0052] Optionally, step S7 specifically includes:

[0053] The system takes execution feedback data and strategy parameters as input, extracts inventory status, water supply status, arrival time and compliance status from the execution feedback data, generates a water supply interruption duration sequence based on the water supply interruption duration in the water supply status, and generates a late arrival timeout sequence based on the arrival time and read the upper limit of the time window parameter used for late arrival calculation from the strategy parameters, forming a risk sample set and outputting the risk sample set and strategy parameters.

[0054] Using the risk sample set and strategy parameters as input, the risk level parameter is read from the strategy parameter. The quantile threshold corresponding to the risk level parameter is calculated for the supply interruption duration sequence and the late arrival timeout sequence, respectively. The average value of the part exceeding the quantile threshold is further calculated as the conditional risk loss. The conditional risk loss set is obtained and the conditional risk loss set and strategy parameters are output.

[0055] The conditional risk loss set and strategy parameters are taken as input. The risk weight parameters are read from the strategy parameters. The conditional risk loss set is weighted and synthesized to form the total risk loss. The total risk loss and strategy parameters are then output.

[0056] Using the total risk loss and policy parameters as input, CVaR risk reinforcement learning is used to update the high-level and low-level policy parameters of the hierarchical multi-agent graph Transformer. The updated policy parameters are obtained by calculating the gradient of the total risk loss and performing backpropagation and iteration. The updated policy parameters are output for policy inference in subsequent time steps in S4 of claim 1.

[0057] The beneficial effects of this invention are:

[0058] 1. Continuous water supply guarantee and tail risk reduction: By using the distribution prediction of the world model and CVaR risk reinforcement learning, and implementing hard shielding and soft scaling risk gating in the constraint-aware hypergraph Transformer, the conditional tail expectation of the interruption duration and the arrival delay is significantly reduced, thereby improving the continuous water supply guarantee rate.

[0059] 2. Improved compliance and executability: Fire safety red lines and constraint parameters (restricted areas, time windows, minimum inventory, sequence dependence, continuous water supply) are embedded into attention gating, and candidate actions are modified by the nearest projection of secondary planning to ensure that the strategy does not violate hard constraints, thereby improving on-site executability and compliance rate.

[0060] 3. Enhanced scheduling efficiency and topology representation: The spatiotemporal hypergraph accurately depicts the sequence of multiple water sources, multiple nodes convergence and connection operations. Combined with the hierarchical multi-agent graph Transformer, it realizes the coordination of high-level tasks, time windows and low-level paths and connections, reduces redundant laying and ineffective connections, shortens arrival time and improves resource utilization efficiency. Attached Figure Description

[0061] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0062] Figure 1 This is a flowchart of a fire emergency dispatch optimization method based on deep reinforcement learning proposed in this invention. Detailed Implementation

[0063] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0064] refer to Figure 1A fire emergency dispatch optimization method based on deep reinforcement learning, characterized by the following steps:

[0065] S1. Collect vehicle positioning, road traffic status, building topology, water source information, water pressure sensor data, consumable inventory data and combat mission progress data. Based on the data, construct an initial spatiotemporal hypergraph. Generate a fire protection red line set and a constraint parameter set based on regulations and operating procedures. Initialize the strategy parameters of the hierarchical multi-agent graph Transformer. Output the initial spatiotemporal hypergraph, fire protection red line set, constraint parameter set and strategy parameters.

[0066] S2. Using the initial spatiotemporal hypergraph as input, the world model is used to predict the distribution of consumable consumption rate, arrival time, water availability and inventory balance, and to obtain the predicted distribution set, quantile set and risk score set.

[0067] S3. The initial spatiotemporal hypergraph, fire red line set, constraint parameter set, strategy parameter, prediction distribution set, quantile set and risk score set are taken as inputs. Constraint gating is performed in the attention calculation of the constraint-aware spatiotemporal hypergraph Transformer to obtain the gated hypergraph representation.

[0068] S4. Using the gated hypergraph representation and policy parameters as input, a hierarchical multi-agent graph Transformer is used for policy reasoning. The higher layer outputs the supply task selection and time window allocation decision, and the lower layer outputs the path planning and connection execution actions. The combination yields a set of candidate actions.

[0069] S5. Taking the candidate action set and constraint parameter set as input, the project solution is performed in the quadratic programming safety filtering module based on the continuous water supply constraint parameters, minimum inventory threshold parameters, time window parameters, restricted area parameters and sequence dependency parameters to obtain safe and feasible actions.

[0070] S6. Execute safe and feasible actions and strategy parameters on-site and collect execution feedback data;

[0071] S7. Using the execution feedback data and policy parameters as input, CVaR risk reinforcement learning is used to calculate the conditional risk loss of the supply interruption duration and late timeout, and to update the policy parameters of the high-level and low-level layers to obtain the updated policy parameters. The updated policy parameters are then output for use in policy reasoning in S4 at subsequent times.

[0072] In this specific embodiment, S1 specifically refers to:

[0073] First, time alignment and quality control are performed on multi-source data, including vehicle positioning, road traffic, building topology, water source information, water pressure sensing, consumable inventory, and combat mission progress. Time alignment uses an offset correction model.

[0074] ;

[0075] in For the original timestamp, For corrected timestamps, For data source Time offset, Index the data source;

[0076] Constructing an initial spacetime hypergraph in a unified coordinate system:

[0077] ;

[0078] in For the initial spacetime hypergraph, This is a collection of nodes (including water sources, temporary connection points, rendezvous points, pump trucks, combat units, and supply stations, etc.). The hyperedge set is used to represent the supply link from the water source to the combat unit. Hyperedge features include water pressure, flow limit, altitude difference, pipe diameter, laying length, number of joints, connection and laying time, and occupancy status. The flow limit can be determined by:

[0079] Make an estimate;

[0080] For super-edge Traffic limit For empirical coefficients, For pipe diameter, For water pressure, For medium density, For superedge index;

[0081] A set of fire safety red lines is generated based on regulations, operating procedures, and on-site control information. And based on this, the set of constraint parameters is derived. The constraints include at least continuous water supply, minimum inventory, time window, restricted areas, and job sequence dependency, wherein the time window can be expressed as:

[0082] ;

[0083] For nodes or tasks Feasible time window For the earliest start time, For the latest completion time, For indexing nodes or tasks;

[0084] Continuous water supply can be achieved through the supply-demand ratio:

[0085] and constrained to Make a judgment;

[0086] For super-edge supply and demand ratio For the super-edge The service's combat unit set, For combat unit indexing, For combat units unit time demand, The minimum supply-demand ratio threshold;

[0087] Finally, initialize the policy parameters of the hierarchical multi-agent graph Transformer:

[0088] With and Maintain consistency;

[0089] in For the set of strategy parameters, For network weight, For the gate control threshold vector corresponding to the fire protection red line, This is the mask tensor corresponding to each constraint.

[0090] In this specific embodiment, S2 specifically refers to:

[0091] Encoding node and hyperedge features based on the initial spatiotemporal hypergraph to form the world model input sequence can be written as:

[0092] ;

[0093] in For a moment Encoding vector, For feature coding function, For a moment Node feature set, For a moment hyperedge feature set, For time indexing;

[0094] Subsequently, temporal inference is performed in the state-space dynamics of the world model to obtain the potential state sequence, denoted as:

[0095] ;

[0096] in For a moment Potential state The potential state of the previous moment, For state transition function, This is a vector of dynamic parameters;

[0097] Using the potential state as input, parameterized probabilistic modeling is performed on the consumption rate of consumables, arrival time, water availability, and inventory balance, which are uniformly represented as:

[0098] and ;

[0099] in For a moment Target variables (including consumable consumption rate, arrival time) Water availability Inventory balance one), For conditional probability distribution, For the selected distribution family, For distribution parameters, For decoding functions, These are the decoding parameters;

[0100] The quantiles are calculated based on the target quantile set and are expressed as follows:

[0101] ;

[0102] in For variables At the quantile level quantiles For variables The cumulative distribution function, The quantile level is taken from the target quantile set;

[0103] A risk assessment of continuous water supply and time windows is conducted by combining fire safety red lines and constraint parameters. The probability of breach of continuous water supply regulations can be expressed as:

[0104] ;

[0105] in For a moment The probability of continuous water supply default exceeding For probability operators, The degree of water supply interruption (e.g., can be determined by...) (Given) The threshold is determined based on constraints and fire safety boundaries;

[0106] The expected value at the end of the time window default condition can be written as:

[0107] and ;

[0108] in To at the quantile level Under the following conditions, tail expectation For expectation operator, For late arrivals and overtime, For arrival time, For nodes or tasks The upper limit of the time window is derived from the set of constraint parameters. For quantiles, For indexing nodes or tasks;

[0109] Ultimately, a set of predicted distributions, quantiles, and risk scores covering each target variable are formed for subsequent constraint gating and policy reasoning.

[0110] In this specific embodiment, S3 specifically refers to:

[0111] The initial spatiotemporal hypergraph, fire safety red line and constraint parameters, strategy parameters, and the predicted distribution, quantiles, and risk scores output by the world model are used as inputs. Constraints are determined for each hyperedge, and a constraint mask is generated. The following method is employed:

[0112] ;

[0113] in For super-edge constraint mask, For indicator functions, For the maximum degree of default, Indicates the degree of breach of continuous water supply constraints. The default rate indicates the minimum inventory threshold. Indicates the default rate within a time window, Indicates the degree of breach of the restricted area, Indicates the degree of default of sequence dependency. For superedge index;

[0114] In risk quantification, quantiles and risk scores are combined to comprehensively measure water supply and lateness. The amount of lateness can be expressed as:

[0115] ;

[0116] in For super-edge The amount of late arrivals and overtime, To predict arrival time, For target node or task The upper limit of the time window;

[0117] The overall risk and soft scaling factor are given as follows:

[0118] and ;

[0119] in For super-edge Comprehensive risk measurement and For risk weights, The probability of exceeding the continuous water supply default, quantile level Under the following conditions, tail expectation For risk level parameters, Risk scaling factor For the Sigmoid function, This refers to the gating sensitivity parameter;

[0120] Subsequently, in the constraint-aware spatiotemporal hypergraph Transformer, the ungated attention weights are first calculated. The gating weights are obtained by implementing hard shielding and soft scaling. ;

[0121] in For super-edge Ungated attention weights For attention scoring function, For hyper-edge embedding, For the set of strategy parameters, The gating weights are used to drive the forward computation to form a gating hypergraph representation for use in hierarchical multi-agent policy inference.

[0122] In this specific embodiment, S4 specifically refers to:

[0123] Gated back hypergraph representation Constraint parameter set With strategy parameters For the input, a hierarchical input embedding is constructed for each agent, denoted as:

[0124] and ;

[0125] in For intelligent agents High-level input embedding, Embedding for low-level input, and For embedding generator functions, For the constraint set Extracted with intelligent agents Related context, For gating hypergraph representation, These are the policy parameters for the hierarchical multi-agent graph Transformer;

[0126] In high-level strategy reasoning, the joint generation of supply tasks and time windows can be written as follows:

[0127] , ;

[0128] in Select distribution for task For candidate tasks, For selected tasks, For task sets, For high-level weight matrix, For time window vector, and These are the earliest start time and the latest finish time, respectively. Assign a function to the time window;

[0129] In low-level policy reasoning, atomic actions such as generating paths and connections are represented as follows:

[0130] and ;

[0131] in For low-level motion distribution, For candidate atomic actions, Define actions for lower levels, For action sets, This is the low-level weight matrix;

[0132] Finally, by combining and aligning high-level tasks and time windows with low-level atomic actions, we obtain:

[0133] and ;

[0134] in For intelligent agents Candidate action vectors, For combination operators, For intelligent agent set, This is a set of candidate actions for subsequent security filtering.

[0135] In this specific embodiment, S5 specifically includes:

[0136] The candidate action set is numerically encoded into a decision vector, and a quadratic programming model is established for safety filtering. First, let the encoded candidate vector be... and with Let be the decision vector to be projected, and let be... As the time window relaxation vector, with For the symmetric positive definite weight matrix of the action deviation, To determine the penalty coefficient for the time window delay, construct the objective function:

[0137] ;

[0138] in For vectors The square of the second norm;

[0139] Constraints, expressed in linear or linearized form, relate continuous water supply, minimum inventory, time window, restricted area, and job sequence, as follows:

[0140] ;

[0141] in A linear mapping matrix from decision vector to arrival / completion time. and These are the lower / upper bound vectors of the time window, respectively. and This represents the coefficient matrix obtained by linearizing the continuous water supply constraint and its lower bound. and The coefficient matrix representing the minimum inventory threshold and its lower bound. and The coefficient matrix representing the no-entry zone constraint and its upper bound, and The coefficient matrix and thresholds representing sequence dependencies (order and minimum margin);

[0142] The preliminary projected solution is obtained by solving the problem. A feasibility test is then performed. If a violation is found due to linearization error or discretization restoration, the nearest projection correction is executed. ;

[0143] in The feasible region defined by the above constraints will ultimately be Based on the defined motion parameters, the movement is restored to a safe and feasible action. ;

[0144] in This is a mapping from vectors to command parameters (including target location, target node, time window, and resource quantity).

[0145] In this specific embodiment, S6 specifically refers to:

[0146] Based on the safety action sequence obtained in step S5 and executed on-site, along with sensor feedback, the strategy is updated using closed-loop risk reinforcement learning and the world model is corrected. First, the instantaneous risk cost at each decision moment is defined as:

[0147] ;

[0148] in For a moment immediate cost, The basic costs consist of travel time, connection time, and energy consumption. For the degree of continuous water supply interruption, For late arrivals exceeding the time window, and The penalties for water supply and lateness are respectively. For time indexing;

[0149] The cumulative discounted cost during the planning period is then expressed as follows:

[0150] ;

[0151] in The cumulative cost from the initial moment, The discount factor has a value range of . The length of the decision scrolling window;

[0152] We construct an optimization objective by minimizing the conditional tail expectation of the Rockafellar-Uryasev form on the risk objective:

[0153] ;

[0154] in To optimize the objective function, For hierarchical multi-agent graph Transformer and constraint gating, the set of policy parameters... Auxiliary variables corresponding to quantile levels The range of risk quantile levels is as follows: For the expected operator on the trajectory distribution induced by policy and environment (including security filters), For positive part operators;

[0155] The strategy parameters and auxiliary variables can be updated using the gradient method as follows:

[0156] and ;

[0157] in For policy learning rate, The learning rate is an auxiliary variable.

[0158] The training data comes from an empirical set consisting of states and actions, and is sampled preferentially based on tail samples. Simultaneously, maximum likelihood correction is applied to the world model parameters to eliminate distribution drift, and the following methods are used:

[0159] Minimize and update;

[0160] in For the world model loss function, For world model parameters, In the state Below the target variable conditional probability, For a moment Environment or hypergraph embedding state, The observed variables include water supply availability and arrival time.

[0161] After completing the above updates, the risk gating threshold and penalty coefficient are simultaneously calibrated, and new policy parameters are output for the next round of closed-loop operation of perception-prediction-gating-decision-security filtering-execution.

[0162] In this specific embodiment, S7 specifically refers to:

[0163] Based on the updated strategy parameters and on-site execution data from step S6, command visualization, task suggestions, and compliance audit results are generated and archived, and the linkage plan is triggered. First, a comprehensive score is calculated in the situation assessment module for commanders to quickly assess, written as:

[0164] ;

[0165] in For comprehensive situational assessment, Weights for each evaluation indicator For water supply security rate, For a moment The extent of water supply interruption For expectation operator, To ensure on-time arrival rate, For probability operators, For arrival time, The upper limit of the time window for a task or node. To constrain compliance rate, For the number of default events, For the number of jobs already executed, Average cost per unit task;

[0166] When the overall score falls below the threshold or a supply-demand imbalance is predicted, the system automatically calculates the reinforcement needs and generates contingency plans. The minimum number of reinforcements is given as follows:

[0167] ;

[0168] in For the number of additional supplies or pump trucks that need to be dispatched, For the floor operator, For the maximum value operator, For total demand, For the assembly of on-site combat units, For combat units unit time demand, For current supply capacity, To effectively replenish the superedge set, For super-edge Traffic limit The nominal supply capacity for a single vehicle or a single supply unit;

[0169] To unify visualization and alarm triggering, the scores are mapped to alarm magnitudes and pushed to the command terminal, the following definition is made:

[0170] ;

[0171] in For alarm amplitude, As a scoring threshold;

[0172] Ultimately, the recommended actions, reinforcement plans, and audit results are packaged into an instruction set, marked with a unique index and timestamp, and pushed to vehicles and combat units according to the permission matrix. At the same time, they are archived in logs and audit repositories for review and compliance filing.

[0173] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A fire emergency dispatch optimization method based on deep reinforcement learning, characterized in that, Includes the following steps: S1. Collect vehicle positioning, road traffic status, building topology, water source information, water pressure sensor data, consumable inventory data and combat mission progress data. Based on the data, construct an initial spatiotemporal hypergraph. Generate a fire protection red line set and a constraint parameter set based on regulations and operating procedures. Initialize the strategy parameters of the hierarchical multi-agent graph Transformer. Output the initial spatiotemporal hypergraph, fire protection red line set, constraint parameter set and strategy parameters. S2. Using the initial spatiotemporal hypergraph as input, the world model is used to predict the distribution of consumable consumption rate, arrival time, water availability and inventory balance, and to obtain the predicted distribution set, quantile set and risk score set. S3. The initial spatiotemporal hypergraph, fire red line set, constraint parameter set, strategy parameter, prediction distribution set, quantile set and risk score set are taken as inputs. Constraint gating is performed in the attention calculation of the constraint-aware spatiotemporal hypergraph Transformer to obtain the gated hypergraph representation. S4. Using the gated hypergraph representation and policy parameters as input, a hierarchical multi-agent graph Transformer is used for policy reasoning. The higher layer outputs the supply task selection and time window allocation decision, and the lower layer outputs the path planning and connection execution actions. The combination yields a set of candidate actions. S5. Taking the candidate action set and constraint parameter set as input, the project solution is performed in the quadratic programming safety filtering module based on the continuous water supply constraint parameters, minimum inventory threshold parameters, time window parameters, restricted area parameters and sequence dependency parameters to obtain safe and feasible actions. S6. Execute safe and feasible actions and strategy parameters on-site and collect execution feedback data; S7. Using the execution feedback data and policy parameters as input, CVaR risk reinforcement learning is used to calculate the conditional risk loss of the supply interruption duration and late timeout, and to update the policy parameters of the high-level and low-level layers to obtain the updated policy parameters for policy inference in subsequent time steps.

2. The fire emergency dispatch optimization method based on deep reinforcement learning according to claim 1, characterized in that, Step S1 is as follows: The system aligns vehicle positioning, road traffic status, building topology, water source information, water pressure sensor data, consumable inventory data, and combat mission progress data with timestamps and coordinates, fills in missing data and removes anomalies to obtain aligned multi-source situational data and outputs the aligned multi-source situational data. Using aligned multi-source situational awareness data as input, an initial spatiotemporal hypergraph is constructed and output. The initial spatiotemporal hypergraph includes a set of nodes and a set of hyperedges. The set of nodes includes water source nodes, temporary connection nodes, hose confluence nodes, pump truck nodes, combat unit nodes, and supply station nodes. The set of hyperedges is used to represent the supply link from the water source node to the combat unit node. Hyperedge features include water pressure, flow limit, altitude difference, pipe diameter, hose laying length, number of joints, connection time, laying time, and current occupancy status. The water pressure and flow limit are calculated or estimated based on water pressure sensor data and pipe diameter. The altitude difference is calculated based on terrain elevation or building elevation. The hose laying length is estimated based on the length of the passable path between nodes. The connection time and laying time are estimated based on distance and operating conditions. The current occupancy status is determined based on the on-site equipment and mission progress. The initial spatiotemporal hypergraph is used as input to generate a set of fire protection red lines based on regulations, operating procedures, and on-site control information, and the set of fire protection red lines is output. Simultaneously, based on the fire protection red line set and the aligned multi-source situational data, a set of constraint parameters is derived and output. The set of constraint parameters includes continuous water supply constraint parameters, minimum inventory threshold parameters, time window parameters, restricted area parameters, and sequence dependency parameters. Among them, the continuous water supply constraint parameters are set according to the relationship between water pressure and flow rate limits and consumable consumption; the minimum inventory threshold parameters are set according to consumable inventory data; the time window parameters are set according to road traffic status and combat mission progress; the restricted area parameters are set according to the spatial range of the fire protection red line set; and the sequence dependency parameters are set according to the operation sequence of connection and laying. The initial spatiotemporal hypergraph, the fire protection red line set, and the constraint parameter set are used as inputs to initialize the policy parameters of the hierarchical multi-agent graph Transformer and output the policy parameters. The initialization includes assigning initial values ​​to the network weights, the gating threshold corresponding to the fire protection red line set, and the mask corresponding to the constraint parameter set in the policy parameters, so that the policy parameters are consistent with the initial spatiotemporal hypergraph, the fire protection red line set, and the constraint parameter set.

3. The fire emergency dispatch optimization method based on deep reinforcement learning according to claim 1, characterized in that, Step S2 is as follows: The initial spatiotemporal hypergraph is used as input. The node features and hyperedge features in the initial spatiotemporal hypergraph are encoded to form the model input sequence for the world model and output the model input sequence. The model input sequence is used as input, and the encoder and state-space dynamics in the world model (DreamerV3) are used to perform time-series inference on the model input sequence to obtain the latent state sequence and output the latent state sequence. Using the potential state sequence as input, parameterized probability modeling of consumable consumption rate, arrival time, water availability and inventory balance is performed through the corresponding decoder to generate a set of predicted distributions covering the variables and output the set of predicted distributions. The predicted distribution set is used as input, and the quantile set is calculated based on the target quantile set. Combined with the fire protection red line set and the constraint parameter set, the risk scores of continuous water supply default and time window default are calculated based on the quantile exceedance probability or the conditional tail expectation, forming a risk score set and outputting the quantile set and risk score set. The predicted distribution set, quantile set, and risk score set are output together with the initial spatiotemporal hypermap, fire protection red line set, constraint parameter set, and strategy parameters.

4. The fire emergency dispatch optimization method based on deep reinforcement learning according to claim 1, characterized in that, Step S3 is as follows: The initial spatiotemporal hypergraph, fire protection red line set, constraint parameter set, strategy parameter, prediction distribution set, quantile set, and risk score set are taken as input. Based on the fire protection red line set and constraint parameter set, the constraints of each hyperedge in the initial spatiotemporal hypergraph are judged to determine whether they violate the continuous water supply constraint, minimum inventory threshold, time window constraint, restricted area constraint, and sequence dependency constraint, forming a constraint mask set and outputting the constraint mask set. Simultaneously, risk scaling coefficients are calculated for each hyperedge based on the quantile set and risk score set, forming a risk scaling coefficient set, which is then output together with the initial spatiotemporal hypergraph, constraint parameter set, and strategy parameters. The initial spatiotemporal hypergraph, constraint parameter set, policy parameter, constraint mask set, and risk scaling factor set are taken as input. Attention is calculated on the initial spatiotemporal hypergraph in the constraint-aware spatiotemporal hypergraph Transformer to generate an ungated attention weight set. The ungated attention weight set, constraint mask set, risk scaling factor set, initial spatiotemporal hypergraph, constraint parameter set, and policy parameter set are output together. The ungated attention weight set, constraint mask set, risk scaling factor set, initial spatiotemporal hypergraph, constraint parameter set and policy parameters are taken as input. The attention weights corresponding to hyperedges that violate the constraint mask set are hard masked and the corresponding weights are reset to zero. The attention weights that are not hard-masked are soft-scaled according to the risk scaling factor set to obtain the gated attention weight set. The gated attention weight set is then output together with the initial spatiotemporal hypergraph, the constraint parameter set, and the policy parameters. The gated attention weight set, the initial spatiotemporal hypergraph, and the policy parameters are taken as input and forward computation is performed in the constraint-aware spatiotemporal hypergraph Transformer to form the gated hypergraph representation. The gated hypergraph representation, the constraint parameter set, and the policy parameters are then output.

5. The fire emergency dispatch optimization method based on deep reinforcement learning according to claim 1, characterized in that, Step S4 is as follows: Using the gated hypergraph representation, constraint parameter set, and policy parameters as input, input embeddings are constructed for each agent in the hierarchical multi-agent graph Transformer based on the gated hypergraph representation. These are used to form high-level input embeddings for high-level reasoning and low-level input embeddings for low-level reasoning, and the outputs are high-level input embeddings, low-level input embeddings, constraint parameter set, and policy parameters. The high-level input embedding, constraint parameter set, and policy parameters are used as inputs. Policy reasoning is performed at the high level of the hierarchical multi-agent graph Transformer to generate supply task selection and time window allocation decisions, forming a high-level decision set and outputting the high-level decision set, constraint parameter set, and policy parameters. The high-level decision set, as well as the low-level input embedding, constraint parameter set, and policy parameter set, are used as inputs. Policy reasoning is performed at the low level of the hierarchical multi-agent graph Transformer to generate path planning and connection execution actions, forming a low-level action set and outputting the low-level action set, constraint parameter set, and policy parameter set. The set of low-level actions, the set of high-level decisions, the set of constraint parameters, and the set of policy parameters are taken as input, combined and aligned to obtain a set of candidate actions, and the set of candidate actions, the set of constraint parameters, and the set of policy parameters are output.

6. The fire emergency dispatch optimization method based on deep reinforcement learning according to claim 1, characterized in that, Step S5 is as follows: The candidate action set, constraint parameter set, and strategy parameters are taken as input. In the quadratic programming safety filtering module, the candidate action set is numerically encoded to establish a quadratic programming model and output the quadratic programming model and strategy parameters. The decision variables of the quadratic programming model are used to represent the action parameters of the candidate action set. The objective function is used to minimize the deviation of the decision variables from the candidate action set and simultaneously minimize the weighted value of the time window delay. The constraints are set according to the constraint parameter set, including continuous water supply constraint parameters, minimum inventory threshold parameters, time window parameters, restricted area parameters, and sequence dependency parameters. The quadratic programming model and policy parameters are used as inputs. The projection is solved in the quadratic programming security filtering module to obtain the preliminary projection action vector, and the preliminary projection action vector and policy parameters are output. The initial projection action vector, the set of constraint parameters, and the strategy parameters are taken as input. The feasibility of the initial projection action vector is checked based on the set of constraint parameters. If it is found that the continuous water supply constraint parameters, minimum inventory threshold parameters, time window parameters, restricted area parameters, or sequence dependency parameters are violated, the nearest projection correction to the feasible domain boundary is performed in the secondary planning safety filtering module to obtain the feasible action vector and output the feasible action vector and strategy parameters. The action vector and policy parameters are taken as input, and the action parameters are restored and aligned according to the action parameter definition of the candidate action set to form safe actions. The safe actions and policy parameters are then output.

7. The fire emergency dispatch optimization method based on deep reinforcement learning according to claim 1, characterized in that, Step S6 is as follows: Taking safe and feasible actions and policy parameters as input, the safe and feasible actions are parsed into instructions to generate a set of execution instructions containing target location, target node, time window, resource quantity and job type. A unified timestamp is generated for each instruction, and the set of execution instructions and policy parameters are output. The execution instruction set and strategy parameters are taken as input. The execution instruction set is sent to the site and status acquisition is started. The site status data related to the execution is collected to form raw execution data. The raw execution data includes at least execution instruction identifier, timestamp, location record, water supply status record and inventory sampling record. The raw execution data and strategy parameters are output. The original execution data and strategy parameters are used as input. Timestamp alignment and anomaly removal are performed. The execution result indicators are calculated based on the execution instruction identifier and timestamp in the original execution data to obtain inventory status, water supply status, arrival time and compliance status. Among them, inventory status represents the remaining quantity of key consumables and the remaining operation time; water supply status represents the continuous water supply indication, water supply interruption duration and water supply capacity assessment value; arrival time represents the time interval from the issuance of the instruction to the arrival at the target location; and compliance status represents whether the execution violates the gating threshold and the spatial and time window constraints corresponding to the mask encoded in the strategy parameters. The output is a set of indicators and strategy parameters. The system takes a set of indicators and strategy parameters as input, aggregates and formats the indicator set to form execution feedback data, which includes at least inventory status, water supply status, arrival time and compliance status. The system outputs execution feedback data and strategy parameters.

8. The fire emergency dispatch optimization method based on deep reinforcement learning according to claim 1, characterized in that, Step S7 is as follows: The system takes execution feedback data and strategy parameters as input, extracts inventory status, water supply status, arrival time and compliance status from the execution feedback data, generates a water supply interruption duration sequence based on the water supply interruption duration in the water supply status, and generates a late arrival timeout sequence based on the arrival time and read the upper limit of the time window parameter used for late arrival calculation from the strategy parameters, forming a risk sample set and outputting the risk sample set and strategy parameters. Using the risk sample set and strategy parameters as input, the risk level parameter is read from the strategy parameter. The quantile threshold corresponding to the risk level parameter is calculated for the supply interruption duration sequence and the late arrival timeout sequence, respectively. The average value of the part exceeding the quantile threshold is further calculated as the conditional risk loss. The conditional risk loss set is obtained and the conditional risk loss set and strategy parameters are output. The conditional risk loss set and strategy parameters are taken as input. The risk weight parameters are read from the strategy parameters. The conditional risk loss set is weighted and synthesized to form the total risk loss. The total risk loss and strategy parameters are then output. Using the total risk loss and policy parameters as input, CVaR risk reinforcement learning is used to update the high-level and low-level policy parameters of the hierarchical multi-agent graph Transformer. The updated policy parameters are obtained by calculating the gradient of the total risk loss and performing backpropagation and iteration. The updated policy parameters are then output for policy inference in subsequent time steps.