Urban waterlogging rapid prediction and disaster-causing mechanism analysis method

By utilizing expert models and hard gating mechanisms in urban flooding prediction, and based on dynamic rainfall characteristics and static physical properties, this method rapidly generates flooding prediction results and determines the disaster-causing mechanisms. This solves the problems of low computational efficiency and poor interpretability in existing technologies, and achieves second-level prediction and quantitative analysis.

CN121724217BActive Publication Date: 2026-04-17NANJING HYDRAULIC RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING HYDRAULIC RES INST
Filing Date
2026-02-11
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies struggle to balance real-time computational efficiency with the interpretability of physical mechanisms when dealing with complex urban environments. Traditional hydrodynamic models are computationally complex and cannot meet the timeliness requirements of rolling forecasts for multiple scenarios under extreme weather conditions. Meanwhile, data-driven models lack physical interpretability and are difficult to directly translate into engineering decision-making basis.

Method used

By acquiring dynamic rainfall characteristics and static physical properties of the target area, spatial clusters are determined based on the static physical properties of computing nodes. Expert models are invoked to predict urban flooding. Hard gating mechanisms are used to reduce computational overhead. Feature contribution values ​​are calculated through interpretable algorithms. The system dominance of each physical system is aggregated to determine the type of disaster-causing mechanism.

Benefits of technology

It achieves second-level prediction of urban flooding and quantitative analysis of the disaster-causing mechanism, provides engineering-operable flood control strategies, and solves the problems of low computational efficiency and poor physical interpretability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121724217B_ABST
    Figure CN121724217B_ABST
Patent Text Reader

Abstract

The application discloses a kind of urban waterlogging rapid prediction and disaster-causing mechanism analysis method.The method obtains dynamic rainfall characteristics and the static physical properties of multiple computing nodes;Determine the spatial cluster to which the computing node belongs based on the static physical properties of the computing node, call the preset expert model corresponding to the spatial cluster, generate the waterlogging prediction result of the computing node based on the dynamic rainfall characteristics and the static physical properties;Based on the expert model, calculate the feature contribution value of dynamic rainfall characteristics and static physical properties to the waterlogging prediction result, aggregate it according to the physical system attribution of static physical properties, obtain the system dominance corresponding to each physical system;According to the comparative relationship of the system dominance of each physical system, determine the disaster-causing mechanism type of the computing node.The application solves the problems of low computing efficiency and poor physical interpretability, realizes the second-level prediction of urban waterlogging and the quantitative analysis of disaster-causing mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to flood disaster prediction technology, and in particular to methods for rapid prediction of urban flooding and analysis of its causative mechanisms. Background Technology

[0002] Currently, urban flooding is frequent, threatening the safe operation of cities. Real-time flood forecasting, which identifies the causes of disasters and enables the development of targeted flood prevention strategies, can advance the process of smart city safety management.

[0003] Currently, urban flood risk assessment mainly adopts the following technical approaches: First, hydrodynamic models based on physical mechanisms (such as InfoWorks Integrated Watershed Management Model and SWMM) simulate surface runoff generation and drainage and pipeline hydraulic processes by solving the Saint-Venant equations; Second, data-driven deep learning proxy models (such as Convolutional Neural Network (CNN) and Long Short-Term Memory Network (LSTM)) train neural networks using historical rainfall-inundation data to establish a nonlinear mapping from meteorological input to water depth.

[0004] However, existing technologies struggle to balance real-time computational efficiency with the interpretability of physical mechanisms when dealing with complex urban environments. Therefore, further research and innovation are needed to address these issues with current technologies. Summary of the Invention

[0005] Objective of the invention: In view of the above-mentioned problems in the prior art, this application provides a method for rapid prediction of urban flooding and analysis of its disaster-causing mechanisms.

[0006] According to one aspect of this application, a method for rapid prediction and disaster-causing mechanism analysis of urban flooding includes:

[0007] The dynamic rainfall characteristics of the target area and the static physical properties distributed across multiple computing nodes within the target area are obtained. The static physical properties belong to multiple physical systems.

[0008] The spatial cluster to which the computing node belongs is determined based on the static physical properties of the computing node. The expert model corresponding to the spatial cluster is called, and the waterlogging prediction result of the computing node is generated based on the dynamic rainfall characteristics and static physical properties.

[0009] Based on the expert model, the characteristic contribution values ​​of dynamic rainfall characteristics and static physical attributes to the urban flooding prediction results are calculated. The characteristic contribution values ​​are aggregated according to the physical system affiliation of the static physical attributes to obtain the system dominance degree of each physical system.

[0010] Based on the comparison of the system dominance of each physical system, the disaster-causing mechanism type of the computing node is determined.

[0011] Beneficial effects: This invention solves the problems of low computational efficiency and poor physical interpretability, achieving second-level prediction of urban flooding and quantitative analysis of its disaster-causing mechanisms. The related technical effects will be described in detail below with reference to specific embodiments. Attached Figure Description

[0012] Figure 1 A flowchart illustrating a method for rapid prediction and disaster mechanism analysis of urban flooding, provided in an embodiment of this application.

[0013] Figure 2 This is a flowchart illustrating how to determine the spatial cluster to which a computing node belongs based on its static physical attributes, as provided in an embodiment of this application.

[0014] Figure 3 This is a flowchart illustrating how feature contribution values ​​are aggregated according to the physical system affiliation of static physical attributes to obtain the system dominance of each physical system, as provided in this embodiment of the application.

[0015] Figure 4 A flowchart for determining the system dominance of a physical system based on the system influence amplitude, provided for an embodiment of this application.

[0016] Figure 5 This is a flowchart illustrating how a disaster mechanism type of a computing node is determined based on a comparison of the system dominance of each physical system, as provided in an embodiment of this application. Detailed Implementation

[0017] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0018] It should be understood that the terms "first," "second," etc., in the specification and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such process, method, product, or apparatus.

[0019] To address the aforementioned issues, the applicant conducted in-depth searches and analyses, and discovered:

[0020] Traditional hydrodynamic models have high computational complexity, with a single simulation taking several hours, which cannot meet the timeliness requirements of rolling forecasts for multiple scenarios under extreme weather conditions. Although existing data-driven models are computationally fast, they are mostly black-box models, and conventional interpretability methods (such as SHAP and Shapley interpretation methods) only focus on the numerical contribution analysis at the microscopic feature level.

[0021] At the physical system level, it is impossible to decouple the two completely different disaster-causing mechanisms: one is insufficient pipeline transportation capacity, which is a structural bottleneck, and the other is river water level backflow, which is a functional bottleneck. This makes it difficult to directly transform the prediction results into decision-making basis for optimizing blue-green-grey engineering facilities.

[0022] To solve this problem, we need to combine... Figures 1 to 5 The present invention will be specifically described through the following embodiments.

[0023] In some embodiments, an exemplary scheme is provided for rapid prediction and disaster mechanism analysis of urban flooding. This can simultaneously solve the problems of low computational efficiency in traditional hydrodynamic models and the lack of physical interpretability in purely data-driven models. Specifically, this scheme includes:

[0024] Step 101: Obtain the dynamic rainfall characteristics of the target area and the static physical attributes of multiple computing nodes distributed within the target area. The static physical attributes are assigned to multiple predefined physical systems according to their physical meaning.

[0025] Alternatively, it can be described as obtaining the dynamic rainfall characteristics of the target area and the static physical properties distributed across multiple computing nodes within the target area, where the static physical properties belong to multiple physical systems.

[0026] In this embodiment, dynamic rainfall characteristics refer to meteorological driving data that changes dynamically over time. Specifically, these characteristics may include rainfall duration, total rainfall, cumulative rainfall within a specific period, and statistical characteristics of the rainfall process. These characteristics collectively describe the input intensity of external meteorological conditions on urban flooding.

[0027] Computational nodes refer to the basic units after discretizing continuous urban space. Their specific forms can be nodes of unstructured grids, center points of regular grids, or centroids of catchment areas. Static physical attributes refer to inherent properties related to the geographical environment and remaining stable in the short term. Specifically, these attributes are explicitly categorized into predefined physical systems, such as topographic systems, pipe network systems, and water systems. This clear classification based on physical meaning lays the data foundation for subsequent analysis of the causes of urban flooding from a physical mechanism perspective. For example, elevation data belongs to the topographic system, while pipe diameter belongs to the pipe network system. This classification is not only a data organization method but also a logical premise for subsequent attribution analysis.

[0028] In this embodiment, the dynamic rainfall characteristics specifically include the following 10 dimensions: total rainfall (mm), rainfall duration (h), peak rainfall intensity (mm / h), peak occurrence time ratio (peak time / total duration), cumulative rainfall in the first hour (mm), cumulative rainfall in the first 3 hours (mm), cumulative rainfall in the first 6 hours (mm), standard deviation of rainfall intensity (mm / h), skewness coefficient of the rainfall process curve, and kurtosis coefficient of the rainfall process curve. These characteristics characterize the dynamic properties of the rainfall event from three dimensions: temporal distribution, intensity characteristics, and statistical properties.

[0029] Step 102: Determine the spatial cluster to which the computing node belongs based on the static physical properties of the computing node, call the pre-set expert model corresponding to the spatial cluster, and generate the waterlogging prediction result of the computing node based on dynamic rainfall characteristics and static physical properties.

[0030] In this step, spatial partitioning is performed using static physical properties rather than dynamic response data. Specifically, based on the prior knowledge that regions with similar physical properties will have similar hydrological responses under the same rainfall excitation, the global computing nodes are divided into several spatial clusters.

[0031] For each spatial cluster, the system pre-configures a specially trained expert model. During the online prediction phase, the system determines the cluster to which a computing node belongs based on its static attributes and directly calls the expert model corresponding to that cluster for inference. This mechanism is known as a hard gating mechanism.

[0032] Compared to a single global model, this zonal modeling strategy can better capture the nonlinear characteristics of local terrain and pipe networks; compared to a hybrid expert model with dynamic gating, hard gating based on static attributes avoids running complex gating networks during the inference phase, reducing computational overhead and achieving second-level prediction. Urban flooding prediction results typically refer to the maximum inundation depth of the computing nodes under a specific rainfall scenario.

[0033] Step 103: Calculate the feature contribution values ​​of dynamic rainfall characteristics and static physical attributes to the urban flooding prediction results based on the expert model, and aggregate the feature contribution values ​​according to the physical system affiliation of the static physical attributes to obtain the system dominance degree of each physical system.

[0034] Specifically, the system uses an interpretability algorithm to calculate the marginal contribution of each input feature to the prediction result, i.e., the feature contribution value. Further, the system performs a physical system-level aggregation operation. Based on physical affiliation, the system aggregates the contribution values ​​of all static features belonging to the same physical system (such as a pipeline system) to obtain the overall influence of that physical system, i.e., the system dominance.

[0035] In summary, this process achieves a semantic leap from the feature layer to the physical mechanism layer. For example, by aggregating the contributions of features such as pipe density and pipe distance, the degree of control of the pipe network system as a whole over urban flooding can be quantified.

[0036] Step 104: Determine the disaster-causing mechanism type of the computing node based on the comparison of the system dominance of each physical system.

[0037] Building upon this foundation, this step identifies the primary disaster-causing factors by comparing the system dominance of different physical systems. Specifically, if the system dominance of a particular physical system is significantly higher than that of other systems, it is determined that the flooding in that area is mainly caused by that physical system. For example, if the pipe network system has the highest system dominance, it is determined that the flooding is caused by an internal structural bottleneck; if the water system system has the highest system dominance, it is determined that the flooding is caused by an external functional bottleneck. This determination directly reflects the engineering problems in urban flood control and drainage, providing a basis for subsequent implementation of targeted blue, green, and gray engineering measures.

[0038] Furthermore, the method for rapid prediction and disaster-causing mechanism analysis of urban flooding can also be as follows: obtain dynamic rainfall characteristics and static physical attributes of terrain, pipe network, and water system classified according to physical meaning; cluster the computing nodes into spatial clusters based on static physical attributes, and use a hard gating mechanism to call the corresponding expert model to quickly generate flooding prediction results; adopt a hierarchical dominance attribution framework to calculate the contribution value of input features and aggregate them according to physical system affiliation to obtain the system dominance of each physical system; determine the disaster-causing mechanism type based on the dominance comparison, such as internal structural bottlenecks, external functional bottlenecks, etc.

[0039] In other embodiments, optional implementations of offline data preparation and expert model construction are described for achieving rapid online prediction and attribution. Accordingly, this embodiment includes:

[0040] Step 201, the physical system includes the terrain system, the pipeline system, and the water system;

[0041] Among them, the static physical properties belonging to the terrain system include elevation, slope and impermeability coefficient;

[0042] Static physical properties belonging to the pipeline network system include distance from manholes, pipe length density per unit area, and pipe volume density per unit area.

[0043] The static physical properties belonging to a water system include the distance to the nearest river, the water area per unit area of ​​the river channel, and the reservoir area per unit area.

[0044] In this embodiment, nine static physical properties across three categories were selected to characterize the physical properties of the urban underlying surface. Among them, the topography system reflects the natural conditions for surface runoff generation and confluence, with elevation determining the potential energy distribution of water, slope affecting runoff velocity, and impermeability coefficient determining runoff volume.

[0045] The pipeline system reflects the capacity of underground artificial drainage facilities. The distance from the inspection well reflects the path length of surface runoff into the underground pipeline network. The pipe length density and pipe volume density per unit area directly quantify the drainage, transportation, and storage capacity of the region.

[0046] Among these, the water system reflects the boundary conditions of the receiving water body. The distance to the nearest river affects the backwater effect of the drainage outlet, while the water area per unit area of ​​the river channel and the reservoir area per unit area reflect the region's regulation and flood control potential. These attributes can typically be extracted from urban geographic information system (GIS) data, digital elevation models (DEMs), and pipeline survey data.

[0047] Step 202: The expert model is pre-trained using supervised learning with the training dataset; wherein, the labels of the waterlogging prediction results in the training dataset are generated based on the simulation calculation results of the coupled hydrological-hydrodynamic physical model under multiple rainfall scenarios, and the simulation calculation results include the maximum inundation depth of the calculation nodes.

[0048] To obtain high-quality data for training expert models, physical model simulations can be used to generate ground truth values. Specifically, a mechanistic model coupling hydrological processes and two-dimensional hydrodynamic processes can be constructed, such as a model built on the InfoWorks ICM platform. This model integrates a rainfall-runoff module, a one-dimensional pipe network flow module, and a two-dimensional surface runoff module, which can realistically simulate the complex interaction processes of rainfall, pipe network drainage, and surface inundation.

[0049] Using this physical model, multiple sets of rainfall scenarios with different return periods and durations were designed for simulation. For example, a total of 45 combined scenarios could be designed, including short-duration (e.g., 1 to 3 hours) sudden heavy rain and long-duration (e.g., 6 to 72 hours) continuous rainfall.

[0050] By simulating calculations, the maximum inundation depth of each computing node under each scenario is extracted and used as the target label for supervised learning. This method of generating labels based on a physical model overcomes the problem of scarce grid-level urban flooding monitoring data in reality, ensuring that the patterns learned by the expert model conform to hydrodynamic mechanisms.

[0051] Step 203, the construction of the expert model, includes: performing clustering processing based on the static physical properties of the sample data to generate multiple initial clusters; counting the number of nodes contained in each initial cluster; if there is an initial cluster with fewer nodes than a preset node number threshold, calculating the Euclidean distance between the cluster center of the initial cluster and the cluster centers of the other initial clusters; merging the initial cluster into the target cluster with the minimum Euclidean distance, until the number of nodes in all clusters is not less than the node number threshold, to obtain spatial clusters, and training the corresponding expert model for each spatial cluster.

[0052] In other words, clustering is performed based on the static physical properties of the sample data to generate multiple initial clusters;

[0053] Count the number of nodes contained in each initial cluster;

[0054] If there exists an initial cluster with fewer nodes than the node count threshold, then calculate the Euclidean distance between the cluster center of this initial cluster and the cluster centers of the other initial clusters.

[0055] The initial cluster is merged into the target cluster with the minimum Euclidean distance until the number of nodes in all clusters is not less than the node number threshold, thus obtaining spatial clusters. For each spatial cluster, a corresponding expert model is trained.

[0056] Alternatively, the sample data, including the static physical properties of each computing node and the corresponding urban flooding simulation results, is obtained from historical rainfall scenarios.

[0057] Furthermore, the system can employ a greedy strategy to handle such small clusters. Alternatively, the system processes each initial cluster sequentially in ascending order of the number of nodes. For each small cluster to be merged, the Euclidean distance between its cluster center vector and the center vectors of all other clusters is calculated. The system finds the target cluster with the smallest distance and merges all nodes from that small cluster into that target cluster.

[0058] After merging, the system updates the cluster center of the target cluster based on the number of nodes in the two clusters before the merge. This process is repeated until all clusters have 1000 or more nodes. The resulting spatial clusters ensure both similarity in physical properties within the clusters and sufficient sample size for model training in each cluster. For each merged spatial cluster, an expert model is trained using data from all nodes within that cluster under all rainfall scenarios. The expert model can optionally employ a gradient boosting decision tree-based algorithm, such as the LightGBM (Lightweight Gradient Boosting Machine) model, to handle nonlinear regression problems.

[0059] Optionally, when training an expert model, the Tree Structure Parzen Estimator (TPE) algorithm can be used to automatically optimize the model's hyperparameters, such as optimizing the number of trees, maximum depth, and learning rate, to further improve the model's prediction accuracy.

[0060] Building upon this, this embodiment also involves the storage and application of standardized parameters. Accordingly, before calculating the Euclidean distance between the static physical properties of the computing nodes and the cluster center vectors, the embodiment further includes: obtaining the mean and standard deviation of all pre-stored samples on each static physical property; and performing Z-score standardization on the static physical properties of the computing nodes based on the mean and standard deviation.

[0061] Specifically, for the j-th static physical attribute, let μ be the mean of this attribute for all samples in the training set. _j The standard deviation is σ _j For any compute node's original attribute value v _j Its standardized value z _j The calculation formula is: z _j =(v _j -μ _j ) / σ _j This type of statistic μ _j and σ _j The data is calculated and stored offline, and then used in the subsequent online phase to perform the same transformation on new input static physical properties, ensuring the mathematical validity of the distance calculation.

[0062] In other words, the Z-score standardization formula can be described as:

[0063] z _j =(v _j -μ _j ) / σ _j ;

[0064] Among them, z _j v is the standardized value of the j-th static physical property; _j μ is the original value of the computation node on the j-th static physical property;_j σ is the arithmetic mean of all samples in the training dataset on the j-th static physical attribute; _j Let $\begin{pmatrix}$ be the standard deviation of all samples in the training dataset on the $j$-th static physical attribute.

[0065] This step describes the spatial partitioning process, particularly the greedy merging strategy for small clusters. Accordingly, the static physical properties of all computing nodes are Z-score standardized, and the K-means clustering algorithm can be used for initial clustering. It is assumed that the initial number of clusters is set to 50. Due to the spatial heterogeneity of cities, clustering may produce small clusters with a small number of nodes, leading to the risk of overfitting due to insufficient data in subsequent models trained on these clusters. Therefore, this embodiment introduces a node count threshold, for example, set to 1000. The system counts the number of nodes in each initial cluster and identifies all small clusters with fewer than 1000 nodes.

[0066] Optionally, the choice of the node number threshold needs to balance the factors of avoiding overfitting and preserving spatial resolution. On the one hand, to avoid overfitting, sufficient training data for the expert model must be ensured. According to statistical experience, the effective number of samples should be at least 10 to 50 times the feature dimension. This scheme uses 19-dimensional input features, so each cluster requires a minimum of approximately 200 to 1000 samples.

[0067] On the other hand, to preserve spatial resolution, if the threshold is too high, too many small clusters will be merged, resulting in a loss of spatial heterogeneity information. In this embodiment, the node number threshold is set to 1000, which is approximately 60% of the average cluster size (approximately 1600 nodes), thus filtering out smaller clusters.

[0068] At this point, each cluster has approximately 800 training samples, which, with an 80% training set ratio, meets the requirements for 19-dimensional feature modeling, achieving a good balance between model reliability and spatial precision. In practical applications, it is recommended to perform cross-validation within the following set range, selecting the configuration with the lowest root mean square error (RMSE) on the test set.

[0069] The above refers to the total number of samples that meet the requirements of data integrity, have no abnormal deviations (such as removing samples with missing key parameters or those exceeding the reasonable range of values), and can accurately reflect the characteristics of the physical system.

[0070] Alternatively, the cluster center weighted update formula can be expressed as:

[0071] μ _new =(|C _target |×μ _target +|C _i |×μ _i ) / (|C _target |+|C_i |);

[0072] Where, μ _new The cluster center vector of the newly generated spatial cluster after merging; |C _target | represents the number of nodes contained in the target cluster (i.e., the large cluster) before merging; μ _target |C is the cluster center vector of the target cluster before merging; _i | represents the number of nodes contained in the cluster to be merged; μ _i Let be the cluster center vector of the small clusters to be merged; |...| represents taking the absolute value.

[0073] In this embodiment, a Physically Guided Hybrid Expert Clustering (CMoE) architecture is employed. By using offline spatial partitioning based on static physical attributes such as terrain, pipelines, and water systems, combined with a hard gating mechanism during online inference, the system can directly call the best-matching local expert model based on physical attributes. This divide-and-conquer strategy avoids the high-dimensional solution process of the global physical model, compressing the computation time from hours to seconds, and enabling real-time multi-scenario rolling forecasts.

[0074] In another embodiment, the specific implementation process of the hard-gated online prediction method is described, particularly in the online application stage, how the system determines the applicable expert model based on the static physical properties of the computing nodes to achieve millisecond-level flood prediction. This embodiment can improve computational efficiency while ensuring adaptability to spatial heterogeneity. Accordingly, this embodiment can be carried out using the following steps:

[0075] Step 301, determining the spatial cluster to which the computing node belongs based on its static physical attributes, including:

[0076] Obtain the pre-stored cluster center vectors corresponding to multiple candidate spatial clusters; solve for the Euclidean distance between the static physical properties of the computation nodes and each cluster center vector;

[0077] The candidate spatial clusters corresponding to the cluster center vectors that have the minimum Euclidean distance to the static physical properties are determined as the spatial clusters to which the computation nodes belong. This is used to implement spatial divide-and-conquer prediction.

[0078] Accordingly, the system loads the spatial partitioning information constructed during the offline training phase, specifically including the cluster center vectors of K candidate spatial clusters. These cluster center vectors are the centroids representing the mean of the physical properties of each cluster in the high-dimensional feature space.

[0079] Based on this, it is essential to ensure that the distribution of online data is consistent with that of offline training data. Therefore, before calculating the Euclidean distance between the static physical attributes of the computing nodes and the cluster center vectors, the process includes: obtaining the mean and standard deviation of all pre-stored samples on each static physical attribute; performing Z-score standardization on the static physical attributes of the computing nodes based on the mean and standard deviation; and accordingly, calculating the Euclidean distance, specifically the Euclidean distance between the standardized static physical attributes and the cluster center vectors.

[0080] In practice, for each computation node to be predicted, the system reads its 9-dimensional static physical attribute vector. Further, using the pre-stored global mean and global standard deviation, each dimension of this vector is standardized to generate a standardized feature vector x. Then, the system calculates the Euclidean distance d between this vector and the vector of each cluster center. _k The calculation formula can be expressed as:

[0081] d _k =sqrt(∑((x _j -C _kj ) 2 ));

[0082] In the formula, x _j C represents the normalized value of the node on the j-th attribute. _kj Let represent the value of the k-th cluster center on the j-th attribute, and ∑ represent the summation over all static attribute dimensions.

[0083] The system compares all calculated Euclidean distances and identifies the minimum value. The cluster containing the cluster center vector corresponding to this minimum value is determined as the spatial cluster to which the computation node belongs. This deterministic assignment mechanism based on minimum distance is called hard gating.

[0084] Unlike soft-gating mechanisms, which require calculating and weighting the outputs of all experts, hard-gating mechanisms activate only the best-matching expert model during the inference phase. This not only reduces the computational load of the model's inference but also avoids interference between the outputs of different expert models, making it suitable for urban flood warning scenarios with high real-time requirements. For example, in urban areas containing tens of thousands of nodes, this method can control the prediction time for the entire region to within seconds.

[0085] In other words, the Euclidean distance calculation formula can be described as: d _k =sqrt(∑((x _j -C _kj ) 2 ));

[0086] Where, d _kLet be the Euclidean distance between the static physical attribute vector of the node to be predicted and the cluster center vector of the k-th candidate spatial cluster; sqrt is the square root operation; ∑ is the summation operation over all the static physical attribute dimensions involved in the calculation; x _j Let C be the standardized value of the node to be predicted on the j-th static physical property; _kj Let be the value of the cluster center vector of the k-th candidate spatial cluster in the j-th dimension.

[0087] Step 302: Generate the waterlogging prediction results of the computing nodes based on dynamic rainfall characteristics and static physical properties.

[0088] After determining the spatial cluster to which the computing node belongs, the system retrieves the expert model uniquely corresponding to that cluster ID from a pre-built model library. This expert model is a machine learning model, such as the LightGBM model, pre-trained specifically for this type of physical attribute region. The system inputs the node's complete input vector, including 10-dimensional dynamic rainfall features and 9-dimensional static physical attributes, into the expert model. The model outputs the predicted maximum inundation depth for that node under the current rainfall scenario.

[0089] An example provides a specific implementation process for hierarchical advantage attribution and disaster mechanism analysis. This embodiment achieves a leap from data interpretation to mechanism analysis by reorganizing microscopic feature contributions into macroscopic physical system dominance, generating engineering-operable governance recommendations. Specifically, it includes:

[0090] Step 401: Aggregate the feature contribution values ​​according to the physical system affiliation of static physical attributes to obtain the system dominance degree of each physical system, specifically including:

[0091] Obtain the characteristic contribution values ​​of each static physical attribute belonging to the same physical system;

[0092] The system influence amplitude of the physical system is obtained by summing the absolute values ​​of the characteristic contribution values ​​of each static physical attribute.

[0093] The system dominance of a physical system is determined based on the magnitude of its influence.

[0094] Accordingly, for each computing node, the system calculates the feature contribution value of all input features based on the expert model of its cluster using an interpretable algorithm. In this embodiment, the TreeSHAP (Tree-based Interpretive Additive Attribution) algorithm can be selected because it can accurately and quickly calculate the SHAP value of the tree model. Before aggregation, the system can also calculate the average of the feature contribution values ​​of all nodes within the same spatial cluster to obtain the cluster-level feature contribution value before proceeding with subsequent calculations.

[0095] Furthermore, the system performs physical grouping. Based on a predefined physical system attribution table, static features are divided into three groups: terrain system (T), pipeline system (P), and water system (H). Although dynamic rainfall features (R) participate in prediction, they are considered as external common driving forces in the disaster mechanism analysis and do not participate in the calculation of system dominance. The system extracts the feature contribution values ​​within each group.

[0096] Next, the system calculates the system impact amplitude of each physical system. Taking the pipeline network system as an example, the calculation formula is as follows:

[0097] Mag _P =∑(|φ _j |);

[0098] Where, φ _j Let represent the SHAP value of the j-th feature (such as pipe density or pipe distance) within the pipeline network system, |...| represents taking the absolute value, and ∑ represents summing over all features in the group.

[0099] In this study, the sign of the SHAP value only indicates whether the feature increases or decreases the predicted result. However, when analyzing the disaster-causing mechanism, the focus is on the degree and magnitude of the physical system's influence on the inundation result, rather than the direction of influence. Therefore, absolute value summation is used. Similarly, the system calculates the influence amplitude of the terrain system and the influence amplitude of the water system separately.

[0100] Alternatively, the calculation process for the feature contribution value can be as follows:

[0101] For each cluster's expert model, the TreeSHAP algorithm is used for feature attribution calculation. TreeSHAP is a SHAP value calculation method specifically designed for tree models (such as LightGBM, XGBoost, and Random Forest). Compared to the KernelSHAP algorithm (O(2^3)...), TreeSHAP... M TreeSHAP reduces the computational complexity to O(TLD) by leveraging the characteristics of tree structures, which previously had an exponential complexity of 0.5T. 2 ), where T is the number of trees, L is the number of leaf nodes, and D is the depth of the tree, which improves computational efficiency on large-scale samples.

[0102] Where O(...) is the asymptotic representation of algorithm complexity, used to describe the trend of algorithm execution time as the input size increases; 2ᴹ represents the exponential growth relationship (M is the number of feature dimensions), that is, the computational cost of KernelSHAP will increase exponentially with the increase of feature dimension M, resulting in low computational efficiency.

[0103] Specifically, from the test sample set corresponding to scenario s, all samples belonging to cluster c are extracted, denoted as X. _c_s ={x _1 x _2, ..., x _n_c_s}; where n _c_s x is the number of samples in this cluster under this scenario; in other words, x _n_c_s Let x be the number of samples of cluster c in scenario s; further, for each sample x _i ∈X _c_s The TreeSHAP function is called to calculate the SHAP value vector φ(x) of its 19-dimensional features. _i )=[φ _1 (x _i ), φ _2 (x _i ), ..., φ _19 (x _i )], where φ _j (x _i ) represents the j-th feature pair in sample x. _i The marginal contribution of the prediction results, i.e., j=1, 2, ..., 19; the average SHAP value of all samples of cluster c under scenario s is taken to obtain the cluster-level average SHAP vector, which is used for subsequent physical system grouping calculations. The calculation of the cluster-level average SHAP vector can be expressed as:

[0104] φ _bar_j (c, s) = 1 / (n) _c_s )∑ i=1 n_c_s φ _j (x _i );

[0105] Above, φ _bar_j (c, s) represents the cluster-level average SHAP value of the j-th feature in cluster c under scenario s; n _c_s Let ∑ be the number of valid samples contained in cluster c under scenario s; ∑ is the summation operator, here summing the values ​​from the 1st to the nth in cluster c under scenario s. _c_s Summing is performed sequentially on each sample; φ _j (x _i Let x be the i-th sample x in cluster c under scenario s. _i The SHAP value corresponds to the j-th feature; j is the feature dimension index, which in some scenarios can be set to 1, 2, ..., 19, corresponding to the dimension sequence of 19 features; i is the sample index, with a value ranging from 1 to n. _c_s , which corresponds to all sample sequences of cluster c under scenario s.

[0106] Step 402, determining the system dominance of a physical system based on the system influence amplitude, includes: calculating the sum of the system influence amplitudes of all physical systems to obtain the total influence amplitude; calculating the ratio of the system influence amplitude of any physical system to the total influence amplitude to obtain the local dominance ratio of that physical system; and determining the local dominance ratio as the system dominance.

[0107] Specifically, the absolute amplitudes are normalized to enable horizontal comparisons across different regions and rainfall scenarios. Furthermore, the Local Dominance Ratio (LDR) is defined as a specific quantitative indicator of system dominance. The system calculates the sum of the influence amplitudes of the three types of physical systems, i.e., Total. _Mag =Mag _T +Mag _P +Mag _H ;

[0108] In the formula, Mag _T Mag _P Mag _H The magnitude of the impact on the corresponding terrain system, pipeline system, and water system.

[0109] Furthermore, the proportion of each system is calculated. For example, the local advantage of the pipeline system is greater than that of LDR. _P The calculation formula is: LDR _P =Mag _P / Total _Mag The value of this indicator ranges from 0 to 1 and satisfies LDR. _T +LDR _P +LDR _H =1. The larger the LDR value, the stronger the dominant role that physical system plays in the formation of waterlogging in the region.

[0110] Above, LDR _T LDR _P LDR _H These correspond to the local dominance ratios of the terrain system, pipeline system, and water system, respectively.

[0111] Step 403: Based on the comparison of the system dominance of each physical system, determine the disaster-causing mechanism type of the computing node, including:

[0112] Obtain the preset dominant threshold;

[0113] If the local dominance ratio belonging to the pipeline network system is greater than or equal to the dominance threshold, the disaster-causing mechanism type is determined to be dominated by internal structural bottlenecks.

[0114] If the local dominance ratio belonging to the water system is greater than or equal to the dominance threshold, the disaster-causing mechanism type is determined to be dominated by external functional bottlenecks.

[0115] If the local dominance ratio belonging to the topographic system is greater than or equal to the dominance threshold, the disaster-causing mechanism type is determined to be topographic runoff-dominated.

[0116] In other words, if the local dominance ratio of each physical system is less than the dominant threshold, the disaster-causing mechanism is determined to be a composite disaster-causing mechanism.

[0117] In practical applications, the local dominance ratio satisfies LDR. _T +LDR _P +LDR _H The constraint of 1 typically prevents two or more physical systems from simultaneously reaching or exceeding the dominant threshold (e.g., 0.5). In boundary cases where the threshold is set to 0.4 or lower, if multiple systems do indeed simultaneously meet the threshold condition, the system will prioritize identifying the physical system with the largest local dominance ratio as the dominant type. If the difference between the maximum and second-largest values ​​is less than a preset difference threshold (e.g., 0.05), it is determined to be a composite disaster mechanism.

[0118] This step establishes the discrimination logic from numerical indicators to qualitative conclusions. The system presets a dominant threshold, which typically ranges from 0.40 to 0.70, and can be optionally set to 0.50, meaning that the contribution of a certain physical system accounts for more than half of the total contribution, thus having a dominant role.

[0119] On the one hand, the selection of the dominant threshold needs to balance discrimination sensitivity and discrimination reliability. Possible selection criteria are as follows:

[0120] The lower limit can be set to 0.40. When the three types of physical systems are uniformly dominant, their respective local dominance ratios are approximately 1 / 3 ≈ 0.33. Setting the threshold to 0.40 means that the contribution of a certain system must be about 20% higher than that of a uniform distribution to be considered dominant, which can avoid misjudgments caused by slight differences.

[0121] The upper limit can be set to 0.70, which means that the contribution of a certain system exceeds the sum of the other two categories. This is a relatively conservative strong dominance judgment. However, if the threshold is too high, it may cause a large number of regions to be judged as composite dominance, reducing the practical value of the method.

[0122] The default value is 0.50, which means that the contribution of a certain physical system is equal to the sum of the other two types, that is, the contribution of this system accounts for about half.

[0123] In practical applications, it is recommended to perform sensitivity analysis with a step size of 0.05 for the dominant threshold within the range of [0.40, 0.60] to observe the stability of the discrimination results. If the discrimination results change drastically with the dominant threshold, it indicates that the region indeed belongs to the composite dominant type.

[0124] On the other hand, let's take actual data from the Longgang River Basin in Shenzhen as an example.

[0125] For computing nodes located in high-density urban village areas, such as Longdong Community, belonging to cluster 32, the local advantage ratio of its pipeline system to LDR is calculated. _P It is approximately 0.45 to 0.55, with a relatively high contribution from the pipe volumetric density. If LDR _PIf the value exceeds a preset threshold of 0.5, the system determines that the area is dominated by internal structural bottlenecks. This reflects that although the area has a dense pipe network, its underground distribution capacity remains a bottleneck limiting drainage efficiency relative to the high intensity of surface runoff.

[0126] For computational nodes located in low-lying areas along riverbanks, such as the Low Mountain Community, belonging to cluster 27, the calculated local dominance ratio of its water system is LDR. _H The value is relatively high, approximately 0.40 to 0.50, with significant contributions from distance from the river channel and channel area. If LDR _H If the threshold is exceeded (or the largest of the three and significantly larger), it is determined to be dominated by an external functional bottleneck. This reflects that the backwater effect of the receiving water body (river channel) limits the terminal discharge capacity of the stormwater pipe network.

[0127] For computational nodes located at the boundary between mountainous and hilly areas and urban areas, such as a hillside community, the calculated local dominance ratio (LDR) of its terrain system is... _T The value is relatively high, approximately 0.45 to 0.55, with significant contributions from elevation and slope. If LDR _T If the value exceeds a preset threshold of 0.5, the system determines that the area is dominated by topographic runoff. This reflects the surface water catchment conditions of the area, such as steep slopes and large elevation differences, which are the main factors leading to the rapid accumulation of surface runoff and the formation of water.

[0128] Step 404, determining the disaster-causing mechanism type of the computing node, also includes:

[0129] If the local dominance ratio of each physical system is less than the dominant threshold, or the difference between the largest local dominance ratio and the second largest local dominance ratio is less than the preset difference threshold, then the disaster-causing mechanism type is determined to be a composite disaster-causing mechanism.

[0130] For computing nodes identified as having a complex disaster-causing mechanism, the top three feature contribution values ​​with the largest absolute values ​​are selected from the feature contribution values ​​of the computing nodes and output as supplementary explanatory information.

[0131] In practical engineering, the causes of urban flooding in many areas are complex, and there is no single dominant factor. For example, when LDR (Local Drainage Relief) _T =0.35, LDR _P =0.33, LDR _H When the value is 0.32, no single system exceeds the threshold of 0.4 or 0.5, and the differences in contributions among the systems are minimal. In this case, the system is determined to have a composite disaster-causing mechanism.

[0132] Building upon this, the system delves further into the micro-level characteristics, directly outputting the three features with the largest absolute SHAP values. For example, the output might show: the first influencing factor is elevation (topography), the second is pipeline density (pipeline network), and the third is distance from the river (water system). This suggests that the synergistic effects of low-lying terrain improvement, pipeline capacity expansion, and river regulation need to be comprehensively considered in this area.

[0133] It should be understood that this step is a fallback judgment logic for complex situations.

[0134] On the other hand, the determination rule for composite dominant types can also be set as follows:

[0135] Cluster c is determined to be a composite dominant or hybrid mechanism type in scenario s when any of the following conditions are met:

[0136] The local dominance ratio of all physical systems is less than the dominance threshold, i.e., it satisfies the following relationship: max(LDR) _P LDR _T LDR _H At this point, no physical system's dominance exceeds the threshold.

[0137] Let the local dominance ratios of the three types of systems be arranged in descending order as LDR. _1 ≥LDR _2 ≥LDR _3 If the following relationship is satisfied: LDR _1 -LDR _2 If the value is less than 0.10, the dominant difference is considered to be insignificant.

[0138] In the above, max(...) is the function to find the maximum value, and θ corresponds to the dominant threshold.

[0139] For nodes identified as having a composite dominant type, the system outputs the top three contributing features as supplementary explanatory information, specifically including the absolute value of the cluster-level average SHAP |φ. _bar_j Select the three largest features from (c, s)| and output the feature names and their SHAP values ​​in descending order of contribution.

[0140] Step 405, after determining the disaster-causing mechanism type of the computing node, also includes:

[0141] If the problem is determined to be dominated by internal structural bottlenecks, engineering optimization suggestions will be generated that prioritize gray infrastructure.

[0142] If the problem is determined to be primarily an external functional bottleneck, engineering optimization recommendations will be generated that prioritize blue infrastructure.

[0143] If topographic runoff is determined to be dominant, engineering optimization suggestions that prioritize green infrastructure will be generated.

[0144] Alternatively, if the disaster is determined to be a complex disaster-causing mechanism, engineering optimization suggestions can be generated to comprehensively optimize blue infrastructure, green infrastructure, and gray infrastructure.

[0145] Accordingly, the system automatically matches a pre-set library of engineering measures based on the judgment results. Specifically, if the problem is determined to be dominated by internal structural bottlenecks (limited by the pipe network), the system suggests gray infrastructure measures. For example, it recommends widening and upgrading the drainage pipes in the area, adding rainwater storage tanks to reduce flood peaks, or optimizing the scheduling and operation strategies of pumping stations to solve the problem of insufficient underground distribution capacity.

[0146] If the problem is determined to be primarily due to external functional bottlenecks (water system backwater), the system recommends implementing blue infrastructure measures. For example, it suggests widening and dredging river channels to lower water levels, constructing riverside wetlands to increase flood storage capacity, installing floodgates to prevent backflow, and improving the boundary conditions of the receiving water body.

[0147] If the system determines that topographic runoff is dominant (terrain runoff), it recommends implementing green infrastructure measures, i.e., sponge city measures. For example, it suggests constructing sunken green spaces, permeable paving, rain gardens, or green roofs to increase surface infiltration, reduce runoff generation at the source, or guide runoff direction by combining micro-topographical modifications.

[0148] For areas identified as having a complex disaster-causing mechanism, the system recommends comparing multiple options and comprehensively evaluating the synergistic benefits of blue, green, and gray measures.

[0149] Furthermore, the method of this invention also includes scenario sensitivity analysis. The system executes the aforementioned prediction and attribution processes for multiple different rainfall scenarios, such as short-duration sudden heavy rainfall scenarios and long-duration continuous rainfall scenarios. By comparing the changes in the disaster-causing mechanism type of the same computing node under different scenarios, the system can identify scenario-sensitive areas. For example, a certain area may exhibit pipeline-dominated behavior under short-duration rainfall conditions (i.e., rainfall intensity exceeding pipeline design standards), while under long-duration rainfall conditions, it may shift to river system-dominated behavior (i.e., downstream river levels rising and causing backwater). For such areas, engineering recommendations will emphasize simultaneously meeting the dual standards of high-intensity rapid drainage and high-water-level backwater prevention.

[0150] Optionally, a formula for the average contribution value of cluster-level features is provided, specifically:

[0151] φ _bar_j (c, s) = (1 / n) _c_s )×∑(φ _j (x _i ));

[0152] Where, φ _bar_j(c, s) represents the cluster-level average contribution of the c-th spatial cluster to the j-th feature under rainfall scenario s; n _c_s ∑ represents the total number of computational nodes contained in the spatial cluster c; ∑ represents the summation over all nodes in the cluster; φ _j (x _i ) represents the SHAP contribution value of the i-th computation node within the spatial cluster with respect to the j-th feature.

[0153] Optionally, the formula for calculating the system impact amplitude, taking a pipeline system as an example, can be:

[0154] Mag _P (c, s) = ∑ j∈P |φ _bar_j (c, s)|;

[0155] Among them, Mag _P (c, s) represents the system impact amplitude of the pipeline system in the c-th spatial cluster under rainfall scenario s; ∑ represents the summation of all features belonging to the pipeline system; φ _bar_j (c, s) represents the cluster-level average contribution value of the j-th feature belonging to the pipeline network system.

[0156] Furthermore, Mag _T and Mag _H The calculation method is similar, only the summation range needs to be replaced with features belonging to the terrain system or water system.

[0157] Alternatively, the Local Dominance Ratio (LDR), taking a pipeline network system as an example, can be calculated using the following formula:

[0158] LDR _P =Mag _P / (Mag _T +Mag _P +Mag _H );

[0159] Among them, LDR _P The local dominance ratio of the pipeline network system, i.e., the system dominance of the pipeline network system; Mag _P The system impact amplitude of the pipeline network system; Mag _T The systemic impact amplitude of the terrain system; Mag _H This represents the systemic impact amplitude of the water system.

[0160] This application constructs a Hierarchical Dominance Attribution Framework (HDAF) to address the issues of lack of physical interpretability and difficulty in decoupling disaster-causing mechanisms. This method reorganizes and aggregates the characteristic contribution values ​​(SHAP) at the micro-level according to physical meaning, calculating the system dominance and normalized local advantage ratio (LDR) of each physical system. This enables the black-box model to output discriminant indicators with clear physical meaning, distinguishing between internal structural bottlenecks caused by insufficient pipeline capacity and external functional bottlenecks caused by river backwater.

[0161] Furthermore, by setting a dominant threshold and discrimination logic, the embodiment establishes an automatic mapping relationship from disaster-causing mechanism type to blue-green-grey engineering facilities, and outputs targeted engineering optimization suggestions.

[0162] Another example describes an exemplary scheme for evaluating the Macro Regional Dominance Index (RDI), illustrating how to assess the dominance of each physical system from a global macro perspective. This embodiment addresses the problem that simple arithmetic averaging under non-uniform grid conditions can lead to evaluation results biased towards densely populated areas (such as the core urban area in a high-precision simulation) while neglecting sparsely populated areas (such as suburbs), thus ensuring the physical impartiality of the macro indicator.

[0163] Accordingly, the control volume area corresponding to each computing node is obtained; the system dominance of the computing nodes is calculated by weighted average based on the control volume area of ​​the computing nodes, and the regional dominance index reflecting the dominance of the physical system in the whole domain is obtained.

[0164] This step introduces a spatial weighting mechanism to calculate the Regional Dominance Index (RDI). In urban flooding simulations, to balance computational accuracy and efficiency, unstructured grids are typically used to discretize the computational domain. This results in significant differences in grid density across different regions. For example, in city centers or areas with dense pipe networks, the grid area may be only a few square meters, while in suburbs or mountainous areas, the grid area may reach hundreds of square meters. Directly calculating the arithmetic mean of the system dominance of all nodes would cause the overall index to be severely biased towards areas with dense nodes, failing to accurately reflect the spatial attributes of the entire domain.

[0165] Therefore, this embodiment introduces the concept of control volume area. For the model discretized using an unstructured triangular mesh, the control volume area A of each computation node i is... _i It can be defined as the area of ​​the Voronoi polygon (Thieson polygon) corresponding to the node, or approximately one-third of the sum of the areas of all the triangular units to which the node belongs. This type of area data can usually be calculated and stored as a static attribute during the physical model construction phase.

[0166] Based on this, the system performs area-weighted calculation. Let there be M calculation nodes in the entire region, and let A be the control volume area of ​​the i-th node._i The system dominance (i.e., local dominance ratio, LDR) of the physical system to which this node belongs is LDR. _i, Taking pipeline systems as an example, their Regional Advantage Index (RDI) _P The calculation formula is:

[0167] RDI _P =∑(LDR _P_i ×A _i ) / ∑(A _i );

[0168] Among them, the numerator ∑(LDR) _P_i ×A _i LDR represents the total contribution of the pipeline system considering spatial weights. _P_i The corresponding local advantage ratio of the pipeline network system, the denominator ∑(A _i The formula represents the total area of ​​the entire region. It assigns equal weight to the contribution of each square meter of physical mechanism to the overall regional index, rather than assigning equal weight to each computational node.

[0169] By calculating the RDI index, the evolution of disaster-causing mechanisms at a global scale can be quantitatively analyzed. For example, the system can calculate the RDI values ​​under different rainfall scenarios s. Under a short-duration heavy rainfall scenario, if the calculated RDI... _P A higher RDI, such as 0.4, indicates a strong constraint effect of the pipeline network system across the entire region; however, under long-duration rainfall scenarios, if RDI is found to be lower... _P A significant decrease, for example, to 0.2, while RDI _T (Topography) has risen considerably.

[0170] In other words, as rainfall continues, the pipe network generally reaches full flow, its drainage capacity is submerged or saturated, and the controlling factor of waterlogging shifts to the natural accumulation effect of topography. Furthermore, this evolutionary analysis can be used to formulate basin-level flood control and drainage plans.

[0171] As an alternative implementation, weighted calculation can also be performed at the cluster level. Specifically, the total area A(c) of each spatial cluster c is calculated, which is the sum of the control volume areas of all nodes within that cluster. Weighting is performed using the cluster-level local dominance ratio LDR(c):

[0172] RDI=∑(LDR(c)×A(c)) / ∑(A(c));

[0173] In the formula, LDR(c) is the cluster-level local dominance ratio corresponding to spatial cluster c. This method is mathematically equivalent, but requires less computation and is more convenient for independent evaluation of different functional areas (such as commercial area clusters and residential area clusters).

[0174] Alternatively, the weighted calculation formula for the Regional Advantage Index (RDI), taking the pipeline network system as an example, can also be expressed as the following formula:

[0175] RDI _P (s)=∑ c (LDR _P (c, s) × A(c)) / ∑ c (A(c));

[0176] Among them, RDI _P (s) represents the regional dominance index of the pipeline network system across the entire region under rainfall scenario s; ∑ represents the summation over all spatial clusters c across the entire region; LDR _P (c, s) represents the local dominance ratio of the pipeline system of the c-th spatial cluster under rainfall scenario s; A(c) represents the control volume area of ​​the c-th spatial cluster, which is the sum of the control volume areas of all computing nodes in the cluster.

[0177] Another example provides an optional technical solution for an electronic device and system architecture, particularly a hardware entity structure for executing the above-described method. In this embodiment, the algorithm logic is solidified into a computer system, giving it a physical form suitable for practical application. Accordingly, it specifically includes:

[0178] Step 601, an electronic device, characterized in that it comprises: a memory for storing a computer program; and a processor for executing the computer program to implement the method as described in any one of the present applications.

[0179] Specifically, the electronic device can be a high-performance workstation, a server cluster, or a cloud-based virtual computing node. In terms of hardware architecture, the device mainly includes a processor, memory, communication interfaces, and buses connecting the various components.

[0180] The processor can specifically be a central processing unit (CPU), graphics processing unit (GPU), tensor processor (TPU), or field-programmable gate array (FPGA), etc. In this invention, the processor is configured to execute a series of complex arithmetic instructions, including but not limited to:

[0181] Performing Z-score normalization on massive static physical properties involves floating-point addition, subtraction, multiplication, and division.

[0182] Calculating Euclidean distance in high-dimensional space to implement hard-gated assignment involves vector operations.

[0183] Multiple LightGBM expert models are invoked in parallel for inference, involving tree structure traversal and leaf node summation;

[0184] Feature attribution calculation based on the TreeSHAP algorithm involves path integral and marginal contribution accumulation.

[0185] To meet the requirements of second-level real-time prediction, a multi-core CPU that supports multi-threaded parallel computing or a GPU with large-scale parallel computing capabilities can be selected.

[0186] Memory includes high-speed random access memory (RAM) and non-volatile memory (ROM, hard disk, or SSD). Memory stores not only the computer program instructions used to execute the above methods, but also a large amount of key data and parameters supporting the algorithm's operation. Specifically, this includes:

[0187] The static physical attribute database stores terrain, pipeline, and water system attribute data for all computing nodes, as well as control volume area data used to calculate RDI.

[0188] Spatial partitioning parameter library, storing cluster center vectors (C) generated by K-means clustering. _1 To C _K K represents the mean and standard deviation used for data standardization.

[0189] The expert model library stores LightGBM model files trained for each spatial cluster. These files can be in .txt or .bin format and include the tree structure, split threshold, and leaf node weights.

[0190] The threshold configuration table stores the threshold number of nodes (e.g., 1000), the dominant threshold (e.g., 0.50), and the difference threshold used for disaster hazard mechanism identification.

[0191] The communication interface is used for data exchange with external systems. For example, it can receive gridded numerical weather forecast data from meteorological departments in real time via a network interface, corresponding to dynamic rainfall characteristics, or connect to a display terminal to graphically display the predicted inundation depth distribution map and the generated engineering optimization suggestion report.

[0192] In actual operation, after the electronic device is powered on, the processor loads the operating system and the computer program of this invention from non-volatile memory into memory. Furthermore, the processor is in a standby state; once it receives a trigger signal, such as a new rainfall forecast, through the communication interface, it sequentially executes steps including data acquisition, standardization, cluster assignment, expert reasoning, attribution analysis, mechanism determination, and macro-level evaluation, outputting the final result to the user interface or downstream decision-making system. This hardware-software collaborative architecture ensures that the method of this invention can serve the actual operational needs of urban flood control and drainage.

[0193] In other embodiments, a rapid prediction and disaster mechanism analysis method for urban flooding based on physical-guided hybrid expert clustering and hierarchical attribution (HMoE-HDAF) is provided. Taking the maximum inundation depth of urban computing nodes / grids as the prediction object, the following functional relationship is established:

[0194] d = f(R, T, P, H);

[0195] Where R represents dynamic rainfall characteristics, T represents topographic attributes, P represents drainage network attributes, and H represents water system attributes.

[0196] Optionally, a high-fidelity physical model and ground truth dataset are constructed. High-confidence ground truth is generated from the calibrated and verified mechanistic model to address the problem of difficulty in training surrogate models due to the scarcity of grid-level flood observations.

[0197] Furthermore, a coupled hydrological-hydrodynamic physical model platform is constructed, such as InfoWorksICM or equivalent software. The model should include at least the following modules: rainfall-runoff module, gate / pump station dynamic scheduling module, and hydrodynamic inundation module, to simultaneously characterize the coupled impact of watershed rainfall, drainage network load, and river flow capacity on urban flooding.

[0198] Furthermore, the physical model was calibrated and validated. Observational data, such as measured river water levels, were used to calibrate the runoff coefficient, roughness coefficient, and underlying surface parameters. The reliability of the model was verified using indicators such as NSE.

[0199] Furthermore, we designed and ran multi-scenario rainfall to generate a ground truth dataset: the rainfall duration range needs to cover combinations of short and long duration heavy rainfall; hydraulic boundary conditions such as constant water level can be set downstream, and the maximum inundation depth of each computing node / grid is output as a supervised learning label.

[0200] Optionally, multi-source spatiotemporal features are constructed. Meteorological drivers and static physical constraints are explicitly encoded as computable features, ensuring that the model input is interpretable and reproducible.

[0201] Furthermore, an input vector is constructed to form a feature-target sample. The input can optionally contain 19 variables, including 10 dynamic rainfall features R and 9 static features (grouped as T, P, H). The dynamic rainfall features R include at least: rainfall duration, total rainfall, maximum hourly cumulative rainfall, pre-peak cumulative rainfall, and rainfall statistics, which are used to characterize rainfall intensity and temporal distribution.

[0202] Furthermore, static environment features extracted by grouping them according to physical systems include:

[0203] Topography (T), elevation, slope, impermeability coefficient, etc.;

[0204] Pipeline P includes distance from designated inspection well, pipe length density per unit area, and pipe volume density per unit area.

[0205] Water system H, distance from the nearest river, water area per unit area of ​​the river channel, area per unit area of ​​the reservoir, etc.

[0206] To reduce the impact of multicollinearity and retain representative indicators, a geographic detector (GD) and variance inflation factor (VIF) can be used for feature screening and multicollinearity testing.

[0207] Optionally, computational domain partitioning based on physical properties (CMoE clustering, where CmoE stands for cluster-based expert hybrid model) can be used. To address the spatial heterogeneity and non-stationarity of urban flooding, a spatial divide-and-conquer approach is adopted, combining static physical property-based clustering partitioning with hard-gated hybrid experts. Homogeneous regions are formed using the static physical similarity of topography, pipe networks, and water systems, allowing different experts to learn the local nonlinear response in different areas.

[0208] Accordingly, dry points with a maximum inundation depth less than a threshold are removed. The threshold can be selected as 0.01m (the parameter can be adjusted within the range of 0.005-0.05m) to reduce invalid samples and improve computational efficiency. The static attributes T, P, and H used for clustering are standardized. K-means or an equivalent clustering algorithm is used to cluster the static attributes. The initial number of clusters K can be determined by parameter tuning within the range of 30-80; in one embodiment, K=50 is used.

[0209] Clusters with fewer than a threshold number of nodes are merged. The threshold can be set to 1000 nodes, with a range of 500-2000 nodes. The merging rule is to merge into the nearest larger cluster based on the Euclidean distance between the cluster centers. This yields a spatial partition set and a cluster assignment function, which serve as the deterministic gating basis for subsequent expert models.

[0210] Optionally, a local expert model can be constructed for rapid prediction. By using partitioned learning and hard gating, computation time can be reduced while maintaining accuracy, achieving predictions within seconds.

[0211] Train a separate data-driven expert model E for each spatial cluster c. _c (...). The expert model takes dynamic rainfall features R and static features T, P, H as inputs and outputs the maximum flooding depth.

[0212] Gated inference employs a hard-gating method: for the sample to be predicted, the cluster c(x) to which it belongs is first determined by the cluster assignment function, and then the corresponding expert is called to output the predicted value y(x)=E. _c(x) (x); where c(x) is the spatial partitioning cluster assignment, i.e., the cluster assignment function.

[0213] Based on this, a maximum inundation depth distribution map of all nodes / grids is generated, which can be further thresholded to obtain the inundation range and risk zoning results.

[0214] In addition, in some scenarios, for 72-hour rainfall events, the CMoE-LightGBM surrogate model completes the prediction of the maximum flooding depth across the entire domain, which significantly improves the simulation speed compared with the mechanistic model, accelerates inference, and meets the needs of real-time multi-scenario forecasting.

[0215] Optionally, Hierarchical Dominance Attribution Analysis (HDAF) is performed. This transforms the surrogate model output into an interpretable physical system disaster mechanism. Based on CMoE predictions, the HDAF framework is introduced to reorganize feature-level attributions into physical system-level dominant indicators. Feature importance is elevated to physical system dominance, achieving interpretable quantitative output from prediction to disaster mechanism.

[0216] At the micro level, for each cluster expert model E _c For rainfall scenario s, the marginal contribution value of each input feature is calculated using SHAP (or equivalent attribution method) to obtain the feature-contribution set, which is used to explain the local driving factors under specific scenarios.

[0217] In the mesoscale layer, features are grouped according to physical meaning into topographic system (T), pipeline system (P), and water system (H). The system influence amplitude (Mag) is obtained by summing the contributions (absolute values ​​can be selected) of features within each group. _T (c, s), Mag _P (c, s), Mag _H (c, s). Based on this, the local dominance ratio is defined. Taking a pipe network as an example, it can be specifically described as:

[0218] Mag _P (c, s) = ∑ v_i∈P |SHAP(c, s, v) _i )|;

[0219] LDR _P (c, s) = (Mag _P (c, s)) / (Mag _P (c, s) + Mag _H (c, s) + Mag _T (c, s));

[0220] In the formula, Mag _P (c, s) represents the system impact amplitude of the pipeline system P in spatial cluster c under scenario s; ∑ v_i∈P The summation operator represents summing all characteristics v of the pipeline system P. _i Summing sequentially; SHAP(c, s, v) _i Let v be the i-th feature of the pipeline system in spatial cluster c under scenario s. _i The SHAP value, where |…| represents the absolute value sign, is used to ignore the direction of feature influence and only quantify the intensity of influence; v _i Let i be a characteristic individual of the pipeline system P, where i is the feature index, c is the spatial cluster identifier, and s is the scenario identifier.

[0221] LDR _P(c, s) represents the local dominance ratio of the pipe network system P in spatial cluster c under scenario s; the denominator is the sum of the influence amplitudes of the terrain system, pipe network system, and water system, used for normalization to ensure LDR. _P The value of (c, s) ranges from 0 to 1, reflecting the relative dominance of the pipeline network system in the disaster-causing outcome under this cluster and scenario; Mag _T (c, s), Mag _H (c, s) represent the system influence amplitudes of the terrain system and the water system in spatial cluster c under scenario s, respectively. The calculation logic is consistent with Mag. _P (c, s) are consistent; LDR _T (c, s), LDR _H (c, s) represent the local dominance ratios of topographic system T and hydrological system H, respectively, and are defined logically and in relation to LDR. _P (c, s) are consistent.

[0222] Furthermore, cluster-scale disaster type discrimination is performed to identify bottleneck types. Within cluster c, {LDR} is compared. _P (c, s), LDR _T (c, s), LDR _H The value of (c, s)} can be optionally set with a dominance threshold θ, which ranges from [0.40, 0.70], for example, a value of 0.50. In other words, the dominance threshold corresponds to the dominant threshold.

[0223] If LDR _P (c, s) is the largest of the three and LDR _P If (c, s) ≥ θ, then the cluster is determined to be dominated by internal structural bottlenecks, with drainage network transportation or storage capacity constraints being the main factor.

[0224] If LDR _H (c, s) is the largest of the three and LDR _H If (c, s) ≥ θ, then the cluster is determined to be dominated by external functional bottlenecks and mainly constrained by the boundary conditions / top-support effect of the receiving water body.

[0225] If LDR _T (c, s) is the largest of the three and LDR _T If (c, s) ≥ θ, then the cluster is determined to be dominated by topographic runoff;

[0226] If none of the three reach the threshold or the differences are small, it is determined to be a composite dominant / mixed mechanism, and the top three contributing features are output to assist in the explanation.

[0227] At the macro level, a regional advantage index is obtained by weighting the entire region by area. Taking the pipeline network as an example, it can be expressed by the following formula:

[0228] RDI _P (s)=(∑c=1 N LDR _P (c, s) × A(c)) / (∑ c=1 N A(c));

[0229] Where A(c) is the spatial area of ​​cluster c, used to avoid small clusters having a disproportionate impact on the overall evaluation.

[0230] Or, in other words, RDI _P (s) represents the regional dominance index of the pipeline system P under scenario s, used to quantify the overall dominant contribution of the pipeline system to the disaster-causing process across the entire spatial area under this scenario; ∑ c=1 N LDR is a summation operator, representing the summation of spatial clusters from the 1st to the Nth in the entire domain; _P (c, s) represents the local dominance ratio of the pipeline system P in the c-th spatial cluster under scenario s, reflecting the relative dominance of the pipeline system within the cluster; A(c) represents the total area of ​​the c-th spatial cluster, which is the sum of the control volume areas of all nodes within the cluster, serving as a weighting factor in the weighted calculation, reflecting the influence of the area proportion of different spatial clusters on the overall dominance; N represents the total number of spatial clusters in the entire domain, c represents the identification index of the spatial cluster, with a value range from 1 to N; s represents the scenario identifier, representing specific disaster-causing scenarios, such as rainfall intensity, duration, and other operating conditions.

[0231] Similarly, RDI can be obtained _T (s), RDI _H (s). Used to characterize the relative dominance of physical systems such as pipe networks, topography, and water systems in the overall domain under scenario s and their evolution with rainfall patterns. Not used to distinguish the internal and external bottleneck types of individual spatial units.

[0232] The following is a numerical example to illustrate the specific calculation process of the local dominance ratio.

[0233] Assuming a computing node in the Longdong community belongs to cluster 32, the feature contribution values ​​calculated by the TreeSHAP algorithm are as follows:

[0234] In the topographic system characteristics, the contribution value of elevation is 0.12, the contribution value of slope is -0.08, and the contribution value of impermeability coefficient is 0.05; in the pipeline system characteristics, the contribution value of distance from manhole is -0.15, the contribution value of pipe length density per unit area is -0.22, and the contribution value of pipe volume density per unit area is -0.28; in the water system characteristics, the contribution value of distance from river is 0.06, the contribution value of river area per unit area is -0.04, and the contribution value of reservoir area per unit area is 0.02.

[0235] Input values, calculate the impact amplitude of each system, i.e.: Mag_T =|0.12|+|-0.08|+|0.05|=0.25;Mag _P =|-0.15|+|-0.22|+|-0.28|=0.65; Mag _H =|0.06|+|-0.04|+|0.02|=0.12;

[0236] Given the input values, calculate the local dominance ratio, i.e., Total. _Mag =0.25+0.65+0.12=1.02; LDR _T =0.25 / 1.02≈0.245; LDR _P =0.65 / 1.02≈0.637; LDR _H =0.12 / 1.02≈0.118.

[0237] Due to LDR _P =0.637>0.50 (dominant threshold), therefore the disaster-causing mechanism of this computing node is determined to be dominated by internal structural bottlenecks. Further analysis shows that the absolute value of the contribution from pipe volume density is the largest, indicating that insufficient pipe volume in this area is a key factor restricting drainage capacity.

[0238] The optional embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details of the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solution of the present invention, and such equivalent transformations all fall within the protection scope of the present invention.

Claims

1. A method for rapid prediction and disaster-causing mechanism analysis of urban flooding, characterized in that, include: The dynamic rainfall characteristics of the target area and the static physical properties distributed across multiple computing nodes within the target area are obtained. The static physical properties belong to multiple physical systems. The spatial cluster to which the computing node belongs is determined based on the static physical properties of the computing node. The expert model corresponding to the spatial cluster is called, and the waterlogging prediction result of the computing node is generated based on the dynamic rainfall characteristics and static physical properties. Based on the expert model, the characteristic contribution values ​​of dynamic rainfall characteristics and static physical attributes to the urban flooding prediction results are calculated. The characteristic contribution values ​​are aggregated according to the physical system affiliation of the static physical attributes to obtain the system dominance degree of each physical system. Based on the comparison of the system dominance of each physical system, the disaster-causing mechanism type of the computing node is determined; Specifically, the feature contribution values ​​are aggregated according to the physical system affiliation of static physical attributes to obtain the system dominance of each physical system. This includes: obtaining the feature contribution values ​​of each static physical attribute belonging to the same physical system; taking the absolute value of the feature contribution values ​​of each static physical attribute and summing them to obtain the system influence amplitude of the physical system; and determining the system dominance of the physical system based on the system influence amplitude. The determination of the system dominance of a physical system based on the system influence amplitude includes: calculating the sum of the system influence amplitudes of all physical systems to obtain the total influence amplitude; calculating the ratio of the system influence amplitude of any physical system to the total influence amplitude to obtain the local dominance ratio of that physical system; and determining the local dominance ratio as the system dominance. The construction of the expert model includes: performing clustering based on the static physical properties of the sample data to generate multiple initial clusters; counting the number of nodes contained in each initial cluster; if there is an initial cluster with fewer nodes than the node number threshold, calculating the Euclidean distance between the cluster center of the initial cluster and the cluster centers of the other initial clusters; merging the initial cluster into the target cluster with the minimum Euclidean distance, until the number of nodes in all clusters is not less than the node number threshold, thus obtaining spatial clusters, and training the corresponding expert model for each spatial cluster.

2. The method according to claim 1, characterized in that, The physical system includes the terrain system, the pipeline system, and the water system; Among them, the static physical properties belonging to the terrain system include elevation, slope and impermeability coefficient; Static physical properties belonging to the pipeline network system include distance from manholes, pipe length density per unit area, and pipe volume density per unit area. The static physical properties belonging to a water system include the distance to the nearest river, the water area per unit area of ​​the river channel, and the reservoir area per unit area.

3. The method according to claim 1, characterized in that, The spatial cluster to which a computing node belongs is determined based on its static physical properties, including: Obtain the pre-stored cluster center vectors corresponding to multiple candidate spatial clusters; Solve for the Euclidean distance between the static physical properties of the computation nodes and the cluster center vectors; The candidate spatial clusters corresponding to the cluster center vectors that have the smallest Euclidean distance to the static physical properties are determined as the spatial clusters to which the computing nodes belong.

4. The method according to claim 1, characterized in that, Based on the comparison of the system dominance of each physical system, the disaster-causing mechanism type of the computing node is determined, including: Obtain the preset dominant threshold; If the local dominance ratio belonging to the pipeline network system is greater than or equal to the dominance threshold, the disaster-causing mechanism type is determined to be dominated by internal structural bottlenecks. If the local dominance ratio belonging to the water system is greater than or equal to the dominance threshold, the disaster-causing mechanism type is determined to be dominated by external functional bottlenecks. If the local dominance ratio belonging to the topographic system is greater than or equal to the dominance threshold, the disaster-causing mechanism type is determined to be topographic runoff-dominated.

5. The method according to claim 4, characterized in that, After determining the disaster-causing mechanism type of the computing node, the following is also included: If the problem is determined to be dominated by internal structural bottlenecks, engineering optimization suggestions will be generated that prioritize gray infrastructure. If the problem is determined to be primarily an external functional bottleneck, engineering optimization recommendations will be generated that prioritize blue infrastructure. If topographic runoff is determined to be dominant, engineering optimization suggestions that prioritize green infrastructure will be generated.

6. The method according to claim 1, characterized in that, Also includes: Obtain the control volume area corresponding to each computing node; The system dominance of computing nodes is calculated by weighted average of the control volume area of ​​computing nodes, resulting in a regional dominance index that reflects the degree of dominance of the physical system across the entire domain.

7. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing a computer program to implement the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Urban inland inundation toughness evaluation method and related equipment

    CN120598375A

  • Urban inland inundation simulation prediction method and system based on interpretable machine learning

    CN121168753A