A flood prediction method and apparatus

CN122525696APending Publication Date: 2026-08-07四川省水文水资源勘测中心(四川省量水设施设备计量检测中心)
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
四川省水文水资源勘测中心(四川省量水设施设备计量检测中心)
Filing Date
2026-07-13
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0002]传统洪水预测系统主要依赖统计模型或深度学习模型,这些方法通常仅从单测站的时间序列数据出发进行分析,使得模型无法全面反映影响洪水形成的复杂环境因素,洪水预测结果准确性较差

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122525696A_ABST
    Figure CN122525696A_ABST
Patent Text Reader

Abstract

The application relates to a flood prediction method and device, a Bayesian time sequence decomposition model fusing a river network topological structure is constructed, topological smoothing constraints are applied to a spatial weight matrix, and multi-source climate driving data are integrated, accurate modeling of a flood space-time evolution process and analysis of propagation characteristics are realized, the river network topological structure and external climate driving data can be effectively integrated, cross-basin flood prediction accuracy is significantly improved, and reference data such as uncertainty quantification and physical consistency verification of a prediction result can be provided for user reference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of hydrological prediction technology, and more specifically, to a flood prediction method and apparatus. Background Technology

[0002] Traditional flood forecasting systems mainly rely on statistical models or deep learning models. These methods typically analyze only time series data from a single station, making it impossible for the models to fully reflect the complex environmental factors that influence flood formation, resulting in poor accuracy of flood forecasting results.

[0003] To address the aforementioned issues, existing technologies urgently need improvement. Summary of the Invention

[0004] This application provides a flood prediction method and apparatus that can effectively integrate river network topology and external climate-driven data, significantly improving the accuracy of flood prediction.

[0005] To achieve the above objectives, according to a first aspect of this application, a flood forecasting method is provided, comprising: Spatiotemporal input data are generated based on rainfall and flow data from multiple monitoring stations, as well as external climate-driven data. An adjacency matrix is ​​constructed based on the river network topology among the multiple monitoring stations, and a symmetric graph Laplace matrix is ​​generated. A Bayesian temporal decomposition model is constructed and trained based on the spatiotemporal input data and the symmetric graph Laplace matrix. The Bayesian temporal decomposition model includes a trend term, a seasonal term, a covariate weight matrix for fusing the external climate-driven data, a spatial weight matrix for capturing the spatial correlation between the monitoring stations, and a noise term. The symmetric graph Laplace matrix is ​​used to simultaneously apply topological smoothing constraints to the spatial weight matrix to force the spatial weights of adjacent monitoring stations with direct hydraulic connectivity in the river network to remain smooth. The external climate-driven data for the time to be predicted is input into the trained Bayesian time series decomposition model, which outputs the prediction results of the flow data and the corresponding uncertainty interval. Based on the flow sequences of upstream and downstream monitoring stations, the time lag correlation coefficient is determined, hydrological propagation time and data anomalies are identified, and the lag propagation analysis results are output. The similarity between the weight vectors of adjacent monitoring stations in the spatial weight matrix after training and the time lag correlation coefficient are verified to confirm each other, so as to verify the physical consistency of the prediction results of the traffic data.

[0006] According to a second aspect of this application, a flood forecasting device is provided, comprising: The data module is used to generate spatiotemporal input data based on rainfall and flow data from multiple monitoring stations, as well as external climate-driven data. The matrix module is used to construct an adjacency matrix based on the river network topology among the multiple monitoring stations and generate a symmetric graph Laplace matrix. The training module is used to construct and train a Bayesian temporal decomposition model based on the spatiotemporal input data and the symmetric graph Laplace matrix. The Bayesian temporal decomposition model includes a trend term, a seasonal term, a covariate weight matrix for fusing the external climate-driven data, a spatial weight matrix for capturing the spatial correlation between the monitoring stations, and a noise term. The symmetric graph Laplace matrix is ​​used to simultaneously apply topological smoothing constraints to the spatial weight matrix to force the spatial weights of adjacent monitoring stations with direct hydraulic connectivity in the river network to remain smooth. The output module is used to input the external climate-driven data of the time to be predicted into the trained Bayesian time series decomposition model and output the prediction results of the flow data and the corresponding uncertainty interval. The analysis module is used to determine the time lag correlation coefficient based on the flow sequence of upstream and downstream monitoring stations, identify hydrological propagation time and data anomalies, and output the lag propagation analysis results. The verification module is used to verify whether the similarity of the weight vectors of adjacent monitoring stations in the trained spatial weight matrix and the time lag correlation coefficient corroborate each other, so as to verify the physical consistency of the prediction results of the traffic data.

[0007] This application provides a flood prediction method and apparatus. By constructing a Bayesian temporal decomposition model that integrates river network topology, applying topological smoothing constraints to the spatial weight matrix, and integrating multi-source climate driving data, it achieves accurate modeling and propagation characteristic analysis of the spatiotemporal evolution of floods. It can effectively integrate river network topology and external climate driving data, significantly improve the accuracy of cross-basin flood prediction, and also provide reference data such as uncertainty quantification and physical consistency verification of prediction results for users to refer to.

[0008] Other features and advantages of this application will be described in detail in the following detailed description section. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] To gain a more complete understanding of this application and its beneficial effects, the following description will be provided in conjunction with the accompanying drawings, wherein the same reference numerals in the following description denote the same parts.

[0011] Figure 1 This is a schematic flowchart of a flood prediction method provided in an embodiment of this application; Figure 2 This is a schematic diagram of the input and output of a Bayesian temporal decomposition model provided in an embodiment of this application; Figure 3 This is a schematic diagram of a flood prediction device provided in an embodiment of this application. Detailed Implementation

[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the protection scope of this application.

[0013] For example, in hydrological monitoring practices in major river basins of Sichuan Province, rainfall and flow data from multiple monitoring stations are used for flood prediction. When processing this data using traditional methods, the hydraulic connectivity between the Xzi and Xjiang stations is not encoded into the prediction process because the model does not consider the river network topology between the stations. This results in the prediction of downstream flow failing to accurately reflect the propagation timing of upstream floods. Simultaneously, external climate-driven data such as atmospheric circulation and sea surface temperature anomalies are not included in the input, leading to insufficient model response to extreme rainfall events triggered by climate anomalies. Furthermore, in the basin boundary areas, the lack of physical constraints on spatial weights causes the model's prediction of trans-basin flood propagation paths to deviate from the actual hydrological processes, resulting in distorted decision-making.

[0014] If the aforementioned problems are not addressed, flood forecasting systems will struggle to provide predictions that conform to hydrophysical laws, hindering decision-makers' ability to accurately assess flood risks. In extreme rainfall events, forecast biases can lead to delays or distortions in early warning information, affecting the timely deployment of emergency response measures. Furthermore, the lack of quantification of forecast uncertainty renders risk assessments unscientific, while neglecting flood propagation mechanisms reduces model interpretability, hindering hydrological experts' verification and trust in the forecast results. Therefore, these issues collectively limit the application value of flood forecasting technology in practical flood control and disaster reduction work, necessitating a technical solution that can integrate multi-source data, adhere to the physical laws of river networks, and provide interpretable forecast results.

[0015] For ease of understanding, the following explains some key terms in this embodiment: Monitoring stations are physical locations set up within rivers or watersheds to continuously or periodically collect hydrological and meteorological data. These stations are typically equipped with sensors capable of measuring key hydrological parameters such as rainfall, water level, and flow rate. Rainfall refers to the depth of rainwater falling to the ground within a specific time period, usually measured in millimeters. It is one of the important input data affecting river flow and flood formation. Flow rate data refers to the amount of water passing through a specific cross-section of a river per unit time, usually measured in cubic meters per second. Flow rate is a core indicator for measuring the size of river runoff and a primary target for flood forecasting.

[0016] External climate-driven data refer to large-scale climate factors that influence regional hydrological processes, such as atmospheric circulation indices, sea surface temperature indices, and monsoon indices. These data can reflect changes in the global or regional climate system and have a significant modulating effect on rainfall and runoff.

[0017] Spatiotemporal input data refers to a comprehensive dataset that integrates local time-series data (such as rainfall and flow) from different monitoring stations with large-scale external climate-driven data, forming a dataset that contains both temporal and spatial dimensions. This data structure can provide models with more comprehensive environmental background information.

[0018] River network topology refers to the network structure describing the physical connections between various monitoring stations in a river system. It reflects the path and connectivity of water flow from upstream to downstream. An adjacency matrix is ​​a mathematical matrix used to represent the connections between nodes in a graph structure. In this matrix, the values ​​of the elements typically indicate whether there is a direct hydraulic connection between two monitoring stations; for example, 1 can represent direct connection, and 0 represents no connection. The Laplacian matrix of a symmetric graph is a special matrix calculated from the adjacency matrix and the degree matrix. This matrix can encode the connectivity and smoothness information of the graph and is often used in models such as graph neural networks to impose spatial constraints to reflect the structural relationships between nodes.

[0019] The Bayesian time series decomposition model is a time series analysis model based on Bayesian statistical principles. This model can decompose complex time series into multiple components such as trend terms, seasonal terms, and residual terms, and can provide the probability distribution and uncertainty estimation of the forecast results.

[0020] The trend term refers to the long-term, smooth-flowing component of a time series, reflecting the overall upward or downward trend of the data over time, such as long-term runoff changes caused by climate change. The seasonal term refers to the periodic, recurring component of a time series, such as the annual summer flood peak and the winter dry season.

[0021] The covariate weight matrix is ​​a parameter matrix in a Bayesian time series decomposition model, specifically used to quantify the impact of external climate-driven data on prediction targets (such as airflow). Through this matrix, the model can learn and integrate the effects of external climate factors.

[0022] The spatial weight matrix is ​​another parameter matrix in the Bayesian time series decomposition model, used to capture the strength and direction of spatial associations between different monitoring stations. This matrix reflects the mutual influence of hydrological processes among the stations.

[0023] The noise term is used in a Bayesian time series decomposition model to represent random fluctuations or residuals not explained by the model. It is typically modeled as a random variable with a specific probability distribution, providing a basis for estimating prediction uncertainty.

[0024] Topological smoothness constraints, by utilizing information about the river network topology, force the spatial weights learned by the model to remain similar or smoother among physically connected stations. This constraint ensures that the spatial correlations learned by the model conform to actual hydrophysical laws.

[0025] The prediction result refers to the model's estimate of future traffic data, while the uncertainty interval refers to the confidence range of the prediction result. It quantifies the reliability or risk of the prediction, such as the probability that the predicted traffic will fall within a certain range.

[0026] Upstream and downstream monitoring stations refer to two monitoring stations that have a direct or indirect hydraulic connection along the river's flow direction. A flow series is a set of flow observations at consecutive points in time from a given monitoring station. The time lag correlation coefficient measures the correlation between two time series at different time offsets. By calculating this coefficient, the lag time of the effect of a change in one series on a change in another series can be identified; for example, the time it takes for an upstream flow change to affect a downstream flow change. Hydrological propagation time refers to the time required for a flood wave or flow to propagate from an upstream station to a downstream station. Data anomalies refer to observations or patterns in a flow series that deviate significantly from the normal pattern, potentially indicating measurement errors or actual hydrological events. Lag propagation analysis results are reports or data obtained after analyzing hydrological propagation time, data anomalies, etc.

[0027] Weight vector similarity refers to the degree of similarity between weight vectors representing different monitoring stations in the spatial weight matrix. Physical consistency verification refers to the process of verifying whether the model's prediction results or internal parameters conform to actual hydrological and physical laws, in order to ensure the reliability and interpretability of the model.

[0028] Based on the above description, such as Figure 1 As shown, the flood prediction method proposed in this application includes the following steps: S101 generates spatiotemporal input data based on rainfall and flow data from multiple monitoring stations, as well as external climate-driven data.

[0029] In practical applications, daily or hourly rainfall and flow observations can be collected from various monitoring stations. Simultaneously, various external climate-driven data can be acquired, such as global atmospheric circulation indices, Pacific sea surface temperature indices, and Indian Ocean monsoon indices provided by meteorological agencies. This data can be directly integrated; for example, the time series of rainfall and flow from all monitoring stations can be aligned with the external climate-driven data according to timestamps and then simply stitched together to form a multi-dimensional spatiotemporal dataset. Alternatively, simple linear interpolation or mean imputation can be performed on data from different sources to handle missing data, and then these processed data can be fused to form a unified spatiotemporal input.

[0030] The external climate driving data includes at least one of the following: atmospheric circulation index; sea surface temperature index; monsoon index; wherein the types of external climate driving data are no less than 130. The external climate driving data refers to large-scale climate variables or their combined indicators that affect the Earth's climate system and, consequently, regional hydrological cycles (such as rainfall and runoff). These data typically reflect global or regional climate patterns and anomalies.

[0031] S102, construct an adjacency matrix based on the river network topology among multiple monitoring stations, and generate a symmetric graph Laplace matrix.

[0032] Specifically, direct hydraulic connectivity between monitoring stations can be identified and recorded based on Geographic Information System (GIS) data or historical hydrological maps of the watershed. For example, if station A is directly upstream of station B, a direct connection is considered to exist between them. Based on these connections, an adjacency matrix can be constructed, where the elements of the matrix indicate whether a connection exists between stations. For example, if station i is directly connected to station j, the corresponding element in the matrix can be set to 1; otherwise, it can be set to 0. Subsequently, based on the constructed adjacency matrix, the Laplace matrix of the symmetric graph is calculated using standard graph theory formulas, for example, by subtracting the adjacency matrix from the degree matrix.

[0033] S103, construct and train a Bayesian temporal decomposition model based on spatiotemporal input data and the Laplacian matrix of the symmetric graph.

[0034] like Figure 2As shown, the Bayesian time-series decomposition model includes time terms such as trend and seasonal terms, a covariate weight matrix for fusing external climate-driven data, a spatial weight matrix for capturing spatial correlations between monitoring stations, and a noise term. A symmetric graph Laplace matrix is ​​used to simultaneously apply topological smoothing constraints to the spatial weight matrix, forcing the spatial weights of adjacent monitoring stations with direct hydraulic connectivity in the river network to remain smooth.

[0035] As one implementation approach, a Gaussian process regression model can be used as the basis for the Bayesian time series decomposition model, with the functional forms of the trend and seasonal terms manually specified. For example, the trend term could be a simple linear function, and the seasonal term could be a fixed-period function. The covariate weight matrix and spatial weight matrix can be initialized with random values. During model training, the maximum likelihood estimation method can be used to optimize the model parameters by minimizing a simple loss function, such as mean squared error, to fit the aforementioned spatiotemporal input data. Topological smoothness constraints can be implemented by adding a regularization term to the loss function, which is based on the product of the aforementioned symmetric graph Laplacian matrix and the aforementioned spatial weight matrix. This constraint term directly adjusts the spatial weight matrix during each parameter update iteration, ensuring it remains similar across physically connected sites.

[0036] S104 inputs the external climate-driven data for the time to be predicted into the trained Bayesian time series decomposition model and outputs the prediction results of the flow data and the corresponding uncertainty interval.

[0037] Specifically, when flood forecasting is required, external climate-driven data for the time period to be predicted can be obtained, such as global climate indices included in weather forecasts for the next few days. This data is input into a pre-trained Bayesian time-series decomposition model. The model estimates the future flow at each monitoring station through its internal mathematical operations and outputs the corresponding predicted value. Simultaneously, the model can approximate the uncertainty range of the prediction based on simple statistics of its internal residual distribution (such as standard deviation). For example, adding or subtracting a fixed multiple of the residual standard deviation from the predicted value yields a prediction range.

[0038] S105 determines the time lag correlation coefficient based on the flow sequences of upstream and downstream monitoring stations, identifies hydrological propagation time and data anomalies, and outputs the lag propagation analysis results.

[0039] As one approach, for each pair of monitoring stations with an upstream-downstream relationship, their past flow sequence data can be collected. By calculating the cross-correlation function of these two flow sequences, the time lag at which the cross-correlation coefficient reaches its maximum value can be identified; this lag is defined as the hydrological propagation time. Simultaneously, a fixed threshold can be set; when the calculated time lag correlation coefficient falls below this threshold, or when the hydrological propagation time deviates from the historical average by more than a fixed percentage, it is marked as data anomaly. The lag propagation analysis results can be presented in a simple table listing the propagation time and correlation coefficient for each pair of upstream and downstream stations.

[0040] S106 verifies whether the similarity between the weight vectors of adjacent monitoring stations in the trained spatial weight matrix and the time lag correlation coefficient corroborate each other, so as to verify the physical consistency of the traffic data prediction results.

[0041] Specifically, after model training is complete, the Euclidean distance or cosine similarity of the weight vectors corresponding to adjacent monitoring stations in the spatial weight matrix can be calculated, and a fixed threshold can be set to determine whether they are sufficiently similar. Simultaneously, the calculated weight vector similarity is compared with the time-lag correlation coefficient obtained from the aforementioned lag propagation analysis. For example, if both the weight vector similarity and the time-lag correlation coefficient are high, they are considered to corroborate each other. When the aforementioned similarity and the aforementioned time-lag correlation coefficient corroborate each other, the prediction results are considered to have passed the physical consistency verification.

[0042] The following example will provide a more detailed explanation of the above technical solution: Suppose that in a region called "Basin A", multiple hydrological monitoring stations, such as stations S1, S2, and S3, are deployed. These stations are distributed across different river sections within the basin, with S1 located upstream of S2 and S2 upstream of S3. The hydrological department in this region needs to provide real-time, high-precision forecasts of flood conditions for the future.

[0043] First, the system continuously collects daily rainfall and flow data from these monitoring stations. Simultaneously, it obtains daily external climate-driving data from global climate databases, such as the Pacific Decadal Oscillation (PDO) and the Southern Oscillation (SOI). This data is integrated into a unified spatiotemporal input dataset, containing hydrological information for each station at different points in time, as well as contemporaneous large-scale climate background information.

[0044] Next, based on the detailed river network map of watershed A, the hydraulic connectivity between stations S1, S2, and S3 is clarified. For example, S1 directly connects to S2, and S2 directly connects to S3. Based on these connectivity relationships, an adjacency matrix is ​​constructed, which intuitively represents the physical connections between stations. Subsequently, the Laplace matrix of the symmetric graph is calculated using this adjacency matrix. This matrix encodes the topological structure information of the river network into a mathematical form, providing structural constraints for subsequent model training.

[0045] Then, a Bayesian time-series decomposition model was constructed and trained using the aforementioned spatiotemporal input data and a symmetric graph Laplace matrix. This model was designed to decompose the flow sequence at each station into long-term trends, seasonal fluctuations, and random noise. The model includes a covariate weight matrix to learn how external climate-driven data (such as PDO and SOI) affect flow within the watershed. Additionally, the model includes a spatial weight matrix to capture spatial relationships between different stations. During training, the aforementioned symmetric graph Laplace matrix was used to impose topological smoothness constraints on the spatial weight matrix. This means that when learning spatial relationships between stations, the model is forced to ensure that if S1 and S2 are directly connected in the river network, their corresponding spatial weight vectors should remain similar, thereby ensuring that the spatial dependencies learned by the model conform to the actual water flow propagation paths and avoiding physically unreasonable spatial relationships.

[0046] When flood forecasting is required, such as predicting flow rates for the next 7 days, the system inputs external climate-driven data for the next 7 days (e.g., predicted PDO and SOI values) into a pre-trained Bayesian time-series decomposition model. The model outputs daily flow rate predictions for each monitoring station (S1, S2, S3) for the next 7 days. More importantly, due to the Bayesian framework, the model also provides a corresponding uncertainty range for each prediction. For example, the predicted flow rate for station S2 is between 100 cubic meters per second and 120 cubic meters per second, providing decision-makers with a quantitative basis for risk assessment.

[0047] Simultaneously, the system calculates the time-lag correlation coefficients of the flow sequences at upstream and downstream monitoring stations (e.g., S1 and S2) based on historical flow data. Through analysis, the typical hydrological propagation time of floodwater from S1 to S2 can be determined; for example, it takes an average of 12 hours. If, within a certain period, the flow at S1 increases significantly, but the flow at S2 does not increase accordingly after 12 hours, or the calculated lag correlation coefficient is abnormally low, the system will identify this as a potential data anomaly or hydrological propagation anomaly event and output the lag propagation analysis results.

[0048] Finally, the system performs physical consistency verification on the prediction results. Specifically, it checks whether the weight vectors corresponding to adjacent monitoring stations (such as S1 and S2) in the trained spatial weight matrix are highly similar. Simultaneously, it compares this similarity with the time lag correlation coefficient between S1 and S2 obtained from the previous lag propagation analysis. If the spatial weights of S1 and S2 learned by the model are highly similar, and the lag propagation analysis also confirms a strong and expected hydrological propagation lag relationship between S1 and S2, then the prediction results are considered to have passed the physical consistency verification. This indicates that the model not only predicts accurately, but its internal understanding of hydrological processes also conforms to actual physical laws.

[0049] In summary, the above technical solution, by integrating multi-source data, introducing river network physical topology constraints, combining the probabilistic characteristics of Bayesian models, and employing hysteresis propagation analysis and physical consistency verification, achieves high-precision flood prediction that conforms to hydrophysical laws, while also possessing better interpretability and risk quantification capabilities. Compared with traditional flood prediction methods, the technical concept of this embodiment represents a significant advancement: In terms of data utilization, traditional methods often rely solely on historical flow data from a single monitoring station for prediction, failing to fully utilize multi-source heterogeneous information. For example, traditional ARIMA models typically only process univariate time series. This embodiment, however, constructs comprehensive spatiotemporal input data by fusing rainfall and flow data from multiple monitoring stations, along with up to 130 external climate driving data points. This enables the model to capture more complex flood formation mechanisms, such as the impact of large-scale climate models on regional rainfall and runoff, which is difficult to achieve with traditional methods.

[0050] Regarding the utilization of spatial connectivity, existing statistical models or deep learning models (such as LSTM) typically treat each station as independent or associate them only through simple spatial averaging when processing multi-site data, failing to effectively utilize the physical connectivity of the river network. This embodiment innovatively constructs an adjacency matrix and a symmetric graph Laplace matrix based on the river network topology between monitoring stations. During the training of the Bayesian time-series decomposition model, the aforementioned symmetric graph Laplace matrix is ​​used to apply topological smoothing constraints to the spatial weight matrix. This constraint forces the spatial associations learned by the model to conform to actual river connectivity, ensuring that the spatial weights of adjacent monitoring stations remain smooth. This allows the model to follow the physical propagation path of floods in the river network, significantly improving the model's physical rationality. For example, in the example of watershed A above, the model can learn the strong association between S1 and S2 caused by hydraulic connectivity, without learning associations that are inconsistent with physics.

[0051] Regarding the quantification of forecast uncertainty, many traditional forecasting models only provide point predictions and lack estimates of forecast uncertainty, preventing decision-makers from comprehensively assessing flood risks. The Bayesian time series decomposition model in this embodiment can output forecast results for flow data and corresponding uncertainty intervals. For example, in the example above, the model not only predicts the flow value at station S2 but also provides a confidence interval. This provides more comprehensive risk assessment information for flood control decisions, enabling decision-makers to formulate more robust response strategies based on the range of uncertainty.

[0052] In characterizing flood propagation effects, traditional models often lack a dedicated mechanism to analyze and explain the propagation process of floods in river networks. This embodiment determines the time lag correlation coefficient based on the flow sequences of upstream and downstream monitoring stations, identifies hydrological propagation time and data anomalies, and outputs the lag propagation analysis results. For example, in the example above, the system can clearly determine the time required for the flood to propagate from S1 to S2 and can identify abnormal propagation patterns. This compensates for the shortcomings of traditional models in characterizing the lag effect of flood propagation and enhances the interpretability of the model.

[0053] Finally, this embodiment verifies the physical consistency of the predicted flow data by checking whether the similarity of the weight vectors of adjacent monitoring stations in the trained spatial weight matrix corroborates the aforementioned time lag correlation coefficient. This step cross-validates the spatial patterns learned by the model with the observed hydrological propagation patterns, ensuring the reliability and physical rationality of the model's predictions. For example, in the above example, if the spatial weights of S1 and S2 learned by the model are highly similar, and actual observations also show a strong lag correlation between S1 and S2, then the effectiveness of the model is further confirmed. This physical consistency verification is not available in traditional black-box models, greatly enhancing the model's credibility and application value.

[0054] In summary, this embodiment provides a high-precision, physically interpretable flood forecasting method with risk quantification capabilities, effectively overcoming the limitations of existing technologies in multi-source data utilization, spatial connectivity modeling, uncertainty estimation, and flood propagation effect characterization, and providing more reliable technical support for basin-level flood management.

[0055] In some schemes of this application, after training the Bayesian time series decomposition model, the method further includes: dividing the watersheds to which the multiple monitoring stations belong into training watersheds and test watersheds; using a masking method to make the update of model parameters depend only on the monitoring station data of the training watershed; and evaluating the cross-watershed generalization ability of the Bayesian time series decomposition model based on the monitoring station data of the test watershed.

[0056] Specifically, the watersheds to which the multiple monitoring stations belong are divided into training and testing watersheds. This aims to provide a basis for evaluating the model's cross-watershed generalization ability and to simulate the model's actual performance in unseen watersheds. This division can be based on Geographic Information System (GIS) data, dividing the entire study area into several independent watershed units according to geographical features such as watershed boundaries and river system distribution. These watershed units are then randomly or according to preset rules (e.g., based on geographical proximity or similarity of hydrological characteristics) assigned to the training and testing watershed groups. Alternatively, hydrological experts can manually specify which watersheds are used for training and which for testing based on historical data, hydrological characteristics, administrative divisions, and other empirical factors. For example, watersheds with significantly different hydrological characteristics can be selected as testing watersheds to more rigorously evaluate the model's generalization ability.

[0057] By masking the model parameters so that updates depend only on monitoring station data from the training watershed, the Bayesian time series decomposition model is ensured not to "see" test watershed data during training, thus guaranteeing the objectivity and accuracy of the generalization assessment. One implementation involves applying a zero or minimal weight to the loss term corresponding to the data from the test watershed monitoring stations during the loss function calculation phase of model training, excluding it from gradient calculation and model parameter updates. Another implementation involves completely removing the test watershed data from the training dataset during the data loading or preprocessing phase, or providing the model with only training watershed data batches during training iterations, excluding test watershed data batches.

[0058] The cross-basin generalization ability of the Bayesian time-series decomposition model is evaluated based on monitoring station data from the test watershed. This aims to quantify the model's predictive performance in untrained watersheds and measure the generalizability of the hydrological patterns learned by the model. This can be achieved by calculating error indices between the model's predictions and actual observations in the test watershed, such as root mean square error (RMSE), mean absolute error (MAE), or Nash efficiency coefficient (NSE). Furthermore, these error indices can be combined to calculate a cross-basin generalization index, which more intuitively reflects the model's transferability across different watersheds.

[0059] The following concrete example illustrates this. Consider a flood forecasting system deployed across multiple watersheds. To evaluate the system's cross-watershed generalization ability, the watersheds to which all monitoring stations belong can be divided into training watersheds (e.g., containing stations in watershed A) and test watersheds (e.g., containing stations in watersheds B and C). When training the Bayesian time series decomposition model, the system uses a masking mechanism to ensure that the model parameter updates depend only on the monitoring station data in watershed A. This means that when the model learns parameters such as the spatial weight matrix, trend term, and seasonal term under the river network topology constraints, it does not utilize any data from watersheds B and C. After training, the model will use the knowledge learned in watershed A to predict flow rates from the monitoring station data in watersheds B and C. By comparing the model's predictions in watersheds B and C with the actual observations, evaluation metrics such as the root mean square error (RMSE) can be calculated, thereby quantifying the generalization ability of the Bayesian time series decomposition model in the untrained watersheds. For example, if the RMSE of the model is significantly lower than that of the traditional model in watersheds B and C, it indicates that it has good cross-watershed generalization ability.

[0060] Through the above technical solution, this application effectively addresses the problem of traditional flood forecasting methods lacking a cross-basin generalization capability assessment mechanism. This solution can simulate the model's actual performance in unfamiliar, untrained basins during the model training phase and quantitatively measure the cross-basin transferability of the hydrological patterns learned by the model, thereby guiding the model to learn general patterns that conform to common hydrophysical properties. This significantly reduces the risk of accuracy degradation in cross-basin applications, greatly ensuring the reliability and practicality of model predictions in unfamiliar basins, and providing a solid foundation for the widespread deployment and application of flood forecasting systems.

[0061] This application further proposes a method for evaluating the cross-basin generalization ability of the Bayesian time series decomposition model, which includes: calculating the time prediction error of the model in the training basin; calculating the cross-basin prediction error of the model in the test basin; and calculating a cross-basin generalization index based on the time prediction error and the cross-basin prediction error, wherein the cross-basin generalization index is negatively correlated with the cross-basin prediction error and positively correlated with the time prediction error.

[0062] The time prediction error of the calculation model on the training watershed refers to the deviation generated when the model predicts flow data on the training watershed that has been used for parameter updates. This error quantifies the model's fitting ability and short-term prediction accuracy on known data. Various statistical indicators can be used to calculate the time prediction error. For example, the root mean square error (RMSE) can be calculated, which is obtained by averaging the squares of the differences between the predicted and observed values ​​and then taking the square root, effectively reflecting the dispersion of the predicted values; or the mean absolute error (MAE) can be calculated, which is obtained by averaging the absolute values ​​of the differences between the predicted and observed values, and can intuitively reflect the average magnitude of the prediction error.

[0063] The cross-basin prediction error of the computational model in the test basin refers to the deviation generated when the model predicts flow data in the test basin where it has not participated in parameter updates. This error directly reflects the model's generalization ability and transfer performance in unseen basins. Similar to time prediction error, cross-basin prediction error can also be quantified using metrics such as root mean square error (RMSE) or mean absolute error (MAE). For example, the RMSE between the predicted values ​​and actual observations at all monitoring stations in the test basin can be calculated to evaluate the model's overall prediction performance across the entire test basin.

[0064] Based on the time prediction error and the cross-basin prediction error, a cross-basin generalization index is calculated. The cross-basin generalization index is negatively correlated with the cross-basin prediction error and positively correlated with the time prediction error. The cross-basin generalization index is a comprehensive quantitative indicator used to measure the ability of a Bayesian temporal decomposition model to migrate and adapt to new environments across different watersheds. The calculation method aims to make it negatively correlated with the cross-basin prediction error (i.e., the smaller the cross-basin prediction error, the higher the generalization index); and positively correlated with the time prediction error (i.e., the smaller the time prediction error (the better the model performs in the training watershed), the higher the generalization index). As a specific implementation, the index can be calculated using the formula `G = 1 - E_cross / E_temporal`, where `E_cross` represents the cross-basin prediction error and `E_temporal` represents the time prediction error. This formula intuitively reflects the model's performance in the test watershed relative to its performance in the training watershed. In addition, other composite indicators can be designed, such as through weighted averages or ratios, to ensure that they meet specific correlations with the two types of errors, thereby providing a unified standard for evaluating generalization ability.

[0065] This scheme establishes a performance benchmark for the model in known environments by calculating the temporal prediction error of the model in the training watershed. The data from the training watershed is used for model parameter updates; therefore, its temporal prediction error reflects the model's ability to fit the learned hydrological data. Based on this, the cross-watershed prediction error of the model in the test watershed is calculated. The data from the test watershed is not used for model parameter updates; therefore, its prediction error directly reflects the model's actual transfer and adaptation capabilities in unseen watersheds. Finally, a cross-watershed generalization index is calculated based on these two types of errors. This index is designed to be negatively correlated with the cross-watershed prediction error—that is, the smaller the prediction error in the test watershed, the stronger the model's generalization ability, and the higher the index value; and positively correlated with the temporal prediction error in the training watershed—that is, the better the model's fit in the training watershed, the higher the index value. This design ensures that the generalization index can comprehensively and objectively reflect the overall performance of the model, considering both the model's basic fitting ability and quantifying its adaptability in unfamiliar environments.

[0066] As a specific implementation method, when evaluating the cross-basin generalization ability of a Bayesian time-series decomposition model, a large watershed can first be divided into multiple sub-basins, for example, dividing multiple hydrological monitoring stations into three sub-basins: A, B, and C. During the model training phase, watershed A can be designated as the training watershed, while watersheds B and C are designated as test watersheds. Model parameter updates are based solely on monitoring station data from watershed A, while data from watersheds B and C are used for subsequent generalization ability evaluation. Specifically, after model training is complete, the model's time prediction error in watershed A (training watershed) can be calculated. For example, the root mean square error (RMSE_A) of the model's flow data prediction in watershed A can be calculated. Subsequently, the cross-basin prediction errors of the model in watersheds B and C (test watersheds) are calculated, for example, RMSE_B and RMSE_C, respectively. Then, the cross-basin generalization index is calculated based on these errors. For example, the formula `G = 1 - E_cross / E_temporal` can be used. Here, `E_temporal` can be RMSE_A, and `E_cross` can be RMSE_B or RMSE_C (or the average RMSE of the test domain). This yields a generalization index G between 0 and 1. For example, if the model's RMSE_A in training domain A is 10 m³ / s and its RMSE_B in test domain B is 12 m³ / s, then the generalization index G_B for domain B is 1 - 12 / 10 = -0.2. If RMSE_B is 8 m³ / s, then G_B = 1 - 8 / 10 = 0.2. The higher this index, the better the model performs in the test domain compared to the training domain, i.e., the stronger its generalization ability. This quantification method allows for a direct comparison of the generalization performance of different models or different training strategies.

[0067] Through the above technical solution, this application provides an effective means to quantitatively evaluate the cross-basin generalization ability of Bayesian time-series decomposition models. By calculating the model's time prediction error in the training basin and its cross-basin prediction error in the test basin, and constructing a cross-basin generalization index based on this, it solves the problem of the lack of a unified and quantifiable evaluation standard in traditional methods. This index can intuitively and clearly measure the model's adaptability in unseen basins, allowing for standardized comparisons of the generalization performance of different models or optimization strategies. This not only provides clear guidance for model optimization but also provides a reliable basis for the actual deployment and application decisions of flood prediction systems, thereby significantly improving the reliability and practical value of flood prediction models in cross-basin applications.

[0068] In some schemes of this application, training a Bayesian temporal decomposition model based on the spatiotemporal input data and the symmetric graph Laplacian matrix includes: calculating the gradient between the symmetric graph Laplacian matrix and the spatial weight matrix; iteratively correcting the spatial weight matrix based on the gradient during model training; and maintaining the spatial continuity of spatial weights of adjacent monitoring stations in high-flow-variability regions through the iterative correction.

[0069] The calculation of the gradient between the symmetric graph Laplacian matrix and the spatial weight matrix aims to quantify the deviation between the current spatial weight matrix and the smoothing constraints implicit in the river network topology, and to indicate how to adjust the spatial weight matrix to better conform to these constraints. The symmetric graph Laplacian matrix encodes the connectivity information of the river network, and its product or related operations with the spatial weight matrix reflect the smoothness of the spatial weights in the topology. This gradient can be directly calculated in the model training framework using automatic differentiation techniques, such as using the automatic differentiation function provided by a deep learning framework to differentiate the loss function containing the topology smoothing term with respect to the spatial weight matrix. Alternatively, the partial derivatives of the topology smoothing constraint term in the loss function with respect to the spatial weight matrix can be derived analytically, and then calculated according to the derived formula.

[0070] During model training, iteratively correcting the spatial weight matrix based on the gradient is an optimization technique. By repeatedly applying the correction steps, the parameters (in this case, the spatial weight matrix) are gradually brought closer to the target state (which better conforms to the topological smoothness constraint). Gradient-based correction means that the direction of correction is related to the gradient direction, typically reducing loss or bias along the opposite direction of the gradient. Integrating the correction into the model training process ensures that the spatial weight matrix is ​​continuously guided to conform to the physical topology of the river network while fitting the data. This can be achieved using gradient descent or its variants, updating the spatial weight matrix with a certain learning rate based on the calculated gradient in each iteration. Alternatively, the topological smoothness constraint term can be added as a regularization term to the model's joint loss function. When minimizing the joint loss function, the optimizer automatically corrects the spatial weight matrix based on the gradient of this regularization term.

[0071] Maintaining spatial continuity of spatial weights for adjacent monitoring stations in high-flow-variability regions through iterative correction means that geographically adjacent monitoring stations exhibit similar or smoothly transitioning spatial weights. In high-flow-variability regions, hydrological processes are complex, and flow changes are drastic; traditional static topological smoothing constraints may be insufficient to capture this complexity. Through iterative correction, the model can dynamically adjust spatial weights to maintain smoothness consistent with the river network topology even in highly variable regions. The iterative correction process can include an adaptive learning rate mechanism, allowing the correction step size to be adjusted based on local hydrological characteristics in high-flow-variability regions, thus guiding the spatial weights towards smoothness more precisely. Alternatively, a momentum term or adaptive gradient scaling can be introduced into the iterative correction to accelerate convergence and avoid oscillations in highly variable regions, thereby more stably maintaining the continuity of spatial weights.

[0072] As a specific implementation, when training a Bayesian time series decomposition model, a joint loss function can be constructed, which includes a data fitting term (e.g., measuring the error between model predictions and actual observed flow) and a topology smoothing constraint term. This topology smoothing constraint term can be represented as `λ_topo × trace(W^TLW)`, where `W` is the spatial weight matrix, `L` is the Laplacian matrix of the symmetric graph, and `λ_topo` is a hyperparameter balancing data fitting and topology smoothing. In each iteration of model training, the optimizer (e.g., a variational inference-based optimizer) computes the gradient of this joint loss function with respect to the spatial weight matrix `W`. This gradient includes contributions from the data fitting term and contributions from the topology smoothing constraint term. For the topology smoothing constraint term `trace(W^TLW)`, its gradient with respect to `W` is `2LW`. The optimizer integrates these gradients and updates the spatial weight matrix `W` according to a preset learning rate. For example, if an element of `W` causes a significant increase in the topological smoothness constraint (i.e., the spatial weights are not smooth at that point), the gradient indicates a direction that allows that element of `W` to reduce this non-smoothness after the update. Through this continuous, gradient-based iterative correction, even when the flow in a certain river segment changes drastically, leading to large differences in data between adjacent stations, the model can still use gradient information to forcibly adjust the spatial weights, keeping them smooth and continuous in physically connected river segments. This prevents the spatial weight learning results from deviating from the actual hydraulic connectivity in highly variable regions.

[0073] Through the above technical solution, this application effectively addresses the shortcomings of traditional fixed topological smoothing constraints in the face of hydrological complexity in areas with high flow variability. By calculating the gradient between the Laplacian matrix of the symmetric graph and the spatial weight matrix, and iteratively correcting the spatial weight matrix based on this gradient, the model can dynamically and accurately adjust the spatial weights, maintaining spatial continuity of spatial weights between adjacent monitoring stations even in areas with high flow variability. This ensures that even under complex conditions of drastic hydrological changes, the spatial weights learned by the model always conform to the physical laws of hydraulic connectivity in river networks, thus significantly improving the physical consistency and robustness of the Bayesian time series decomposition model under complex hydrological conditions, providing a more reliable and accurate foundation for flood prediction.

[0074] In one improved embodiment of this application, before generating the spatiotemporal input data, the method further includes: identifying months with more than a preset threshold of missing days; constructing a probabilistic model using the temporal correlation of adjacent monitoring stations in the same watershed; and probabilistically reconstructing the missing values ​​through Bayesian inference to generate the imputed complete data for use in generating the spatiotemporal input data.

[0075] The feature identifies months with more than a preset threshold of missing data days, aiming to accurately filter out periods with severe data loss that may significantly impact subsequent analysis. The preset threshold can be set based on the actual application scenario, data quality requirements, or historical experience; for example, it can be set to 5, 10, or 15 consecutive days of missing data. By setting a threshold, unnecessary complex processing of a small amount of sporadic missing data can be avoided, thereby optimizing the efficiency of the data processing workflow and focusing computational resources on resolving critical missing data issues that affect data integrity.

[0076] Constructing probabilistic models based on the temporal correlations of adjacent monitoring stations within the same watershed is based on the physical characteristics of hydrological processes. Within the same watershed, hydrological processes (such as rainfall and flow) at adjacent monitoring stations are often influenced by similar climatic conditions and geographical environments, thus exhibiting significant correlations between their time-series data. Various methods can be employed to construct probabilistic models. For example, regression analysis (such as linear and nonlinear regression) can be used based on historical data to establish a mapping relationship between missing stations and adjacent stations; time-series models (such as Vector Autoregression, VAR) can be used to capture dynamic dependencies among multiple stations; or machine learning methods (such as random forests and support vector machines) can be used to learn complex nonlinear correlations. These models can effectively utilize known data to infer potential patterns in missing data.

[0077] By using Bayesian inference to probabilistically reconstruct missing values ​​and generate imputed complete data for use in generating the spatiotemporal input data, the core principle is to generate a probability distribution representing the possible range of missing values ​​rather than providing a single deterministic imputed value. Bayesian inference methods can employ techniques such as Markov Chain Monte Carlo (MCMC) algorithms (e.g., Gibbs sampling, Metropolis-Hastings algorithm) or variational inference to sample or approximately infer the posterior distribution of missing values ​​from the constructed probabilistic model. This probabilistic reconstruction can more realistically reflect the inherent uncertainty of the missing data itself, providing a more robust input for subsequent flood prediction models and helping to quantify the uncertainty of prediction results.

[0078] The following is a concrete example to illustrate this. Suppose that at monitoring station A in a certain river basin, the flow data for July 2020 is missing for 10 consecutive days due to equipment failure. According to this solution, the system first identifies that the number of missing days (10 days) for station A in July 2020 exceeds a preset threshold (e.g., set to 5 days), thus determining that imputation processing is needed for that month. Next, the system uses the flow data of station A and neighboring monitoring stations B and C in the same historical period (e.g., July 2019, July 2021, etc.), as well as the flow data of station A itself during the non-missing period, to construct a probabilistic model. This model can be a multiple regression model based on historical correlation, or a Bayesian network model considering time series characteristics. Through this model, combined with Bayesian inference methods, the system will probabilistically reconstruct the 10 days of missing flow data for station A in July 2020. This means that the system will not simply provide a single estimate, but will generate a probability distribution containing multiple possible values. For example, it may obtain a mean and standard deviation, or a series of Monte Carlo sampling values, which together describe the possible range of missing flow. Ultimately, these probabilistically reconstructed missing values ​​are combined with the original observation data to form the imputed complete flow data, which is then used to generate spatiotemporal input data that includes rainfall, flow, and external climate-driven data.

[0079] Through the above technical solution, this application can effectively solve the serious problem of missing data in actual hydrological monitoring data and avoid the bias introduced by traditional simple interpolation methods. This solution, through targeted missing data detection and Bayesian probabilistic interpolation that conforms to hydrophysical laws, provides high-quality, highly reliable, and complete data for subsequent generation of spatiotemporal input data, significantly improving the training effect and final prediction accuracy of the flood prediction model, and enhancing the robustness of the entire flood prediction process.

[0080] In some schemes of this application, the trend term and seasonal term in the Bayesian time series decomposition model are constructed as follows: a Matérn-1.5 kernel function is assigned to the trend term to characterize smooth long-term climate change; a periodic kernel function is assigned to the seasonal term, and periodic parameters of 12 months and 60 months are set to capture seasonal and interannual fluctuations; during model training, the hyperparameters of the Matérn-1.5 kernel function and the periodic kernel function are jointly optimized.

[0081] As a specific implementation method, Gaussian process regression framework can be used to model the trend and seasonal terms when constructing a Bayesian time series decomposition model. For the trend term, a Matérn-1.5 kernel function can be instantiated, for example, using `MaternKernel(nu=1.5)` from Gaussian process libraries such as GPyTorch or GPflow. This kernel function will automatically include a length scale hyperparameter to control the smoothness of trend changes. For the seasonal term, a periodic kernel function can be instantiated, for example, using `PeriodicKernel()`. This periodic kernel function will include a period length hyperparameter and a length scale hyperparameter. To capture the periodicity of 12 months and 60 months, two periodic kernel functions can be superimposed, i.e., `PeriodicKernel(period=12) + PeriodicKernel(period=60)`, or the period parameter of a periodic kernel function can be set to learnable, and it can be guided to converge to 12 months and 60 months during training. During model training, the hyperparameters of these kernel functions can be optimized by maximizing the log-marginal likelihood. Specifically, a joint loss function can be constructed, which includes a data fitting term and a regularization term, where the hyperparameters of the kernel functions are the variables to be optimized. This joint loss function is then iteratively optimized using a gradient descent-based optimizer, such as the Adam optimizer. In each iteration, the gradient of the loss function with respect to all kernel function hyperparameters is calculated, and the hyperparameters are updated based on the gradient. For example, an initial learning rate can be set and gradually decayed as training progresses to ensure the stability and convergence of the optimization process. Through this joint optimization, the model can adaptively learn the kernel function parameters that best reflect the trends and seasonal characteristics of hydrological time series, thus making the decomposition results of the trend and seasonal terms more accurate.

[0082] Through the above technical solutions, this application effectively addresses the limitations of traditional flood prediction models in modeling trend and seasonal terms. By specifying a Matérn-1.5 kernel function for the trend term, the model can characterize long-term climate change with appropriate smoothness, avoiding overly rigid or flexible trend fitting, and making the trend component more closely resemble actual hydrological and physical processes. Simultaneously, by specifying a periodic kernel function for the seasonal term and setting periodic parameters of 12 months and 60 months, the model can accurately capture intra-annual seasonal variations and inter-annual periodic fluctuations in hydrological data, comprehensively covering periodic patterns at different scales. During model training, the hyperparameters of the two kernel functions are jointly optimized to ensure that the fitting results of the trend and seasonal terms are mutually coordinated and adapted, reducing the bias in time series decomposition. This refined modeling approach enables the Bayesian time series decomposition model to more accurately and physically conform to the different components in the flow time series, providing a reliable foundation for subsequent flood prediction and significantly improving prediction accuracy and the physical interpretability of the model. Combining the overall framework of the Bayesian time series decomposition model in the above flood prediction methods, this optimization of the trend and seasonal terms enables the model to more accurately capture the inherent dynamics of hydrological processes by fusing multi-source heterogeneous data and applying river network topological constraints, thereby outputting more reliable flow data prediction results and corresponding uncertainty intervals.

[0083] In some schemes of this application, the method for verifying whether the similarity of the weight vectors of adjacent monitoring stations in the trained spatial weight matrix and the time lag correlation coefficient mutually corroborate each other includes: determining whether the similarity of the weight vectors of adjacent monitoring stations is greater than a preset similarity threshold; determining whether the peak value of the time lag correlation coefficient is greater than a preset correlation coefficient threshold; when the similarity is greater than the preset similarity threshold and the peak value is greater than the preset correlation coefficient threshold, determining that the similarity and the time lag correlation coefficient mutually corroborate each other, and the prediction result is verified by physical consistency.

[0084] In the Bayesian time-series decomposition model, each monitoring station corresponds to a spatial weight vector, which represents the station's contribution to flow prediction in the spatial dimension or the strength of its association with other stations. The similarity of weight vectors between adjacent monitoring stations refers to the degree of proximity in the feature space between the corresponding spatial weight vectors of two monitoring stations with direct hydraulic connectivity in the river network. This similarity quantifies whether the spatial associations learned by the model conform to the hydraulic connectivity constraints defined by the river network topology. For example, it can be measured by calculating the cosine similarity of two weight vectors; the closer the cosine similarity value is to 1, the more consistent the directions of the two vectors, and the higher the similarity. Alternatively, the similarity can be represented by calculating the Euclidean distance between two weight vectors and then taking its reciprocal or transforming it using an exponential function; the smaller the distance, the higher the similarity. A preset similarity threshold is a predetermined value used to determine whether the calculated similarity meets the expected physical reasonableness standard. This threshold can be set based on historical hydrological data analysis, expert experience, or the model's performance on the validation set. The peak value of the time-lag correlation coefficient refers to the maximum correlation value reached by the cross-correlation function at a specific time lag when cross-correlation analysis is performed on the flow sequences of upstream and downstream monitoring stations. This peak value directly reflects the correlation strength between the two at the optimal propagation time point during the propagation of upstream flow changes downstream. For example, the Pearson correlation coefficient can be used to calculate the correlation of flow sequences under different time lags, and its maximum value can be taken as the peak value. In addition, when there are nonlinear relationships or outliers in the flow sequences, nonparametric methods such as Spearman's rank correlation coefficient or Kendall's rank correlation coefficient can also be used for calculation. The preset correlation coefficient threshold is a pre-set value used to determine whether the identified time-lag correlation is significant enough to confirm that there is a clear hydraulic connectivity and propagation relationship between upstream and downstream stations. This threshold can be determined based on hydrophysical laws, statistical principles, or historical data analysis. The final judgment step is the final judgment logic for physical consistency verification. Its role is to comprehensively evaluate whether the spatial correlation patterns learned by the model (reflected by the similarity of spatial weight vectors) and the actually observed hydrological propagation laws (reflected by the peak value of the time-lag correlation coefficient) support and corroborate each other. A model's predictions are considered physically reasonable and consistent only when both conditions are met: the spatial weights of adjacent sites learned by the model are sufficiently similar, and there is a sufficiently strong lag correlation between the observed upstream and downstream flow sequences. This logical judgment can be achieved using the logical AND operation in programming languages, ensuring that a model's predictions are considered reliable only when they conform to physical laws in both the "learning" and "observation" dimensions.

[0085] The following is a concrete example to illustrate this. Suppose there is an upstream monitoring station A and a downstream monitoring station B in a river basin, with a direct hydraulic connection between them. After the Bayesian time series decomposition model is trained, the model learns spatial weight vectors for stations A and B respectively. To verify the physical consistency of the model, firstly, the cosine similarity between these two spatial weight vectors is calculated, for example, the result is 0.93. Meanwhile, the preset similarity threshold is 0.85. Since 0.93 is greater than 0.85, the first condition is satisfied. Secondly, cross-correlation analysis is performed on the historical flow sequences of stations A and B, and the peak value of the time-lag correlation coefficient is 0.89. This peak occurs at a lag of 2 days, indicating that the flow change at station A affects station B approximately 2 days later. Meanwhile, the preset correlation coefficient threshold is 0.70. Since 0.89 is greater than 0.70, the second condition is also satisfied. Given that both of the above conditions are met—namely, the similarity of the spatial weight vectors is greater than the preset similarity threshold, and the peak value of the time lag correlation coefficient is greater than the preset correlation coefficient threshold—this scheme determines that the similarity between site A and site B and the time lag correlation coefficient mutually corroborate each other, thereby confirming that the traffic prediction results for this area have passed the physical consistency verification.

[0086] Through the above technical solution, this application provides a clear and executable physical consistency verification criterion, effectively solving the problem of the lack of a stable and accurate verification mechanism in traditional schemes. This scheme, by simultaneously considering the similarity of the spatial weight vectors learned by the model and the time lag correlation coefficients observed in actual data, can stably and accurately confirm whether flood prediction results conform to the actual hydrological and physical laws of the river network. This avoids situations where prediction results may meet numerical accuracy standards but fail to conform to hydrological and physical logic, thus significantly ensuring the physical rationality and interpretability of the prediction results and improving the reliability of flood prediction.

[0087] In some schemes of this application, the steps of training a Bayesian temporal decomposition model include: constructing a joint loss function, which includes a data fitting term for fitting the observed data and a topological smoothing constraint term constructed based on the Laplacian matrix of the symmetric graph and the spatial weight matrix; optimizing the joint loss function by maximizing the lower bound of evidence using a variational inference algorithm; and synchronously iteratively updating the model parameters of the Bayesian temporal decomposition model and the spatial weight matrix during the optimization process, so that the spatial weight matrix, while fitting the spatiotemporal input data, follows the hydraulic connectivity relationship defined by the river network topology.

[0088] As a specific implementation method, when training the aforementioned Bayesian time-series decomposition model, a joint loss function L can be constructed, which has the form: `L = L_data + λ * L_topo`. Here, `L_data` is the data fitting term, which can be a negative log-likelihood function. Assuming the observed data follows a Gaussian distribution, then `L_data = -Σ log P(y_t | x_t,θ)`, where `y_t` is the observed flow data at time t, `x_t` is the spatiotemporal input data, and `θ` is the model parameter. `L_topo` is the topological smoothing constraint term, which can be specifically represented as `trace(W^TLW)`, where W is the spatial weight matrix, and L is the symmetric graph Laplacian matrix. `λ` is a hyperparameter used to balance the importance of data fitting and topological smoothing constraints. During model training, a stochastic variational inference algorithm can be used to maximize the lower bound of evidence for this joint loss function. Specifically, in each iteration, a small batch of data is randomly sampled from the spatiotemporal input data. The gradients of `L_data` and `L_topo` corresponding to this small batch of data are calculated, and then these two gradients are weighted and summed to obtain the gradient of the joint loss function. Next, gradient descent algorithms such as the Adam optimizer are used to synchronously update the model parameters of the Bayesian time-series decomposition model (such as the hyperparameters of the trend and seasonal terms, the parameters of the covariate weight matrix, etc.) and the elements of the spatial weight matrix W. Through this synchronous iterative update, while the spatial weight matrix W fits the observed data, the weight differences between adjacent monitoring stations are penalized by the topological smoothing constraint term, thereby forcing it to conform to the hydraulic connectivity relationships defined by the river network topology.

[0089] Through the above technical solution, this application effectively addresses the problem of simultaneously optimizing data fitting requirements and river network topological smoothness constraints during the training of Bayesian time-series decomposition models. By constructing a joint loss function that includes data fitting terms and topological smoothness constraint terms, and optimizing it using a variational inference algorithm, the synchronous iterative update of the model parameters and the spatial weight matrix of the Bayesian time-series decomposition model is achieved. This ensures that the spatial weight matrix, while fitting the spatiotemporal input data, strictly adheres to the hydraulic connectivity relationships defined by the river network topology, thereby ensuring that the spatial correlations learned by the model conform to actual hydrophysical laws. Ultimately, this significantly improves the accuracy and physical rationality of flood prediction, providing a foundation for more reliable flood warnings and management.

[0090] In other embodiments of this application, training a Bayesian temporal decomposition model based on spatiotemporal input data and a symmetric graph Laplacian matrix includes the following specific operations: During each iteration of the online update of the model parameters using a variational inference algorithm, the following operations are performed synchronously. This mechanism ensures that the application of topological smoothing constraints and the correction of the spatial weight matrix are not completed all at once or statically, but are tightly coupled with the model parameter learning process in each iteration of the variational inference algorithm, thereby maintaining the model's accurate learning of hydraulic connectivity relationships.

[0091] Specifically, firstly, a topology smoothing constraint term is calculated based on the current spatial weight matrix and the Laplacian matrix of the symmetric graph. The core function of this constraint term is to measure the difference between the spatial weight vectors of adjacent monitoring stations with direct hydraulic connectivity, thus intuitively reflecting the degree to which the current spatial weight matrix adheres to the physical connectivity of the river network. For example, the topology smoothing constraint term can be expressed in the form `trace(W^TLW)`, where `W` is the spatial weight matrix and `L` is the Laplacian matrix of the symmetric graph; alternatively, for each pair of directly connected monitoring stations, the Euclidean distance or cosine similarity between their corresponding spatial weight vectors is calculated, and the sum of the squares of these differences is used as the constraint term. Subsequently, the topology smoothing constraint term is added as a regularization term to the loss function during model training, and the gradient of the topology smoothing constraint term with respect to the spatial weight matrix is ​​calculated. Adding it as a regularization term to the loss function ensures that the model, during optimization, not only strives for the best fit to the observed data but also endeavors to reduce the difference between the spatial weight vectors of adjacent monitoring stations. Calculating its gradient with respect to the spatial weight matrix indicates how to adjust the spatial weight matrix to most effectively reduce the value of the topology smoothing constraint term. Finally, the spatial weight matrix is ​​modified according to the gradient to force the spatial weights of adjacent monitoring stations to gradually smooth out during iterative updates. This modification process is iterative and continuous, forcing the spatial weights of adjacent monitoring stations to gradually smooth out in each iteration, thereby ensuring that the spatial associations learned by the model remain highly consistent with the physical connectivity of the river network throughout the entire training process.

[0092] As a specific implementation method, the following steps can be used to dynamically apply topological smoothing constraints when training a Bayesian temporal decomposition model. Assume the Bayesian temporal decomposition model has been initialized, and its model parameters (including the spatial weight matrix) are being updated online using a variational inference algorithm. At the beginning of each iteration of the variational inference, the spatial weight matrix for the current iteration step is first obtained. Simultaneously, a symmetric graph Laplacian matrix, pre-constructed based on the river network topology between monitoring stations, is used. Next, the topological smoothing constraint term is calculated. For example, this constraint term can be defined as `λ_topo *trace(W^TLW)`, where `λ_topo` is a preset hyperparameter used to control the strength of the topological smoothing constraint. This expression effectively measures the difference between the spatial weight vectors of adjacent monitoring stations with direct hydraulic connectivity; the greater the difference, the larger the value of the constraint term. Subsequently, this calculated topological smoothing constraint term is added as a regularization term to the joint loss function of the current iteration. For example, if the original loss function is `Loss_data` (which measures the fit between the model's predictions and the actual observed data), then the new joint loss function will become `Loss_total = Loss_data + λ_topo * trace(W^TLW)`. Next, to optimize this joint loss function, its gradient with respect to the spatial weight matrix needs to be calculated. Specifically, the gradient of the topological smoothing constraint term `λ_topo * trace(W^TLW)` with respect to the spatial weight matrix is ​​calculated, resulting in `2 * λ_topo * L * W`. This gradient is combined with the gradient of `Loss_data` with respect to the spatial weight matrix (if it exists) to obtain the total gradient of the joint loss function with respect to the spatial weight matrix. Finally, based on this total gradient, the spatial weight matrix is ​​corrected using gradient descent or other optimization algorithms. For example, the spatial weight matrix can be updated as `W_new = W_old - learning_rate * Gradient_total`. In this way, the spatial weight matrix is ​​forced to be adjusted in each iteration, so that the spatial weight vectors of its adjacent monitoring stations gradually become smoother, thereby ensuring that the model always follows the physical connectivity of the river network during the learning process.

[0093] This scheme effectively solves the problem that traditional static topological constraints cannot consistently guarantee the smoothness of spatial weights by synchronously calculating the topological smoothness constraint term in each iteration of the variational inference online update of the Bayesian time-series decomposition model, adding it as a regularization term to the loss function, and adjusting the spatial weight matrix according to the gradient. Through this dynamic, iterative correction mechanism, the model can continuously encode the hydraulic connectivity relationships defined by the river network topology into the spatial weight matrix, ensuring that the spatial weights of adjacent monitoring stations remain smooth throughout the training process, thus avoiding deviations from physical requirements during iterations. This enables the model to more accurately learn and capture the true spatial relationships between monitoring stations, significantly improving the physical consistency and accuracy of flood prediction results, especially in complex river network environments, providing more reliable predictions.

[0094] In some embodiments of this application, the method for outputting the prediction results of traffic data and the corresponding uncertainty interval includes: estimating the probability distribution of the prediction results based on the noise term in the Bayesian time series decomposition model; calculating the confidence interval at a preset confidence level according to the probability distribution; and outputting the confidence interval as the uncertainty interval to quantify the risk range of the prediction results.

[0095] As a specific implementation method, when predicting the flow rate of a river basin for the next 24 hours, a pre-trained Bayesian time-series decomposition model is first used. This model has been trained based on historical rainfall, flow data, and external climate-driven data, and the statistical parameters of its internal noise term have been determined. When the external climate-driven data for the time to be predicted is input, the Bayesian time-series decomposition model outputs a probability distribution of the flow rate for the next 24 hours, such as a Gaussian distribution with the predicted point value as the mean and a specific variance as its characteristic. Assuming the flood control command needs a 90% confidence level to assess flood risk, the system calculates the confidence interval at the 90% confidence level based on the mean and variance of the Gaussian distribution. For example, if the predicted point value is 1000 cubic meters per second, the calculated confidence interval might be [950 cubic meters per second, 1050 cubic meters per second]. Finally, the system outputs this interval [950, 1050] as the uncertainty interval. This means that, according to the model's prediction, there is a 90% probability that the flow rate in the next 24 hours will fall between 950 and 1050 cubic meters per second. Decision-makers can use this range of information, combined with the risk levels corresponding to different flow thresholds, to more accurately assess the likelihood and potential impact of flooding, thereby taking more prudent and effective preventative measures.

[0096] Through the above technical solution, this application effectively solves the problems of traditional flood forecasting methods that only output deterministic results, cannot quantitatively estimate forecast uncertainty, and lack risk assessment dimensions. By fully utilizing the inherent noise term in the Bayesian time series decomposition model, this solution can naturally and efficiently estimate the probability distribution of the forecast results, avoiding additional complex uncertainty estimation modules, thereby simplifying the calculation process and ensuring the rationality of the probability distribution estimation. Calculating confidence intervals based on preset confidence levels allows for clear quantification of the risk range of the forecast results and flexible adaptation to different levels of flood control early warning needs. Outputting the confidence interval as the uncertainty interval provides decision-makers with an intuitive risk reference, enabling them to more comprehensively grasp the reliability of the forecast results. Therefore, in basin-level flood control decision-making, in addition to the predicted value, key risk dimension information is obtained, significantly enhancing the value and guiding significance of flood forecast results in practical applications.

[0097] Some schemes in this application also include: continuously monitoring the long-term stability and abrupt changes of the time lag correlation coefficient determined by the flow sequences of the upstream and downstream monitoring stations; determining whether the time lag correlation coefficient exhibits a continuous and significant abnormal pattern that cannot be explained by normal hydrological variations; and when the abnormal pattern is detected, triggering a topological change early warning signal and marking the river segment where a topological change may occur.

[0098] The following is a concrete example. Assume that in a certain watershed, there is a stable hydraulic connection between upstream monitoring station A and downstream monitoring station B. The time-lag correlation coefficient of their flow series typically fluctuates between 0.7 and 0.9, with a stable lag time of approximately 3 days. The proposed solution continuously monitors the time-lag correlation coefficient between these two stations. For example, the system can calculate the latest time-lag correlation coefficient daily and compare it with the moving average over the past 30 days. If monitoring reveals that within five consecutive days, the time-lag correlation coefficient between station A and station B suddenly drops below 0.2, and the lag time becomes uncertain or significantly prolonged, and historical data analysis confirms that the current period is not a dry or flood season, which would normally cause changes in correlation, then the system will determine that this is a persistent and significant anomalous pattern that cannot be explained by normal hydrological variability. In this case, the system will immediately trigger a topological change early warning signal, for example, displaying a red alert on the control panel of the flood warning platform and notifying on-duty personnel via SMS. At the same time, the system will highlight the river segment between station A and station B on the geographic information system (GIS) map and mark information such as "may be blocked" or "abnormal hydraulic connectivity" to mark the river segment that may have topological changes.

[0099] In some embodiments of this application, after triggering a topology change early warning signal, the method further includes: acquiring near-real-time remote sensing images and / or ground sensor data related to the marked river segment; identifying river network physical change information from the remote sensing images and / or ground sensor data through image processing and data change detection; and judging the identified river network physical change information by combining a historical river network topology change event database and a rule engine to determine the nature of the topology change.

[0100] As a specific implementation method, when the system detects a sudden and significant increase in the flood propagation time and a significant decrease in the correlation coefficient of the flow sequences determined by continuously monitoring upstream and downstream monitoring stations, thus triggering a topological change early warning signal and marking the river segment, this scheme will immediately initiate the verification and discrimination process. First, the system will automatically request the latest high-resolution optical satellite imagery and / or synthetic aperture radar (SAR) imagery of the marked river segment from the remote sensing data platform. Simultaneously, if IoT water level gauges or surface displacement sensors are deployed near the river segment, the system will also acquire near-real-time data from these sensors. Next, for the acquired optical satellite imagery, the system can use image segmentation algorithms based on convolutional neural networks (CNNs), such as the U-Net model, to accurately extract the river channel boundary and water area, and compare it with historical images of the river segment to identify morphological features such as whether the river channel width has changed significantly, whether there are new water areas, or whether existing water areas have disappeared. For SAR imagery, the system can leverage its ability to penetrate clouds and fog to detect minute surface deformations using interferometric SAR (InSAR) technology, thereby determining whether events such as landslides or foundation subsidence may affect the river channel. Simultaneously, for ground sensor data, the system can employ a cumulative summation (CUSUM) control chart algorithm to continuously monitor water level and surface displacement sequences. If a sustained abrupt change in these parameters exceeding normal fluctuation ranges is detected, it is considered an indication of physical change. Finally, the system inputs this identified information on river network physical changes, such as "disappearance of water in a section of the river channel," "significant surface subsidence," and "abnormal rise in upstream water level but no significant change in downstream water level," into a pre-defined rule engine. The rule engine can include rules such as: "If remote sensing imagery shows the disappearance of water in the river channel and InSAR detects signs of landslides in the upstream area, it is determined as 'complete river channel blockage';" and "If remote sensing imagery shows the emergence of a new water channel beside the river and downstream flow sensor data shows an abnormal increase, it is determined as 'formation of a new flood diversion channel'." In addition, the system will query a historical river network topology change event database to find historical events similar to the currently identified features to aid in the judgment. Through these steps, the system can ultimately clearly determine the nature of the topological change in the marked river segment, such as determining it as "the river channel is completely blocked due to a landslide," thus providing a precise basis for the subsequent dynamic adjustment of model parameters.

[0101] Through the aforementioned technical solution, this application can accurately identify genuine river network physical change information from these multi-source data, effectively eliminating false anomalies caused by normal hydrological variations. Finally, by combining a historical river network topology change event database with a rule engine, the identified physical change information is intelligently judged, clearly determining the specific nature of the topology change, such as river channel blockage, diversion, or the formation of a new flood diversion channel. This clear judgment provides an accurate and reliable basis for the subsequent dynamic adjustment of the adjacency matrix and symmetric graph Laplace matrix in the Bayesian time-series decomposition model, enabling the model to adapt promptly to the changed river network topology. This ensures the accuracy and physical consistency of flood prediction results when the river network physical structure changes, avoiding prediction bias caused by discrepancies between the model's understanding of the river network topology and the actual physical conditions.

[0102] Furthermore, some embodiments of this application, after determining the nature of the topological change, further include: dynamically modifying the connection relationships of corresponding upstream and downstream monitoring stations in the adjacency matrix according to the determined nature of the topological change, and recalculating the symmetric graph Laplace matrix based on the modified adjacency matrix; passing the recalculated symmetric graph Laplace matrix to the Bayesian time-series decomposition model; using a variational inference algorithm, reapplying topological smoothing constraints to the spatial weight matrix based on the recalculated symmetric graph Laplace matrix, and recalibrating the model parameters to adapt to the changed river network topology.

[0103] As a specific implementation, suppose that monitoring stations A and B in a certain watershed originally had a direct hydraulic connectivity, represented as non-zero values ​​in the adjacency matrix. However, through analysis of remote sensing imagery and verification of ground sensor data, it is determined that at a certain point in time, a local debris flow caused complete blockage of the river channel between monitoring stations A and B, resulting in a topological change property of "complete river blockage." At this point, the system dynamically modifies the connection relationships between monitoring stations A and B in the adjacency matrix based on this determined property. For example, it sets the values ​​of elements representing A to B and B to A in the adjacency matrix to zero, thereby reflecting the interruption of physical connectivity. Subsequently, based on this modified adjacency matrix, the system recalculates the symmetric graph Laplace matrix. This new symmetric graph Laplace matrix no longer contains connection information between monitoring stations A and B. Then, this recalculated symmetric graph Laplace matrix is ​​passed to the Bayesian temporal decomposition model. Inside the model, a variational inference algorithm is activated, using this new symmetric graph Laplace matrix to reapply topological smoothing constraints. This means that when optimizing the spatial weight matrix, the model will no longer attempt to force a smooth transition between monitoring stations A and B, but will instead learn based on the new topology. Simultaneously, the model will use variational inference algorithms to recalibrate all model parameters, including trend terms, seasonal terms, and covariate weight matrices. For example, if flow changes at monitoring station A no longer directly affect monitoring station B, the model will adjust the weights corresponding to A and B in its spatial weight matrix, and may also adjust other parameters to better fit the new data pattern. Through this series of operations, the Bayesian time series decomposition model can quickly adapt to the physical change of river blockage, avoiding predictions based on incorrect topology, thereby ensuring the accuracy and physical consistency of flood prediction results.

[0104] This scheme addresses the problem of traditional models failing to dynamically adapt to changes in the river network's topology, leading to decreased prediction accuracy and lack of physical consistency. By dynamically modifying the adjacency matrix and recalculating the Laplace matrix of the symmetric graph after determining the nature of river network topological changes, and then passing this information to a Bayesian time-series decomposition model, variational inference algorithms are used to reapply topological smoothing constraints to the spatial weight matrix and recalibrate the model parameters. Through this mechanism, the model maintains consistency between its internal topological representation and the actual river network connectivity, ensuring accurate and reliable flood prediction results even under extreme events such as river blockage or diversion. This fundamentally solves the prediction bias problem caused by changes in physical structure, providing more timely, accurate, and physically accurate data for flood control decisions.

[0105] The following supplementary explanation is provided using a specific scenario: The monitoring stations upstream of Station A and downstream of Station B in the X River Basin are directly hydraulically connected. External climate-driven data: using the Southern Oscillation Index (SOI); Forecast objective: Predict Bilibili's average daily traffic tomorrow.

[0106] In this scenario, this application includes the following steps: First, historical data is retrieved from the database and merged into the spatiotemporal input data shown in Table 1 below: June 1 A 10 100 +5.2 June 1 B 5 250 +5.2 June 2 A 20 150 +3.1 June 2 B 15 280 +3.1 June 3 A 25 180 +1.8 June 3 B 20 320 +1.8 ... ... ... ... ... Table 1 The spatiotemporal input data contains information from both the time dimension and the spatial dimension (multiple sites).

[0107] Secondly, an adjacency matrix is ​​constructed based on the river network topology to generate a symmetric graph Laplace matrix.

[0108] In this step, the river network topology is very simple: A → B (A is upstream, B is downstream).

[0109] The adjacency matrix is ​​shown in Table 2: A 0 1 B 1 0 Table 2 In Table 2, 0 indicates no connection and 1 indicates direct connection.

[0110] Based on this adjacency matrix, the system generates a symmetric graph Laplacian matrix L through mathematical calculation (normalization after subtracting the adjacency matrix from the degree matrix). This L matrix is ​​not a feature input, but a "physical rule constraint operator." It internally encodes the information that "A and B are connected," which is used during subsequent training to tell the model: you must keep the spatial weights of A and B similar.

[0111] Next, a Bayesian temporal decomposition model was constructed and trained.

[0112] The system constructs the FloodBayOTIDE model, which contains: Trend indicator: Capture long-term trends in traffic increase or decrease; Seasonal items: Capturing patterns during flood season / dry season; Covariate weight matrix: Learn how the SOI index affects traffic; Spatial weight matrix: Learn the spatial relationship between site A and site B; Noise term: Quantification of prediction uncertainty.

[0113] The training process is as follows: the model reads the spatiotemporal input data and begins training. During each round of parameter updates, two things happen simultaneously: Variational inference update: The model uses a variational inference algorithm to update all parameters (including the spatial weight matrix) based on the data fit.

[0114] Topological smoothing constraint application: The model uses the Laplacian matrix of the symmetric graph to construct a constraint term \( \text{trace}(W^TLW) \) and adds it to the loss function. The effect of this constraint term is to force the spatial weight vector of station A to be similar to that of station B. Because the L matrix has already told the model that "A and B are connected", if the model allows their weights to differ greatly, it will be penalized.

[0115] Training results: After training, the model learned to make the spatial weights of station A and station B highly similar (e.g., the cosine similarity reaches 0.95) while maintaining prediction accuracy.

[0116] Next, input the external climate driving data for the time to be predicted, and output the prediction results and uncertainty range.

[0117] This step is to predict the traffic on Bilibili tomorrow (June 4th).

[0118] Input: Simply feed tomorrow's predicted SOI index (e.g., +0.5) into the trained model.

[0119] Output: As shown in Table 3: Bilibili's average daily traffic tomorrow 310 m³ / s [275, 350] m³ / s Table 3 The predicted value of 310 m³ / s is the flow rate that the model considers most likely to occur.

[0120] The uncertainty interval [275, 350] m³ / s indicates that there is a 90% probability that the actual flow rate will fall within this range, which is estimated by the noise term in the model.

[0121] Subsequently, the time lag correlation coefficient was determined based on the flow sequences of upstream and downstream monitoring stations.

[0122] The system does not rely on a model; it directly performs statistical analysis on the historical traffic sequences of stations A and B. The traffic sequence of station A is progressively shifted forward (1 hour, 2 hours, 3 hours...), and the correlation coefficient with the traffic sequence of station B is calculated after each shift.

[0123] The analysis results are shown in Table 4: A site Bilibili 6 hours 0.92 Table 4 As shown in Table 4, the flow change at station A leads that at station B by about 6 hours, and the correlation is extremely strong. This is the result of the delayed propagation analysis, which reflects the actual physical laws of flood propagation.

[0124] Finally, we verify whether spatial weight similarity and time lag correlation coefficient corroborate each other.

[0125] This step extracts the following data: Spatial relationship: Extract the spatial weight vectors of station A and station B from the trained model, and calculate the similarity = 0.95; Time relationship: The lag analysis results of station A to station B are obtained from Table 4. The optimal lag period is 6 hours and the peak correlation coefficient is 0.92. Based on spatial relationships, the model suggests that A and B have a strong spatial connection; based on temporal relationships, the data also shows that A and B have a stable propagation relationship over a period of 6 hours.

[0126] Verification conclusion: Both spatial and temporal dimensions corroborate each other, pointing to the same physical fact: "It takes approximately 6 hours for the floodwater to flow from A to B." Therefore, the prediction result for station B (310 m³ / s) in Table 3 has passed the physical consistency verification, demonstrating high hydrological rationality and reliability.

[0127] According to the second aspect of this application, such as Figure 3 As shown in the illustration, this application also provides a flood forecasting device, which includes: Data module 301 is used to generate spatiotemporal input data based on rainfall and flow data from multiple monitoring stations, as well as external climate-driven data; Matrix module 302 is used to construct an adjacency matrix based on the river network topology among the multiple monitoring stations and generate a symmetric graph Laplace matrix; Training module 303 is used to construct and train a Bayesian temporal decomposition model based on the spatiotemporal input data and the symmetric graph Laplace matrix. The Bayesian temporal decomposition model includes a trend term, a seasonal term, a covariate weight matrix for fusing the external climate-driven data, a spatial weight matrix for capturing the spatial correlation between the monitoring stations, and a noise term. The symmetric graph Laplace matrix is ​​used to simultaneously apply topological smoothing constraints to the spatial weight matrix to force the spatial weights of adjacent monitoring stations with direct hydraulic connectivity in the river network to remain smooth. The output module 304 is used to input the external climate driving data of the time to be predicted into the trained Bayesian time series decomposition model and output the prediction results of the flow data and the corresponding uncertainty interval. Analysis module 305 is used to determine the time lag correlation coefficient based on the flow sequence of upstream and downstream monitoring stations, identify hydrological propagation time and data anomalies, and output the lag propagation analysis results. The verification module 306 is used to verify whether the similarity of the weight vectors of adjacent monitoring stations in the trained spatial weight matrix and the time lag correlation coefficient corroborate each other, so as to verify the physical consistency of the prediction results of the traffic data.

[0128] The flood forecasting device has all the beneficial effects of the above-mentioned methods, which will not be elaborated further in this application.

[0129] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0130] The embodiments, implementation methods, and related technical features of this application can be combined and substituted for each other without conflict.

[0131] The above are merely preferred embodiments of this application and are not intended to limit this application in any way. Although the descriptions of each embodiment in this application have different focuses, and parts not described in detail in a certain embodiment can be referred to the relevant descriptions of other embodiments, any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of this application without departing from the content of the technical solution of this application shall still fall within the scope of the technical solution of this application.

Claims

1. A flood prediction method, characterized in that, include: Spatiotemporal input data are generated based on rainfall and flow data from multiple monitoring stations, as well as external climate-driven data. An adjacency matrix is ​​constructed based on the river network topology among the multiple monitoring stations, and a symmetric graph Laplace matrix is ​​generated. Construct and train a Bayesian temporal decomposition model based on the spatiotemporal input data and the symmetric graph Laplacian matrix; The Bayesian time-series decomposition model includes a trend term, a seasonal term, a covariate weight matrix for fusing the external climate-driven data, a spatial weight matrix for capturing the spatial correlation between the monitoring stations, and a noise term; the symmetric graph Laplace matrix is ​​used to simultaneously apply topological smoothing constraints to the spatial weight matrix to force the spatial weights of adjacent monitoring stations with direct hydraulic connectivity in the river network to remain smooth. The external climate-driven data for the time to be predicted is input into the trained Bayesian time series decomposition model, which outputs the prediction results of the flow data and the corresponding uncertainty interval. Based on the flow sequences of upstream and downstream monitoring stations, the time lag correlation coefficient is determined, hydrological propagation time and data anomalies are identified, and the lag propagation analysis results are output. The similarity between the weight vectors of adjacent monitoring stations in the spatial weight matrix after training and the time lag correlation coefficient are verified to confirm each other, so as to verify the physical consistency of the prediction results of the traffic data.

2. The method according to claim 1, characterized in that, After training the Bayesian temporal decomposition model, the following steps are also included: The watersheds to which the multiple monitoring stations belong are divided into training watersheds and test watersheds; By using a masking method, the update of model parameters depends only on the monitoring station data of the training watershed; The cross-basin generalization ability of the Bayesian time-series decomposition model is evaluated based on monitoring station data from the test watershed.

3. The method according to claim 1, characterized in that, The step of training a Bayesian temporal decomposition model based on the spatiotemporal input data and the symmetric graph Laplacian matrix includes: Calculate the gradients of the Laplacian matrix of the symmetric graph and the spatial weight matrix; During model training, the spatial weight matrix is ​​iteratively corrected based on the gradient; Through the iterative correction, the spatial continuity of spatial weights of adjacent monitoring stations is maintained in areas with high flow variation.

4. The method according to claim 1, characterized in that, The verification of whether the similarity of the weight vectors of adjacent monitoring stations in the spatial weight matrix after training corroborates the time lag correlation coefficient includes: Determine whether the similarity of the weight vectors of the adjacent monitoring stations is greater than a preset similarity threshold; Determine whether the peak value of the time lag correlation coefficient is greater than a preset correlation coefficient threshold; When the similarity is greater than the preset similarity threshold and the peak value is greater than the preset correlation coefficient threshold, it is determined that the similarity and the time lag correlation coefficient mutually confirm each other, and the prediction result is verified by physical consistency.

5. The method according to claim 1, characterized in that, The step of training a Bayesian temporal decomposition model based on spatiotemporal input data and the Laplacian matrix of a symmetric graph includes: In each iteration of the online update of the model parameters using the variational inference algorithm, the following operations are performed simultaneously: Based on the current spatial weight matrix and the symmetric graph Laplace matrix, a topology smoothing constraint term is calculated, which is used to measure the difference between the spatial weight vectors of adjacent monitoring stations with direct hydraulic connectivity. The topological smoothing constraint term is added as a regularization term to the loss function of model training, and the gradient of the topological smoothing constraint term with respect to the spatial weight matrix is ​​calculated. The spatial weight matrix is ​​modified according to the gradient to force the spatial weights of adjacent monitoring stations to gradually become smoother during the iterative update process.

6. The method according to claim 1, characterized in that, The output includes the prediction results of the traffic data and the corresponding uncertainty interval, including: Based on the noise term in the Bayesian time series decomposition model, the probability distribution of the prediction results is estimated; Based on the probability distribution, calculate the confidence interval at the preset confidence level; The confidence interval is output as the uncertainty interval to quantify the risk range of the prediction result.

7. The method according to any one of claims 1 to 6, characterized in that, Also includes: The long-term stability and abrupt changes of the time lag correlation coefficient determined by continuously monitoring the flow sequences of the upstream and downstream monitoring stations; Determine whether the time lag correlation coefficient exhibits a persistent and significant anomalous pattern that cannot be explained by normal hydrological variability; When the abnormal pattern is detected, a topology change warning signal is triggered, and the river segment where a topology change may occur is marked.

8. The method according to claim 7, characterized in that, After triggering the topology change warning signal, it also includes: Acquire near real-time remote sensing images and / or ground sensor data related to the marked river segment; Information on physical changes in the river network is identified from the remote sensing images and / or ground sensor data through image processing and data change detection; By combining a historical river network topology change event database with a rule engine, the identified physical change information of the river network is judged to determine the nature of the topology change.

9. The method according to claim 8, characterized in that, After determining the properties of the topological change, the following is also included: Based on the determined topological change properties, the connection relationships of the corresponding upstream and downstream monitoring stations in the adjacency matrix are dynamically modified, and the Laplace matrix of the symmetric graph is recalculated based on the modified adjacency matrix. The recalculated symmetric graph Laplacian matrix is ​​passed to the Bayesian temporal decomposition model; Using a variational inference algorithm, topological smoothing constraints are reapplied to the spatial weight matrix based on the recalculated Laplacian matrix of the symmetric graph, and the model parameters are recalibrated to adapt to the changed river network topology.

10. A flood prediction device, characterized in that, include: The data module is used to generate spatiotemporal input data based on rainfall and flow data from multiple monitoring stations, as well as external climate-driven data. The matrix module is used to construct an adjacency matrix based on the river network topology among the multiple monitoring stations and generate a symmetric graph Laplace matrix. The training module is used to construct and train a Bayesian temporal decomposition model based on the spatiotemporal input data and the symmetric graph Laplacian matrix; The Bayesian time-series decomposition model includes a trend term, a seasonal term, a covariate weight matrix for fusing the external climate-driven data, a spatial weight matrix for capturing the spatial correlation between the monitoring stations, and a noise term; the symmetric graph Laplace matrix is ​​used to simultaneously apply topological smoothing constraints to the spatial weight matrix to force the spatial weights of adjacent monitoring stations with direct hydraulic connectivity in the river network to remain smooth. The output module is used to input the external climate-driven data of the time to be predicted into the trained Bayesian time series decomposition model and output the prediction results of the flow data and the corresponding uncertainty interval. The analysis module is used to determine the time lag correlation coefficient based on the flow sequence of upstream and downstream monitoring stations, identify hydrological propagation time and data anomalies, and output the lag propagation analysis results. The verification module is used to verify whether the similarity of the weight vectors of adjacent monitoring stations in the trained spatial weight matrix and the time lag correlation coefficient corroborate each other, so as to verify the physical consistency of the prediction results of the traffic data.