Urban flood prediction and risk assessment method based on multivariate data fusion
Through multivariate data fusion and random forest model, combined with Bayesian method and Monte Carlo simulation technology, the problem of single data sources and insufficient processing of uncertainty in the existing technology is solved, and high-precision flood prediction and risk assessment are achieved, and urban flood prevention and disaster reduction decisions are supported.
Patent Information
- Application Number
- CN202510769296.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-06-10
AI Technical Summary
In the prior art, urban flood prediction and risk assessment methods rely on a single data source, resulting in low prediction accuracy, incomplete risk assessment and insufficient multi-source uncertainty processing, making it difficult to meet high-precision and real-time requirements.
The multivariate data fusion method is adopted to build a random forest model and risk assessment index system through the integration of meteorological, hydrological, geographic information and socio-economic data, and the uncertainty processing is carried out in combination with Bayesian method and Monte Carlo simulation technology, and dynamic evaluation is carried out using real-time monitoring and geographic information systems.
It significantly improves the accuracy of flood prediction and comprehensiveness of risk assessment, provides scientific decision-making support and timely flood prevention and disaster reduction information, and enhances the reliability and stability of assessment results.
Smart Images

Figure CN120296584B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of urban flood prediction and risk assessment, and in particular to an urban flood prediction and risk assessment method based on multivariate data fusion. Background Art
[0002] Traditional flood forecasting and risk assessment methods primarily rely on single data sources, such as meteorological or hydrological data. These methods often face limitations in accuracy and reliability when faced with complex and changing climate and environmental conditions. Furthermore, existing methods struggle to handle multi-source uncertainty, making it difficult to comprehensively assess flood risk. This results in low reliability of forecasts and an inability to provide effective decision support for flood prevention and disaster reduction.
[0003] Existing technologies have significant shortcomings in data fusion. Traditional data processing methods are often unable to effectively integrate multi-source data such as meteorological, hydrological, geographic, and socioeconomic data, resulting in data redundancy and information loss. Furthermore, existing technologies also have flaws in model construction and uncertainty handling. A single prediction model struggles to capture the complex nonlinear relationships underlying flood occurrence, while inadequate quantification and handling of uncertainty further reduces the reliability of assessment results. These limitations make it difficult for existing technologies to meet the high-precision requirements of urban flood prediction and risk assessment in practical applications.
[0004] Existing risk assessment technologies often lack a comprehensive indicator system and fail to fully consider multiple dimensions of flood risk, including hazard, exposure, vulnerability, and resilience. This results in biased risk assessments that fail to fully reflect the impact of floods on urban socioeconomics and infrastructure. Furthermore, existing technologies also lack dynamic assessment and real-time early warning capabilities, making it difficult to update risk assessment results in a timely manner and meeting the real-time needs of urban flood prevention and disaster reduction.
[0005] In order to overcome the shortcomings of existing technologies, an urban flood prediction and risk assessment method based on multivariate data fusion is needed. Summary of the Invention
[0006] The purpose of this invention is to solve the problems in the prior art of single data source, low prediction accuracy, incomplete risk assessment indicators and insufficient multi-source uncertainty processing, and to propose an urban flood prediction and risk assessment method based on multivariate data fusion.
[0007] To achieve the above objectives, the present invention adopts the following technical solution: a method for urban flood prediction and risk assessment based on multivariate data fusion, comprising the following steps:
[0008] Step S1, collecting data from multiple data sources, including meteorological data, hydrological data, geographic information data, and socioeconomic data;
[0009] Step S2: pre-process the collected data and perform dimensionality reduction and fusion on the pre-processed data using data fusion technology;
[0010] Step S3, constructing a flood prediction model based on random forest to predict the probability of flood occurrence;
[0011] Step S4, constructing a risk assessment indicator system, including hazard indicators, exposure indicators, vulnerability indicators and resilience indicators;
[0012] Step S5, using the analytic hierarchy process and random forest method to determine the weight of each indicator and build a comprehensive risk assessment model;
[0013] Step S6, using real-time monitoring data combined with geographic information system technology to achieve dynamic assessment and visualization of flood risks;
[0014] In step S7, the Bayesian method and Monte Carlo simulation technology are used to quantify and process the multi-source uncertainty.
[0015] Furthermore, in step S1, the following sub-steps are also included:
[0016] S1-1, obtaining meteorological data through weather stations, satellite remote sensing and weather radar, wherein the meteorological data includes rainfall, rainfall intensity, temperature and humidity data;
[0017] S1-2, obtaining hydrological data through hydrological monitoring stations, water level sensors and groundwater monitoring wells, wherein the hydrological data includes river water level, flow and groundwater level data;
[0018] S1-3, obtaining geographic information data through digital elevation models, aerial photogrammetry, and satellite remote sensing images, wherein the geographic information data includes terrain and land use type data;
[0019] S1-4, obtaining socioeconomic data through population census, mobile communication data, urban planning data and municipal engineering data, wherein the socioeconomic data includes population distribution, building distribution and infrastructure distribution data.
[0020] Furthermore, in step S2, the following sub-steps are also included:
[0021] S2-1, cleaning and processing the collected data, including removing duplicates, removing erroneous records, removing irrelevant records, filling missing values and processing outliers;
[0022] S2-2, standardize the preprocessed data, normalize the mean of each feature to 0, and normalize the standard deviation to 1;
[0023] S2-3, calculate the covariance matrix of the standardized data, extract the eigenvalues and eigenvectors, select the first k eigenvectors with the largest eigenvalues as the principal components, project the original data onto the principal components, and achieve data dimensionality reduction. The specific formula of the covariance matrix is:
[0024]
[0025] in, represents the covariance matrix, is the i-th sample vector, is the sample mean vector, n is the number of samples, and T represents the transpose operation of the matrix;
[0026] S2-4, using a temporal interpolation method and a spatial interpolation method to synchronize data of different temporal resolutions and spatial resolutions onto a unified temporal and spatial grid, wherein the temporal interpolation method includes linear interpolation and polynomial interpolation, and the spatial interpolation method includes Kriging interpolation and nearest neighbor interpolation;
[0027] S2-5, integrating meteorological data, hydrological data, geographic information data and socio-economic data through data fusion technology to form a unified data set, wherein the data fusion technology includes weighted averaging method, Bayesian fusion and machine learning fusion.
[0028] Furthermore, in step S3, the following sub-steps are also included:
[0029] S3-1, determining parameters of the random forest model, wherein the parameters include the number of trees, the depth of the trees, and the minimum number of samples for splitting a node;
[0030] S3-2, divide the data into training set and validation set, use the training set data to train the random forest model, the training process of the random forest model is expressed as:
[0031]
[0032] in, is the predicted value, Q is the number of trees, q represents the qth decision tree in the random forest, a represents the feature vector input to the random forest model, is the predicted value of the qth tree;
[0033] S3-3, using k-fold cross validation, divide the training data into k subsets, where each subset is used as a validation set in turn, and the remaining data is used as a training set, and multiple training and validation are performed. The cross validation formula is:
[0034]
[0035] Among them, CV is the cross-validation result, k is the number of folds, is the mean square error of the b-th fold;
[0036] S3-4, based on the results of cross-validation, select the model parameter combination for the flood prediction model.
[0037] Furthermore, in step S4, the following sub-steps are also included:
[0038] S4-1: Determine risk indicators for the likelihood and intensity of floods, including flood inundation frequency, average annual rainy season rainfall, and average annual number of rainy days;
[0039] S4-2, determine exposure indicators for assessing the number of people and properties that may be affected by flooding, including population density and building distribution;
[0040] S4-3, determine vulnerability indicators for assessing the extent of potential losses incurred by floods, including the flood resistance of buildings and residents' awareness of flood prevention;
[0041] S4-4. Identify resilience indicators for assessing the speed and extent of post-flood recovery, including socioeconomic resilience.
[0042] Furthermore, in step S5, the following sub-steps are also included:
[0043] S5-1, through expert scoring and consistency test, use the hierarchical analysis method to construct a judgment matrix and determine the subjective weight of each indicator. The specific formula is:
[0044]
[0045] Among them, W represents the calculated weight, m is the number of indicators, is the maximum eigenvalue, Represents the element in the u-th row and v-th column of the judgment matrix, which indicates the importance of the u-th indicator relative to the v-th indicator. It represents the sum of all elements j in the vth column of the judgment matrix, that is, the total importance of the vth indicator;
[0046] S5-2, through random forest model training and feature importance analysis, determine the objective weight of each indicator. The specific formula is:
[0047]
[0048] in, is the importance of the oth feature, Q is the number of trees, is the mean square error change of the o-th feature on the q-th tree;
[0049] S5-3, combining subjective weights and objective weights, constructing a comprehensive risk assessment model to assess flood risk. The formula of the comprehensive risk assessment model is:
[0050]
[0051] Among them, R is the comprehensive risk, m is the number of indicators, is the weight of the g-th indicator, is the value of the g-th indicator.
[0052] Furthermore, in step S6, the following sub-steps are also included:
[0053] S6-1, obtain real-time meteorological, hydrological and geographic information data through IoT sensors and real-time data transmission technology;
[0054] S6-2: Use geographic information system technology to combine flood risk data with geographic information to perform geospatial visualization, generate risk maps, and display risk distribution. The geospatial visualization formula is:
[0055]
[0056] Among them, Risk Map is the risk map, Location is the geographical location, and R is the comprehensive risk, which is calculated by the comprehensive risk assessment model;
[0057] S6-3, based on real-time data, updates the input of the comprehensive risk assessment model and recalculates the risk value to achieve dynamic monitoring and assessment. The dynamic assessment process is expressed as:
[0058]
[0059] in, is the risk value at time s, m is the number of indicators, is the weight of the g-th indicator, is the value of the gth indicator at time s.
[0060] Furthermore, in step S7, the following sub-steps are also included:
[0061] S7-1. Identify and categorize sources of uncertainty that affect flood prediction and risk assessment, including data uncertainty, model uncertainty, and prediction uncertainty;
[0062] S7-2, using Bayesian theorem, quantify the uncertainty of model parameters and update the posterior distribution of model parameters. The specific formula is:
[0063]
[0064] in, is the posterior probability, is the likelihood function, is the prior probability, is the marginal probability of the data;
[0065] S7-3, randomly sample a large number of parameters from the posterior distribution, use Monte Carlo simulation techniques to simulate the propagation of uncertainty in the model, and evaluate the uncertainty distribution of the model output. The specific formula is:
[0066]
[0067] Where U is the uncertainty distribution, N is the number of samples, is the model output;
[0068] S7-4, based on the uncertainty analysis results, reduce the impact of uncertainty on the evaluation results by increasing the amount of data.
[0069] The beneficial effects brought about by the technical solution provided by the present invention include at least:
[0070] The present invention effectively improves the integrity and accuracy of data by integrating multiple data sources. The random forest prediction model constructed on this basis can capture the complex nonlinear relationship of flood occurrence and significantly improve the accuracy of flood prediction.
[0071] The present invention constructs a comprehensive risk assessment indicator system covering multiple dimensions. The constructed comprehensive risk assessment model can comprehensively assess flood risks. The risk assessment results are more comprehensive and accurate, and can provide more scientific decision-making support for urban flood prevention and disaster reduction.
[0072] The present invention utilizes real-time monitoring data and geographic information system technology to achieve dynamic assessment and visual display of flood risks, can provide real-time dynamic early warning information, promptly reflect changes in flood risks, and provide more timely decision-making support for flood prevention and disaster reduction.
[0073] The present invention adopts Bayesian method and Monte Carlo simulation technology to quantify and process multi-source uncertainty, which can effectively reduce the impact of uncertainty on the evaluation results and improve the reliability and stability of the evaluation results. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0075] Figure 1 A flowchart of a method provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0076] To further illustrate the technical means and effectiveness of the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, describes in detail the specific implementation, structure, features, and effectiveness of the urban flood prediction and risk assessment method based on multivariate data fusion proposed by the present invention. In the following description, different references to "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics of one or more embodiments may be combined in any suitable manner.
[0077] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0078] The following examples are for illustrative purposes only and are not intended to limit the scope of the present invention.
[0079] The specific scheme of the urban flood prediction and risk assessment method based on multivariate data fusion provided by the present invention is described in detail below with reference to the accompanying drawings.
[0080] See also Figure 1 , which shows a method flow chart of a method for urban flood prediction and risk assessment based on multivariate data fusion provided by an embodiment of the present invention, the method comprising the following steps:
[0081] Step S1, collecting data from multiple data sources, including meteorological data, hydrological data, geographic information data, and socioeconomic data;
[0082] Wherein, step S1 further includes the following sub-steps:
[0083] S1-1, meteorological data are obtained through weather stations, satellite remote sensing and weather radar. The meteorological data include rainfall, rainfall intensity, temperature and humidity data;
[0084] S1-2, obtain hydrological data through hydrological monitoring stations, water level sensors and groundwater monitoring wells. The hydrological data include river water level, flow and groundwater level data;
[0085] S1-3, obtain geographic information data through digital elevation models, aerial photogrammetry and satellite remote sensing images, and the geographic information data includes terrain and land use type data;
[0086] S1-4, obtain socioeconomic data through population census, mobile communication data, urban planning data and municipal engineering data. Socioeconomic data include population distribution, building distribution and infrastructure distribution data.
[0087] It should be noted that weather stations are the main source of ground-based meteorological data, located in cities and surrounding areas. They can provide high-temporal-resolution meteorological data, which helps to understand rainfall patterns and climate change.
[0088] Satellite remote sensing: It can provide meteorological data over a wide range, with wide coverage and high data acquisition frequency. It is particularly important for supplementing the data from ground meteorological stations, especially in areas where meteorological stations are sparsely distributed.
[0089] Meteorological radar: It can monitor the intensity and distribution of rainfall in real time and provide rainfall data with high spatial resolution. It is critical for monitoring and early warning of short-term heavy rainfall and can help to promptly detect rainfall events that may cause floods.
[0090] Precipitation: The total amount of rainfall per unit time, usually measured in millimeters (mm).
[0091] Rainfall intensity: The rate of rainfall per unit time, usually expressed in millimeters per hour (mm / h).
[0092] Air temperature: The temperature of the atmosphere, usually measured in degrees Celsius (°C).
[0093] Humidity: The amount of water vapor in the air, usually expressed as relative humidity (%) or absolute humidity (g / m³).
[0094] Hydrological monitoring stations: installed near water bodies such as rivers and lakes, they monitor water levels and flow rates and can provide high-precision real-time hydrological data.
[0095] Water level sensor: installed in rivers and lakes to monitor water level changes in real time and transmit data to the data center via wireless network to ensure the timeliness and accuracy of the data.
[0096] Groundwater monitoring wells: used to monitor changes in groundwater levels, which are important for assessing groundwater recharge and discharge. Changes in groundwater levels can affect river flow and flood formation.
[0097] River level: The height of the water level in a river, usually measured in metres (m).
[0098] Flow rate: The amount of water passing through a certain section per unit time, usually measured in cubic meters per second (m³ / s).
[0099] Groundwater level: The height of groundwater, usually measured in metres (m).
[0100] Digital Elevation Model (DEM): A digital model that represents terrain elevation and can provide high-precision terrain information. DEM data is generated through aerial photogrammetry or satellite remote sensing technology and is used for flood inundation simulation and risk assessment.
[0101] Aerial photogrammetry: Aerial photography is used to obtain high-resolution topographic and land use information, which can provide detailed topographic and land use data.
[0102] Satellite remote sensing images: They can provide large-scale geographic information and are critical for assessing the impact of floods on different land use types.
[0103] Terrain: Elevation and slope information of the terrain, usually represented in the form of a Digital Elevation Model (DEM).
[0104] Land use type: land use classification, including urban, farmland, forest, and water, obtained through satellite remote sensing images and aerial photogrammetry.
[0105] Population census: provides detailed information on resident distribution, including population size, density, and age distribution, which helps to assess the impact of floods on residents.
[0106] Mobile communication data: can provide real-time population distribution and mobility information, especially the dynamic changes in population during floods, obtained through signal records from mobile communication base stations.
[0107] Urban planning data: including building distribution and land use planning information, helps to assess the impact of floods on urban infrastructure.
[0108] Municipal engineering data: including drainage system and flood control facility information, which helps to assess the city’s flood control capabilities.
[0109] Population distribution: The geographical distribution of residents, usually expressed as population density (people per square kilometer).
[0110] Building distribution: Information on the location and type of buildings, represented by geographic information system data.
[0111] Infrastructure distribution: includes the location and status information of roads, bridges, and drainage system infrastructure, represented by municipal engineering data and GIS data.
[0112] Step S2: pre-process the collected data and perform dimensionality reduction and fusion on the pre-processed data using data fusion technology;
[0113] Wherein, in step S2, the following sub-steps are also included:
[0114] S2-1, cleaning and processing the collected data, including removing duplicates, removing erroneous records, removing irrelevant records, filling missing values and processing outliers;
[0115] S2-2, standardize the preprocessed data, normalize the mean of each feature to 0, and normalize the standard deviation to 1;
[0116] S2-3, calculate the covariance matrix of the standardized data, extract the eigenvalues and eigenvectors, select the first k eigenvectors with the largest eigenvalues as the principal components, project the original data onto the principal components to achieve data dimensionality reduction. The specific formula of the covariance matrix is:
[0117]
[0118] in, represents the covariance matrix, is the i-th sample vector, is the sample mean vector, n is the number of samples, and T represents the transpose operation of the matrix;
[0119] S2-4, using temporal interpolation methods and spatial interpolation methods to synchronize data of different temporal and spatial resolutions onto a unified temporal and spatial grid. Temporal interpolation methods include linear interpolation and polynomial interpolation, and spatial interpolation methods include Kriging interpolation and nearest neighbor interpolation.
[0120] S2-5, integrate meteorological data, hydrological data, geographic information data and socio-economic data through data fusion technology to form a unified data set. Data fusion technologies include weighted averaging, Bayesian fusion and machine learning fusion.
[0121] It should be noted that duplicate records are removed: through data deduplication algorithms, duplicate records in the data set are identified and deleted to ensure the uniqueness of the data.
[0122] Remove erroneous records: Use data validation rules to check the rationality and accuracy of data, and delete or correct records with obvious errors.
[0123] Removal of irrelevant records: Based on the research objectives and needs, data related to flood prediction and risk assessment were screened out and irrelevant records were removed.
[0124] Fill missing values: Use interpolation method to fill missing values.
[0125] Handling outliers: Identify and handle outliers through statistical analysis methods to ensure the accuracy and reliability of data.
[0126] Data normalization is to adjust data with different characteristics to the same dimension and distribution range.
[0127] Principal component analysis (PCA) is a commonly used data dimensionality reduction method that reduces data dimensions by extracting the main features of the data.
[0128] Linear interpolation is a simple time series interpolation method used to estimate unknown data points between two known data points. It can be used when the data changes relatively slowly.
[0129] Polynomial interpolation is a more complex interpolation method that can better fit the variation of the data by constructing a polynomial function to estimate the unknown data points.
[0130] Kriging interpolation is a statistically based spatial interpolation method that is particularly suitable for processing spatial data. It assumes that the spatial correlation between data points can be described by a semivariogram model.
[0131] Nearest neighbor interpolation is a simple and intuitive spatial interpolation method that assigns the value of an unknown point to the value of the nearest known point.
[0132] The weighted average method is a simple and effective way to aggregate data by assigning weights to each data source and then calculating the weighted average. The weights are assigned based on the reliability and importance of the data source.
[0133] Bayesian fusion is based on Bayes' theorem and combines prior knowledge and new data to update data estimates. It can handle data uncertainty and incompleteness and is particularly suitable for scenarios where data sources are uncertain.
[0134] Machine learning fusion uses machine learning models (random forest, neural network) to fuse multi-source data, which can automatically learn the complex relationships between data sources and improve the accuracy and robustness of the fusion.
[0135] Step S3, constructing a flood prediction model based on random forest to predict the probability of flood occurrence;
[0136] Wherein, in step S3, the following sub-steps are also included:
[0137] S3-1, determine the parameters of the random forest model, including the number of trees, tree depth, and the minimum number of samples for splitting a node;
[0138] S3-2, divide the data into training set and validation set, use the training set data to train the random forest model, the training process of the random forest model is expressed as:
[0139]
[0140] in, is the predicted value, Q is the number of trees, q represents the qth decision tree in the random forest, a represents the feature vector input to the random forest model, is the predicted value of the qth tree;
[0141] S3-3, using k-fold cross validation, divide the training data into k subsets, where each subset is used as a validation set in turn, and the remaining data is used as a training set, and multiple training and validation are performed. The cross validation formula is:
[0142]
[0143] Among them, CV is the cross-validation result, k is the number of folds, is the mean square error of the b-th fold;
[0144] S3-4, based on the results of cross-validation, select the model parameter combination for the flood prediction model.
[0145] It should be noted that the number of trees: the number of decision trees in the random forest. Increasing the number of trees can improve the stability and prediction accuracy of the model, but at the same time it will increase the computational cost.
[0146] Tree depth: The maximum depth of each decision tree. Limiting the depth of the tree can prevent overfitting, but a tree that is too shallow may not be able to capture complex relationships in the data.
[0147] Minimum number of samples for splitting nodes: The minimum number of samples required to split internal nodes can control the growth of the tree and prevent overfitting.
[0148] Data partitioning: The dataset is divided into a training set and a validation set, with a ratio of 70% training set and 30% validation set.
[0149] K-fold cross validation is an effective method to evaluate model performance and ensure that the model performs well on unseen data.
[0150] The specific steps for selecting the optimal model parameter combination based on the cross-validation results include:
[0151] 1. Parameter selection: Select the optimal model parameter combination through grid search or random search method combined with cross-validation results.
[0152] 2. Model optimization: Retrain the model using the selected parameter combination to ensure that the model performs well on the validation set.
[0153] The following is the code for building a random forest-based flood prediction model:
[0154] import numpy as np
[0155] import pandas as pd
[0156] from sklearn.model_selection import train_test_split, GridSearchCV,cross_val_score
[0157] from sklearn.ensemble import RandomForestRegressor
[0158] from sklearn.metrics import mean_squared_error
[0159] # Assume the data has been loaded into a DataFrame
[0160] # data = pd.read_csv('flood_data.csv')
[0161] # Features and target variables
[0162] X = data.drop('flood_probability', axis=1)
[0163] y = data['flood_probability']
[0164] # Data partitioning
[0165] X_train, X_val, y_train, y_val = train_test_split(X, y, test_size=0.3, random_state=42)
[0166] # Random Forest Model
[0167] rf = RandomForestRegressor(random_state=42)
[0168] # Parameter grid
[0169] param_grid = {
[0170] 'n_estimators': [100, 200, 300],
[0171] 'max_depth': [10, 20, 30],
[0172] 'min_samples_split': [2, 5, 10]
[0173] }
[0174] # Grid Search
[0175] grid_search = GridSearchCV(estimator=rf, param_grid=param_grid, cv=5,scoring='neg_mean_squared_error')
[0176] grid_search.fit(X_train, y_train)
[0177] # Optimal parameters
[0178] best_params = grid_search.best_params_
[0179] print(f"Best parameters: {best_params}")
[0180] # Retrain the model using the optimal parameters
[0181] best_rf = RandomForestRegressor(**best_params, random_state=42)
[0182] best_rf.fit(X_train, y_train)
[0183] # predict
[0184] y_pred = best_rf.predict(X_val)
[0185] # Evaluate
[0186] mse = mean_squared_error(y_val, y_pred)
[0187] print(f"Mean Squared Error: {mse}")
[0188] # Cross-validation evaluation
[0189] cv_scores = cross_val_score(best_rf, X_train, y_train, cv=5, scoring='neg_mean_squared_error')
[0190] cv_mse = -np.mean(cv_scores)
[0191] print(f"Cross-validated Mean Squared Error: {cv_mse}")
[0192] Step S4, constructing a risk assessment indicator system, including hazard indicators, exposure indicators, vulnerability indicators and resilience indicators;
[0193] Wherein, in step S4, the following sub-steps are also included:
[0194] S4-1: Determine the risk indicators for the likelihood and intensity of floods, including flood inundation frequency, average annual rainy season rainfall, and average annual number of heavy rain days;
[0195] S4-2, determine exposure indicators for assessing the number of people and properties that may be affected by flooding, including population density and building distribution;
[0196] S4-3, determine vulnerability indicators for assessing the extent of potential losses incurred by floods, including the flood resistance of buildings and residents' awareness of flood prevention;
[0197] S4-4. Identify resilience indicators to assess the speed and extent of post-flood recovery, including socioeconomic resilience.
[0198] It should be noted that the flood inundation frequency: the ratio of the number of times an area is flooded within a specific period of time to the total time, reflects the frequency of floods in the area.
[0199] Average annual rainy season rainfall: the average rainfall during the rainy season (June to September) each year, reflecting the impact of rainfall on flood occurrence.
[0200] Average annual number of days with heavy rain: the number of days with daily rainfall exceeding 50 mm each year, reflecting the impact of heavy rain on flood occurrence.
[0201] Population density: The number of people per unit area, usually measured in persons per square kilometer, which reflects the population's exposure to floods.
[0202] Building distribution: The location and number of buildings, typically represented through geographic information system (GIS) data, reflects the exposure of buildings to flooding.
[0203] Building flood resilience: A building's ability to withstand floods, usually assessed by the building's structural type and flood-resistant design.
[0204] Residents’ awareness of flood prevention: Residents’ awareness of and ability to cope with floods, usually assessed through questionnaires or social survey data.
[0205] Socioeconomic resilience: The speed and extent to which socioeconomic activities resume after a flood, usually assessed through economic data and recovery time.
[0206] Step S5, using the analytic hierarchy process and random forest method to determine the weight of each indicator and build a comprehensive risk assessment model;
[0207] Wherein, in step S5, the following sub-steps are also included:
[0208] S5-1, through expert scoring and consistency test, use the hierarchical analysis method to construct a judgment matrix and determine the subjective weight of each indicator. The specific formula is:
[0209]
[0210] Among them, W represents the calculated weight, m is the number of indicators, is the maximum eigenvalue, Represents the element in the u-th row and v-th column of the judgment matrix, which indicates the importance of the u-th indicator relative to the v-th indicator. It represents the sum of all elements j in the vth column of the judgment matrix, that is, the total importance of the vth indicator;
[0211] S5-2, through random forest model training and feature importance analysis, determine the objective weight of each indicator. The specific formula is:
[0212]
[0213] in, is the importance of the oth feature, Q is the number of trees, is the mean square error change of the o-th feature on the q-th tree;
[0214] S5-3, combining subjective weights and objective weights, construct a comprehensive risk assessment model to assess flood risk. The formula of the comprehensive risk assessment model is:
[0215]
[0216] Among them, R is the comprehensive risk, m is the number of indicators, is the weight of the g-th indicator, is the value of the g-th indicator.
[0217] It should be noted that expert scoring is a key step in the Analytic Hierarchy Process (AHP), which is used to quantify experts' judgment on the relative importance of different indicators. The specific steps are as follows:
[0218] 1. Determine the evaluation indicators: First, determine the indicator system that needs to be evaluated, including hazard indicators, exposure indicators, vulnerability indicators and resilience indicators.
[0219] 2. Expert selection: Invite experts with relevant knowledge and experience to participate in the scoring. The experts should have an in-depth understanding of flood risk assessment.
[0220] 3. Pairwise comparison: Experts compare each pair of indicators and give a relative importance score.
[0221] 4. Construct a judgment matrix based on the experts’ scores.
[0222] Consistency test is an important step in the AHP process, which is used to verify the consistency of expert scores.
[0223] The analytic hierarchy process (AHP) is a multi-criteria decision-making method based on expert scoring and consistency testing, which is used to determine the weight of each indicator.
[0224] Model training: Random forest is an ensemble learning method that improves the accuracy and stability of predictions by building multiple decision trees.
[0225] Feature Importance Analysis: Random forests can provide importance scores for each feature, which reflect the contribution of the feature in the model.
[0226] The code for building a comprehensive risk assessment model by combining subjective and objective weights is:
[0227] # Assume that subjective weight and objective weight have been calculated
[0228] # W_AHP = [0.3, 0.2, 0.4, 0.1] # Example subjective weight
[0229] # W_RF = [0.2, 0.3, 0.4, 0.1] # Example objective weight
[0230] # Combining subjective and objective weights
[0231] alpha = 0.5# subjective weight distribution coefficient
[0232] W_combined = alpha * W_AHP + (1 - alpha) * W_RF
[0233] #Build a comprehensive risk assessment model
[0234] # Assume that X_val is the feature data of the validation set
[0235] R = np.dot(X_val, W_combined)
[0236] # Assessing overall risk
[0237] print(f"Comprehensive risk R: {R}")
[0238] Step S6, using real-time monitoring data combined with geographic information system technology to achieve dynamic assessment and visualization of flood risks;
[0239] Wherein, in step S6, the following sub-steps are also included:
[0240] S6-1, obtain real-time meteorological, hydrological and geographic information data through IoT sensors and real-time data transmission technology;
[0241] S6-2: Use geographic information system technology to combine flood risk data with geographic information to perform geospatial visualization, generate risk maps, and display risk distribution. The geospatial visualization formula is:
[0242]
[0243] Among them, Risk Map is the risk map, Location is the geographical location, and R is the comprehensive risk, which is calculated by the comprehensive risk assessment model;
[0244] S6-3, based on real-time data, updates the input of the comprehensive risk assessment model and recalculates the risk value to achieve dynamic monitoring and assessment. The dynamic assessment process is expressed as:
[0245]
[0246] in, is the risk value at time s, m is the number of indicators, is the weight of the g-th indicator, is the value of the gth indicator at time s.
[0247] It should be noted that IoT sensors: sensors deployed at key locations (rain gauges, water level gauges, weather stations) collect meteorological, hydrological and geographic information data in real time.
[0248] Real-time data transmission: The data collected by the sensor is transmitted to the data center in real time via wireless networks (4G / 5G, LoRa).
[0249] The GIS generates an intuitive risk map. The specific steps include:
[0250] 1. Data integration: Combine real-time monitoring data with geographic information data (elevation, land use, population distribution) to form a complete risk assessment data set.
[0251] 2. Geospatial visualization: Use GIS software (ArcGIS, QGIS) to generate risk maps to show risk distribution.
[0252] The code to implement dynamic monitoring and evaluation includes:
[0253] import numpy as np
[0254] import pandas as pd
[0255] import geopandas as gpd
[0256] import matplotlib.pyplot as plt
[0257] from sklearn.ensemble import RandomForestRegressor
[0258] # Assume that real-time data has been loaded into the DataFrame
[0259] # real_time_data = pd.read_csv('real_time_data.csv')
[0260] # Features and target variables
[0261] X_real_time = real_time_data.drop('flood_probability', axis=1)
[0262] y_real_time = real_time_data['flood_probability']
[0263] # Load geographic information data
[0264] # geo_data = gpd.read_file('geo_data.shp')
[0265] # Random Forest model initialization
[0266] rf = RandomForestRegressor(n_estimators=100, max_depth=20, random_state=42)
[0267] # Model training
[0268] rf.fit(X_real_time, y_real_time)
[0269] # Predict risk value
[0270] y_pred_real_time = rf.predict(X_real_time)
[0271] # Combine the prediction results with geographic information data
[0272] geo_data['risk'] = y_pred_real_time
[0273] Geospatial Visualization
[0274] fig, ax = plt.subplots(figsize=(10, 10))
[0275] geo_data.plot(column='risk', ax=ax, legend=True, cmap='viridis')
[0276] plt.title('Flood Risk Map')
[0277] plt.show()
[0278] Step S7, using Bayesian method and Monte Carlo simulation technology to quantify and process multi-source uncertainty;
[0279] Wherein, in step S7, the following sub-steps are also included:
[0280] S7-1. Identify and categorize sources of uncertainty that affect flood prediction and risk assessment, including data uncertainty, model uncertainty, and prediction uncertainty.
[0281] S7-2, using Bayesian theorem, quantify the uncertainty of model parameters and update the posterior distribution of model parameters. The specific formula is:
[0282]
[0283] in, is the posterior probability, is the likelihood function, is the prior probability, is the marginal probability of the data;
[0284] S7-3, randomly sample a large number of parameters from the posterior distribution, use Monte Carlo simulation techniques to simulate the propagation of uncertainty in the model, and evaluate the uncertainty distribution of the model output. The specific formula is:
[0285]
[0286] Where U is the uncertainty distribution, N is the number of samples, is the model output;
[0287] S7-4, based on the uncertainty analysis results, reduce the impact of uncertainty on the evaluation results by increasing the amount of data.
[0288] It should be noted that data uncertainty is the uncertainty caused by measurement errors, missing or incomplete data.
[0289] Model uncertainty: uncertainty due to model assumptions, parameter estimates, or the choice of model structure.
[0290] Forecast uncertainty: The uncertainty in forecast results caused by uncertainty in future conditions (meteorological conditions, human activities).
[0291] The specific steps of the Bayesian method include:
[0292] 1. Set prior distribution: Set the prior distribution of model parameters based on prior knowledge .
[0293] 2. Calculate the likelihood function: Calculate the likelihood function based on the observed data , which represents the probability of the data under given parameters.
[0294] 3. Update the posterior distribution: Update the posterior distribution of the parameters through Bayes' theorem .
[0295] The specific steps of the Monte Carlo simulation technique for uncertainty distribution include:
[0296] 1. Random sampling: randomly draw a large number of samples from the posterior distribution of the parameter.
[0297] 2. Uncertainty propagation: These samples are fed into the model to calculate the uncertainty distribution of the model output.
[0298] 3. Uncertainty assessment: Evaluate the uncertainty distribution of the model output through statistical analysis (calculation of mean, variance, confidence interval).
[0299] Increase the amount of data: By increasing the number of observations, the randomness and uncertainty of the data can be reduced.
[0300] The code to increase the amount of data to reduce uncertainty includes:
[0301] import numpy as np
[0302] import pandas as pd
[0303] from sklearn.ensemble import RandomForestRegressor
[0304] from sklearn.model_selection import train_test_split
[0305] from sklearn.metrics import mean_squared_error
[0306] # Assume that the initial data has been loaded into the DataFrame
[0307] # initial_data = pd.read_csv('initial_data.csv')
[0308] # Features and target variables
[0309] X_initial = initial_data.drop('flood_probability', axis=1)
[0310] y_initial = initial_data['flood_probability']
[0311] # Initial model training
[0312] rf = RandomForestRegressor(n_estimators=100, max_depth=20, random_state=42)
[0313] rf.fit(X_initial, y_initial)
[0314] # Assume that new data has been loaded into the DataFrame
[0315] # new_data = pd.read_csv('new_data.csv')
[0316] # Add new data to the initial data
[0317] X_new = new_data.drop('flood_probability', axis = 1)
[0318] y_new = new_data['flood_probability']
[0319] # Update the dataset
[0320] X_updated = pd.concat([X_initial, X_new], ignore_index = True)
[0321] y_updated = pd.concat([y_initial, y_new], ignore_index = True)
[0322] # Retrain the model
[0323] rf.fit(X_updated, y_updated)
[0324] # Evaluate the model performance
[0325] X_train, X_test, y_train, y_test = train_test_split(X_updated, y_updated, test_size = 0.3, random_state = 42)
[0326] rf.fit(X_train, y_train)
[0327] y_pred = rf.predict(X_test)
[0328] mse = mean_squared_error(y_test, y_pred)
[0329] print(f"Mean Squared Error after adding new data: {mse:.4f}")
[0330] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. Urban flood prediction and risk assessment method based on multivariate data fusion, characterized by: The method includes: Step S1, collecting data from multiple data sources, including meteorological data, hydrological data, geographic information data, and socioeconomic data; Step S2: pre-process the collected data and perform dimensionality reduction and fusion on the pre-processed data using data fusion technology; Step S3, constructing a flood prediction model based on random forest to predict the probability of flood occurrence; Step S4, constructing a risk assessment indicator system, including hazard indicators, exposure indicators, vulnerability indicators and resilience indicators; Step S5, using the analytic hierarchy process and random forest method to determine the weight of each indicator and build a comprehensive risk assessment model; Step S6, using real-time monitoring data combined with geographic information system technology to achieve dynamic assessment and visualization of flood risks; Step S7, using Bayesian method and Monte Carlo simulation technology to quantify and process multi-source uncertainty; Wherein, in step S5, the following sub-steps are also included: S5-1, through expert scoring and consistency test, use the hierarchical analysis method to construct a judgment matrix and determine the subjective weight of each indicator. The specific formula is: Among them, W represents the calculated weight, m is the number of indicators, is the maximum eigenvalue, Represents the element in the u-th row and v-th column of the judgment matrix, which indicates the importance of the u-th indicator relative to the v-th indicator. It represents the sum of all elements j in the vth column of the judgment matrix, that is, the total importance of the vth indicator; S5-2, through random forest model training and feature importance analysis, determine the objective weight of each indicator. The specific formula is: in, is the importance of the oth feature, Q is the number of trees, is the mean square error change of the o-th feature on the q-th tree; S5-3, combining subjective weights and objective weights, constructing a comprehensive risk assessment model to assess flood risk. The formula of the comprehensive risk assessment model is: Among them, R is the comprehensive risk, m is the number of indicators, is the comprehensive weight, is the value of the g-th indicator.
2. The urban flood prediction and risk assessment method based on multivariate data fusion according to claim 1 is characterized by: Wherein, step S1 further includes the following sub-steps: S1-1, obtaining meteorological data through weather stations, satellite remote sensing and weather radar, wherein the meteorological data includes rainfall, rainfall intensity, temperature and humidity data; S1-2, obtaining hydrological data through hydrological monitoring stations, water level sensors and groundwater monitoring wells, wherein the hydrological data includes river water level, flow and groundwater level data; S1-3, obtaining geographic information data through digital elevation models, aerial photogrammetry, and satellite remote sensing images, wherein the geographic information data includes terrain and land use type data; S1-4, obtaining socioeconomic data through population census, mobile communication data, urban planning data and municipal engineering data, wherein the socioeconomic data includes population distribution, building distribution and infrastructure distribution data.
3. The urban flood prediction and risk assessment method based on multivariate data fusion according to claim 1 is characterized by: Wherein, in step S2, the following sub-steps are also included: S2-1, cleaning and processing the collected data, including removing duplicates, removing erroneous records, removing irrelevant records, filling missing values and processing outliers; S2-2, standardize the preprocessed data, normalize the mean of each feature to 0, and normalize the standard deviation to 1; S2-3, calculate the covariance matrix of the standardized data, extract the eigenvalues and eigenvectors, select the first k eigenvectors with the largest eigenvalues as the principal components, project the original data onto the principal components, and achieve data dimensionality reduction. The specific formula of the covariance matrix is: in, represents the covariance matrix, is the i-th sample vector, is the sample mean vector, n is the number of samples, and T represents the transpose operation of the matrix S2-4, using a temporal interpolation method and a spatial interpolation method to synchronize data of different temporal resolutions and spatial resolutions onto a unified temporal and spatial grid, wherein the temporal interpolation method includes linear interpolation and polynomial interpolation, and the spatial interpolation method includes Kriging interpolation and nearest neighbor interpolation; S2-5, integrating meteorological data, hydrological data, geographic information data and socio-economic data through data fusion technology to form a unified data set, wherein the data fusion technology includes weighted averaging method, Bayesian fusion and machine learning fusion.
4. The urban flood prediction and risk assessment method based on multivariate data fusion according to claim 1 is characterized by: Wherein, in step S3, the following sub-steps are also included: S3-1, determining parameters of the random forest model, wherein the parameters include the number of trees, the depth of the trees, and the minimum number of samples for splitting a node; S3-2, divide the data into training set and validation set, use the training set data to train the random forest model, the training process of the random forest model is expressed as: in, is the predicted value, Q is the number of trees, q represents the qth decision tree in the random forest, a represents the feature vector input to the random forest model, is the predicted value of the qth tree; S3-3, using k-fold cross validation, divide the training data into k subsets, where each subset is used as a validation set in turn, and the remaining data is used as a training set, and multiple training and validation are performed. The cross validation formula is: Among them, CV is the cross-validation result, k is the fold number, is the mean square error of the b-th fold; S3-4, based on the results of cross-validation, select the model parameter combination for the flood prediction model.
5. The urban flood prediction and risk assessment method based on multivariate data fusion according to claim 1 is characterized by: Wherein, in step S4, the following sub-steps are also included: S4-1: Determine risk indicators for the likelihood and intensity of floods, including flood inundation frequency, average annual rainy season rainfall, and average annual number of rainy days; S4-2, determine exposure indicators for assessing the number of people and properties that may be affected by flooding, including population density and building distribution; S4-3, determine vulnerability indicators for assessing the extent of potential losses incurred by floods, including the flood resistance of buildings and residents' awareness of flood prevention; S4-4. Identify resilience indicators for assessing the speed and extent of post-flood recovery, including socioeconomic resilience.
6. The urban flood prediction and risk assessment method based on multivariate data fusion according to claim 1 is characterized by: Wherein, in step S6, the following sub-steps are also included: S6-1, obtain real-time meteorological, hydrological and geographic information data through IoT sensors and real-time data transmission technology; S6-2: Use geographic information system technology to combine flood risk data with geographic information to perform geospatial visualization, generate risk maps, and display risk distribution. The geospatial visualization formula is: Among them, Risk Map is the risk map, Location is the geographical location, and R is the comprehensive risk, which is calculated by the comprehensive risk assessment model; S6-3, based on real-time data, updates the input of the comprehensive risk assessment model and recalculates the risk value to achieve dynamic monitoring and assessment. The dynamic assessment process is expressed as: in, is the risk value at time s, is the weight of the g-th indicator, is the value of the gth indicator at time s.
7. The urban flood prediction and risk assessment method based on multivariate data fusion according to claim 1 is characterized by: Wherein, in step S7, the following sub-steps are also included: S7-1. Identify and categorize sources of uncertainty that affect flood prediction and risk assessment, including data uncertainty, model uncertainty, and prediction uncertainty; S7-2, using Bayesian theorem, quantify the uncertainty of model parameters and update the posterior distribution of model parameters. The specific formula is: in, is the posterior probability, is the likelihood function, is the prior probability, is the marginal probability of the data; S7-3, randomly sample a large number of parameters from the posterior distribution, use Monte Carlo simulation techniques to simulate the propagation of uncertainty in the model, and evaluate the uncertainty distribution of the model output. The specific formula is: Where U is the uncertainty distribution, N is the number of samples, is the model output; S7-4, based on the uncertainty analysis results, reduce the impact of uncertainty on the evaluation results by increasing the amount of data.
Citation Information
Patent Citations
Urban flood risk assessment method and platform based on multi-source data fusion
CN118886717A
Urban flood prediction method based on Bayesian convolutional neural network
CN119648078A
Rural waterlogging area waterlogging sheet risk grade evaluation and early warning method based on multi-source data
CN119920077A