Urban flood prediction and risk assessment method based on multivariate data fusion

Through multivariate data fusion and random forest model, combined with hierarchical analysis method and Bayesian method, the problem of single data and insufficient processing of uncertainty in urban flood prediction and risk assessment is solved, high-precision flood prediction and comprehensive risk assessment are achieved, and timely support for flood prevention and disaster reduction decisions are provided.

CN120296584AActive Publication Date: 2025-07-11XIANGJIANG LAB

Patent Information

Application Number
CN202510769296.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-11
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

The existing technology has problems such as single data source, low prediction accuracy, incomplete risk assessment indicators and insufficient multi-source uncertainty processing in urban flood forecasting and risk assessment, resulting in low accuracy and reliability of prediction results and cannot meet the needs of flood prevention and disaster reduction.

Method used

Multivariate data fusion method is adopted to construct a random forest model for flood prediction through the integration of meteorological, hydrological, geographical information and socio-economic data, and a comprehensive risk assessment index system is built, combining hierarchical analysis method and Bayesian method to deal with uncertainty, and dynamic assessment and visual display are achieved.

Benefits of technology

It improves the accuracy of flood forecasting and comprehensiveness of risk assessment, provides scientific decision-making support, can promptly reflect changes in flood risk, and enhances the reliability and stability of assessment results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296584A_ABST
    Figure CN120296584A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of urban flood prediction and risk assessment, and discloses an urban flood prediction and risk assessment method based on multivariate data fusion, and the method comprises the steps: collecting data through a plurality of data sources; preprocessing the collected data, and performing dimensionality reduction and fusion on the preprocessed data; a flood prediction model based on the random forest is constructed, and the flood occurrence probability is predicted; constructing a risk assessment index system; determining the weight of each index by using an analytic hierarchy process and a random forest method, and constructing a comprehensive risk assessment model; dynamic assessment and visual display of flood risks are realized by combining real-time monitoring data with a geographic information system technology; the multi-source uncertainty is quantized and processed by adopting a Bayesian method and a Monte Carlo simulation technology. According to the method, multiple data sources can be integrated, the flood risk is comprehensively evaluated, the multi-source uncertainty is effectively processed, and the reliability and stability of the evaluation result are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of urban flood prediction and risk assessment, and particularly to an urban flood prediction and risk assessment method based on multi-source data fusion. Background Art

[0002] Traditional flood prediction and risk assessment methods mainly rely on single data sources, such as meteorological data or hydrological data. When facing complex and changeable climate and environmental conditions, the accuracy and reliability of these methods are often limited. In addition, existing methods have deficiencies in dealing with multi-source uncertainties, making it difficult to comprehensively evaluate flood risks, resulting in low reliability of prediction results and unable to provide effective decision-making support for flood control and disaster reduction.

[0003] There are obvious deficiencies in existing technologies in data fusion. Traditional data processing methods usually cannot effectively integrate multi-source data such as meteorological, hydrological, geographic information, and social economy, resulting in data redundancy and information loss. In addition, existing technologies also have defects in model construction and uncertainty processing. A single prediction model is difficult to capture the complex non-linear relationships of flood occurrence, and the insufficient quantification and processing of uncertainties further reduce the reliability of evaluation results. These limitations make it difficult for existing technologies to meet the high-precision requirements of urban flood prediction and risk assessment in practical applications.

[0004] In terms of risk assessment, existing technologies often lack a comprehensive index system and fail to fully consider multiple dimensions such as the hazard, exposure, vulnerability, and recoverability of flood occurrence. This leads to one-sided risk assessment results and cannot comprehensively reflect the impact of floods on urban social economy and infrastructure. In addition, existing technologies also have deficiencies in dynamic assessment and real-time warning, unable to update risk assessment results in a timely manner and difficult to meet the real-time requirements of urban flood control and disaster reduction.

[0005] In order to overcome the deficiencies of existing technologies, an urban flood prediction and risk assessment method based on multi-source data fusion is needed. Summary of the Invention

[0006] The purpose of the present invention is to propose an urban flood prediction and risk assessment method based on multi-source data fusion to solve the problems of single data source, low prediction accuracy, incomplete risk assessment indicators, and insufficient processing of multi-source uncertainties in existing technologies.

[0007] To achieve the above purpose, the present invention adopts the following technical solutions: An urban flood prediction and risk assessment method based on multi-source data fusion, including the following steps: Step S1, collect data through multiple data sources, including meteorological data, hydrological data, geographic information data, and social economy data; Step S2, preprocess the collected data, and reduce the dimension and fuse the preprocessed data through data fusion technology; Step S3, construct a flood prediction model based on random forest to predict the flood occurrence probability; Step S4, construct a risk assessment index system, including hazard index, exposure index, vulnerability index and resilience index; Step S5, use the analytic hierarchy process and random forest method to determine the weights of each index, and construct a comprehensive risk assessment model; Step S6, through real-time monitoring data combined with geographic information system technology, realize the dynamic assessment and visual display of flood risk; Step S7, use Bayesian method and Monte Carlo simulation technology to quantify and process multi-source uncertainties.

[0008] Furthermore, in step S1, the following sub-steps are also included: S1-1, obtain meteorological data through weather stations, satellite remote sensing and weather radars, and the meteorological data includes rainfall, rainfall intensity, temperature and humidity data; S1-2, obtain hydrological data through hydrological monitoring stations, water level sensors and groundwater monitoring wells, and the hydrological data includes river water level, flow and groundwater level data; S1-3, obtain geographic information data through digital elevation models, aerial photogrammetry and satellite remote sensing images, and the geographic information data includes terrain and land use type data; S1-4, obtain socio-economic data through population census, mobile communication data, urban planning data and municipal engineering data, and the socio-economic data includes population distribution, building distribution and infrastructure distribution data.

[0009] Furthermore, in step S2, the following sub-steps are also included: S2-1, clean and process the collected data, including removing duplicates, removing error records, removing irrelevant records, filling missing values and handling outliers; S2-2, perform standardization processing on the preprocessed data, normalizing the mean of each feature to 0 and the standard deviation to 1; S2-3, calculate the covariance matrix of the standardized data, extract eigenvalues and eigenvectors, select the top k eigenvectors with the largest eigenvalues as the principal components, project the original data onto the principal components, and realize data dimension reduction. The specific formula of the covariance matrix is: Where, represents the covariance matrix, is the i-th sample vector, is the sample mean vector, n is the number of samples, and T represents the transpose operation of the matrix; S2-4, Using time interpolation methods and spatial interpolation methods, synchronize data with different time resolutions and spatial resolutions to a unified time and space grid. The time interpolation methods include linear interpolation and polynomial interpolation, and the spatial interpolation methods include Kriging interpolation and nearest neighbor interpolation; S2-5, Integrate meteorological data, hydrological data, geographical information data, and socioeconomic data through data fusion technology to form a unified dataset. The data fusion technology includes weighted average method, Bayesian fusion, and machine learning fusion.

[0010] Further, in step S3, the following sub-steps are also included: S3-1, Determine the parameters of the random forest model. The parameters include the number of trees, the depth of the trees, and the minimum number of samples for splitting nodes; S3-2, Divide the data into a training set and a validation set, and use the training set data to train the random forest model. The training process of the random forest model is expressed as: where, is the predicted value, Q is the number of trees, q represents the q-th decision tree in the random forest, and a represents the feature vector input into the random forest model, is the predicted value of the q-th tree; S3-3, Use k-fold cross-validation. Divide the training data into k subsets, where each subset is used as the validation set in turn, and the remaining data is used as the training set for multiple training and validation. The formula for cross-validation is: where, CV is the cross-validation result, k is the number of folds, is the mean squared error of the b-th fold; S3-4, According to the results of cross-validation, select the model parameter combination for the flood prediction model.

[0011] Further, in step S4, the following sub-steps are also included: S4-1, Determine the hazard indicators for the likelihood and intensity of flood occurrence. The hazard indicators include flood inundation frequency, average annual rainy season rainfall, and average annual number of heavy rain days; S4-2, Determine the exposure indicators for evaluating the number of people and property that may be affected by floods. The exposure indicators include population density and building distribution; S4-3, Determine the vulnerability indicators for evaluating the possible loss extent under the action of floods. The vulnerability indicators include building flood resistance ability and residents' flood prevention awareness; S4-4. Determine the resilience indicators for evaluating the speed and extent of post-flood recovery, where the resilience indicators include the socio-economic recovery ability.

[0012] Furthermore, in step S5, the following sub-steps are also included: S5-1. Construct a judgment matrix using the analytic hierarchy process through expert scoring and consistency testing to determine the subjective weights of each indicator. The specific formula is: where W represents the calculated weight, m is the number of indicators, is the maximum eigenvalue, represents the element in the u-th row and v-th column of the judgment matrix, indicating the importance of the u-th indicator relative to the v-th indicator, represents the sum of all elements j in the v-th column of the judgment matrix, that is, the total importance of the v-th indicator; S5-2. Determine the objective weights of each indicator through random forest model training and feature importance analysis. The specific formula is: where, is the importance of the o-th feature, Q is the number of trees, is the change in mean squared error of the o-th feature on the q-th tree; S5-3. Combine the subjective weight and the objective weight to construct a comprehensive risk assessment model to evaluate the flood risk. The formula of the comprehensive risk assessment model is: where R is the comprehensive risk, m is the number of indicators, is the weight of the g-th indicator, is the value of the g-th indicator.

[0013] Furthermore, in step S6, the following sub-steps are also included: S6-1. Obtain real-time meteorological, hydrological, and geographical information data through Internet of Things sensors and real-time data transmission technology; S6-2. Combine the flood risk data with geographical information using geographical information system technology for geospatial visualization to generate a risk map to display the risk distribution. The formula for geospatial visualization is: where Risk Map is the risk map, Location is the geographical location, R is the comprehensive risk, which is calculated through the comprehensive risk assessment model; S6-3. Update the input of the comprehensive risk assessment model based on real-time data, recalculate the risk value, and achieve dynamic monitoring and assessment. The dynamic assessment process is expressed as follows: Wherein, is the risk value at time s, m is the number of indicators, is the weight of the g-th indicator, is the value of the g-th indicator at time s.

[0014] Furthermore, in step S7, the following sub-steps are further included: S7-1. Identify and classify the sources of uncertainty affecting flood prediction and risk assessment. The uncertainties include data uncertainty, model uncertainty, and prediction uncertainty; S7-2. Quantify the uncertainty of model parameters through Bayes' theorem and update the posterior distribution of model parameters. The specific formula is: Wherein, is the posterior probability, is the likelihood function, is the prior probability, is the marginal probability of the data; S7-3. Randomly sample a large number of samples from the posterior distribution of parameters, use Monte Carlo simulation technology to simulate the propagation of uncertainty in the model, and evaluate the uncertainty distribution of the model output. The specific formula is: Wherein, U is the uncertainty distribution, N is the number of samples, is the model output; S7-4. According to the uncertainty analysis results, reduce the impact of uncertainty on the evaluation results by increasing the data volume.

[0015] The beneficial effects brought by the technical solution provided by the present invention at least include: By integrating multiple data sources, the present invention effectively improves the integrity and accuracy of data. On this basis, the constructed random forest prediction model can capture the complex non-linear relationship of flood occurrence, and significantly improves the accuracy of flood prediction.

[0016] The present invention constructs a comprehensive risk assessment index system covering multiple dimensions. The constructed comprehensive risk assessment model can comprehensively evaluate flood risks, and the risk assessment results are more comprehensive and accurate, and can provide more scientific decision-making support for urban flood control and disaster reduction.

[0017] The present invention utilizes real-time monitoring data and geographic information system technology to achieve dynamic assessment and visual display of flood risks, capable of providing real-time and dynamic early warning information, promptly reflecting changes in flood risks, and providing more timely decision-making support for flood prevention and mitigation.

[0018] The present invention adopts Bayesian methods and Monte Carlo simulation techniques to quantify and process multi-source uncertainties, which can effectively reduce the impact of uncertainties on assessment results and improve the reliability and stability of assessment results. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0020] Figure 1 It is a flowchart of the method provided for an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] To further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following, in conjunction with the drawings and preferred embodiments, details the specific implementation manners, structures, features, and effects of the urban flood prediction and risk assessment method based on multi-source data fusion proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0023] The following embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention.

[0024] The following specifically describes the specific solution of the urban flood prediction and risk assessment method based on multi-source data fusion provided by the present invention with reference to the drawings.

[0025] Please refer to Figure 1 , which shows a flowchart of the urban flood prediction and risk assessment method based on multi-source data fusion provided by an embodiment of the present invention. The method includes the following steps: Step S1, collect data through multiple data sources, including meteorological data, hydrological data, geographic information data, and socioeconomic data; Among them, in step S1, the following sub-steps are further included: S1-1. Obtain meteorological data through weather stations, satellite remote sensing, and weather radars. The meteorological data includes rainfall, rainfall intensity, air temperature, and humidity data; S1-2. Obtain hydrological data through hydrological monitoring stations, water level sensors, and groundwater monitoring wells. The hydrological data includes river water level, flow rate, and groundwater level data; S1-3. Obtain geographical information data through digital elevation models, aerial photogrammetry, and satellite remote sensing images. The geographical information data includes terrain and land use type data; S1-4. Obtain socio-economic data through population censuses, mobile communication data, urban planning data, and municipal engineering data. The socio-economic data includes population distribution, building distribution, and infrastructure distribution data.

[0026] It should be noted that weather stations: The main source for obtaining surface meteorological data, distributed in urban and surrounding areas, can provide meteorological data with high temporal resolution, which helps to understand rainfall patterns and climate change.

[0027] Satellite remote sensing: Can provide meteorological data over a large range, with a wide coverage area and a high data acquisition frequency, which is particularly important for supplementing the data of ground weather stations, especially in areas where weather stations are sparsely distributed.

[0028] Weather radar: Can monitor the intensity and distribution of rainfall in real time, provide rainfall data with high spatial resolution, which is very crucial for the monitoring and early warning of short-term heavy rainfall, and can help to detect rainfall events that may trigger floods in a timely manner.

[0029] Rainfall: The total amount of rainfall within a unit time, usually measured in millimeters (mm).

[0030] Rainfall intensity: The rate of rainfall within a unit time, usually measured in millimeters per hour (mm / h).

[0031] Air temperature: The temperature of the atmosphere, usually measured in degrees Celsius (°C).

[0032] Humidity: The content of water vapor in the air, usually measured in relative humidity (%) or absolute humidity (g / m³).

[0033] Hydrological monitoring stations: Installed near the water bodies of rivers and lakes, used to monitor hydrological parameters such as water level and flow rate, and can provide high-precision real-time hydrological data.

[0034] Water level sensors: Installed in the water bodies of rivers and lakes, used to monitor the water level changes in real time, and transmit the data to the data center through a wireless network to ensure the timeliness and accuracy of the data.

[0035] Groundwater monitoring well: Used to monitor the changes in groundwater levels, which is very important for evaluating groundwater recharge and discharge. Changes in groundwater levels can affect river flow and the formation of floods.

[0036] River water level: The height of the river water, usually measured in meters (m).

[0037] Discharge: The amount of water passing through a certain cross-section per unit time, usually measured in cubic meters per second (m³ / s).

[0038] Groundwater level: The height of the groundwater, usually measured in meters (m).

[0039] Digital Elevation Model (DEM): A digital model representing terrain elevation that can provide high-precision terrain information. DEM data is generated through aerial photogrammetry or satellite remote sensing technology and is used for simulating flood inundation areas and risk assessment.

[0040] Aerial photogrammetry: Obtaining high-resolution terrain and land use information through aerial photography, which can provide detailed terrain and land use data.

[0041] Satellite remote sensing images: Can provide large-scale geographical information, which is crucial for evaluating the impact of floods on different land use types.

[0042] Terrain: Information on terrain elevation and slope, usually represented in the form of a Digital Elevation Model (DEM).

[0043] Land use types: Classification of land uses, including urban, agricultural, forest, and water areas, obtained through satellite remote sensing images and aerial photogrammetry.

[0044] Census: Provides detailed information on the distribution of residents, including population quantity, density, and age distribution, which helps to evaluate the impact of floods on residents.

[0045] Mobile communication data: Can provide real-time information on population distribution and movement, especially the dynamic changes in population during floods, obtained through signal records of mobile communication base stations.

[0046] Urban planning data: Includes building distribution and land use planning information, which helps to evaluate the impact of floods on urban infrastructure.

[0047] Municipal engineering data: Includes drainage systems and flood control facility information, which helps to evaluate the flood control capacity of the city.

[0048] Population distribution: The geographical distribution of residents, usually represented by population density (persons per square kilometer).

[0049] Building distribution: The location and type information of buildings are represented by geographic information system data.

[0050] Infrastructure distribution: Includes the location and status information of infrastructure such as roads, bridges, and drainage systems, represented by municipal engineering data and GIS data.

[0051] Step S2: Preprocess the collected data, and reduce the dimension and fuse the preprocessed data through data fusion technology. Among them, in step S2, the following sub-steps are also included: S2-1: Clean and process the collected data, including removing duplicates, removing error records, removing irrelevant records, filling in missing values, and handling outliers. S2-2: Standardize the preprocessed data, normalizing the mean of each feature to 0 and the standard deviation to 1. S2-3: Calculate the covariance matrix of the standardized data, extract eigenvalues and eigenvectors, select the top k eigenvectors with the largest eigenvalues as the principal components, and project the original data onto the principal components to achieve data dimensionality reduction. The specific formula for the covariance matrix is: Among them, represents the covariance matrix, is the i-th sample vector, is the sample mean vector, n is the number of samples, and T represents the matrix transpose operation; S2-4: Use time interpolation methods and space interpolation methods to synchronize data with different time resolutions and space resolutions to a unified time and space grid. Time interpolation methods include linear interpolation and polynomial interpolation, and space interpolation methods include Kriging interpolation and nearest neighbor interpolation. S2-5: Integrate meteorological data, hydrological data, geographic information data, and socioeconomic data through data fusion technology to form a unified dataset. Data fusion technologies include weighted average method, Bayesian fusion, and machine learning fusion.

[0052] It should be noted that removing duplicate records: Identify and delete duplicate records in the dataset through a data deduplication algorithm to ensure data uniqueness.

[0053] Removing error records: Check the reasonableness and accuracy of the data through data verification rules, and delete or correct obviously wrong records.

[0054] Removing irrelevant records: According to the research objectives and requirements, screen out data related to flood prediction and risk assessment, and remove irrelevant records.

[0055] Filling in missing values: Use interpolation methods to fill in missing values.

[0056] Handling outliers: Identify and handle outliers through statistical analysis methods to ensure the accuracy and reliability of the data.

[0057] Data standardization is to adjust data with different features to the same dimension and distribution range.

[0058] Principal Component Analysis (PCA) is a commonly used data dimensionality reduction method that reduces the data dimension by extracting the main features of the data.

[0059] Linear interpolation is a simple time - series interpolation method used to estimate unknown data points between two known data points and is applicable to situations where data changes are relatively smooth.

[0060] Polynomial interpolation is a more complex interpolation method that can better fit the data changes by constructing a polynomial function to estimate unknown data points.

[0061] Kriging interpolation is a spatial interpolation method based on statistics, especially suitable for processing spatial data, assuming that the spatial correlation between data points can be described by a semivariogram model.

[0062] Nearest - neighbor interpolation is a simple and intuitive spatial interpolation method that assigns the value of the unknown point as the value of the nearest known point.

[0063] The weighted average method is a simple and effective method that integrates data by assigning weights to each data source and then calculating the weighted average. The assignment of weights is based on the reliability and importance of the data sources.

[0064] Bayesian fusion is based on Bayes' theorem, combines prior knowledge and new data, updates the data estimate, can handle data uncertainty and incompleteness, and is especially suitable for scenarios where data sources have uncertainty.

[0065] Machine - learning fusion uses machine - learning models (random forest, neural network) to fuse multi - source data, can automatically learn the complex relationships between data sources, and improve the accuracy and robustness of fusion.

[0066] Step S3, construct a flood prediction model based on a random forest to predict the probability of flood occurrence; Among them, in step S3, the following sub - steps are also included: S3 - 1, determine the parameters of the random forest model. The parameters include the number of trees, the depth of the trees, and the minimum number of samples for splitting nodes; S3 - 2, divide the data into a training set and a validation set, and use the training set data to train the random forest model. The training process of the random forest model is expressed as: Among them, is the predicted value, Q is the number of trees, q represents the q-th decision tree in the random forest, and a represents the feature vector input into the random forest model. is the predicted value of the q-th tree; S3-3. Use k-fold cross-validation to divide the training data into k subsets. Among them, each subset is used as the validation set in turn, and the remaining data is used as the training set for multiple training and validations. The formula for cross-validation is: Among them, CV is the cross-validation result, and k is the number of folds. is the mean squared error of the b-th fold; S3-4. According to the results of cross-validation, select the model parameter combination for the flood prediction model.

[0067] It should be noted that the number of trees: the number of decision trees in the random forest. Increasing the number of trees can improve the stability and prediction accuracy of the model, but at the same time, it will increase the computational cost.

[0068] The depth of the tree: the maximum depth of each decision tree. Limiting the depth of the tree can prevent overfitting, but a tree that is too shallow may not be able to capture the complex relationships in the data.

[0069] The minimum number of samples for splitting nodes: the minimum number of samples required to split internal nodes, which can control the growth of the tree and prevent overfitting.

[0070] Data partitioning: Divide the dataset into a training set and a validation set, with a ratio of 70% for the training set and 30% for the validation set.

[0071] k-fold cross-validation is an effective method for evaluating the performance of the model, which can ensure that the model performs well on unseen data.

[0072] The specific steps for selecting the optimal model parameter combination according to the results of cross-validation include: 1. Parameter selection: Through grid search or random search methods, combined with the cross-validation results, select the optimal model parameter combination.

[0073] 2. Model optimization: Retrain the model using the selected parameter combination to ensure that the model performs well on the validation set.

[0074] The following is the code for building a flood prediction model based on random forest: import numpy as np import pandas as pd from sklearn.model_selection import train_test_split, GridSearchCV,cross_val_score from sklearn.ensemble import RandomForestRegressor from sklearn.metrics import mean_squared_error # Assume the data has been loaded into a DataFrame # data = pd.read_csv('flood_data.csv') # Features and target variable X = data.drop('flood_probability', axis=1) y = data['flood_probability'] # Data splitting X_train, X_val, y_train, y_val = train_test_split(X, y, test_size=0.3, random_state=42) # Random Forest model rf = RandomForestRegressor(random_state=42) # Parameter grid param_grid = { 'n_estimators': [100, 200, 300], 'max_depth': [10, 20, 30], 'min_samples_split': [2, 5, 10] } # Grid search grid_search = GridSearchCV(estimator=rf, param_grid=param_grid, cv=5,scoring='neg_mean_squared_error') grid_search.fit(X_train, y_train) # Optimal parameters best_params = grid_search.best_params_ print(f"Best parameters: {best_params}") # Retrain the model with the optimal parameters best_rf = RandomForestRegressor(**best_params, random_state=42) best_rf.fit(X_train, y_train) # Make predictions y_pred = best_rf.predict(X_val) # Evaluate mse = mean_squared_error(y_val, y_pred) print(f"Mean Squared Error: {mse}") # Cross-validation evaluation cv_scores = cross_val_score(best_rf, X_train, y_train, cv=5, scoring='neg_mean_squared_error') cv_mse = -np.mean(cv_scores) print(f"Cross-validated Mean Squared Error: {cv_mse}") Step S4, construct a risk assessment index system, including hazard indicators, exposure indicators, vulnerability indicators and resilience indicators; Among them, in step S4, the following sub-steps are also included: S4-1, determine the hazard indicators for the likelihood and intensity of flood occurrence. The hazard indicators include flood inundation frequency, annual average rainy season rainfall and annual average number of heavy rain days; S4-2, determine the exposure indicators for evaluating the number of people and property that may be affected by floods. The exposure indicators include population density and building distribution; S4-3, determine the vulnerability indicators for evaluating the possible loss degree under the action of floods. The vulnerability indicators include the flood resistance ability of buildings and the flood prevention awareness of residents; S4-4, determine the restoration indicators for evaluating the speed and degree of post-flood restoration, where the restoration indicators include the socio-economic restoration capacity.

[0075] It should be noted that the flood inundation frequency: the ratio of the number of times a certain area is inundated by floods within a specific time to the total time, which reflects the frequency of flood occurrence in this area.

[0076] Annual average rainfall during the rainy season: the average rainfall during the rainy season (from June to September) each year, which reflects the impact of rainfall on flood occurrence.

[0077] Annual average number of heavy rain days: the number of days with daily rainfall exceeding 50 mm each year, which reflects the impact of heavy rain on flood occurrence.

[0078] Population density: the number of people per unit area, usually measured in people per square kilometer, which reflects the degree of population exposure to floods.

[0079] Building distribution: the location and number of buildings, usually represented by geographical information system (GIS) data, which reflects the degree of building exposure to floods.

[0080] Building flood resistance capacity: the disaster resistance ability of buildings in floods, usually evaluated by the structural type and flood prevention design of buildings.

[0081] Residents' flood prevention awareness: the awareness and response ability of residents to floods, usually evaluated through questionnaire surveys or social survey data.

[0082] Socio-economic restoration capacity: the speed and degree of the restoration of socio-economic activities after floods, usually evaluated by economic data and restoration time.

[0083] Step S5, use the analytic hierarchy process and random forest method to determine the weights of each indicator and construct a comprehensive risk assessment model; Among them, in step S5, the following sub-steps are also included: S5-1, through expert scoring and consistency test, use the analytic hierarchy process to construct a judgment matrix and determine the subjective weights of each indicator. The specific formula is: Among them, W represents the calculated weight, m is the number of indicators, is the maximum eigenvalue, represents the element in the u-th row and v-th column of the judgment matrix, indicating the importance of the u-th indicator relative to the v-th indicator, represents the sum of all elements j in the v-th column of the judgment matrix, that is, the total importance of the v-th indicator; S5-2, through random forest model training and feature importance analysis, determine the objective weights of each indicator. The specific formula is: Among them, is the importance of the o-th feature, Q is the number of trees, is the change in mean squared error of the o-th feature on the q-th tree; S5-3. Combining subjective weight and objective weight, construct a comprehensive risk assessment model to evaluate flood risk. The formula of the comprehensive risk assessment model is: Among them, R is the comprehensive risk, m is the number of indicators, is the weight of the g-th indicator, is the value of the g-th indicator.

[0084] It should be noted that the expert scoring is a key step in the Analytic Hierarchy Process (AHP), which is used to quantify the experts' judgment on the relative importance of different indicators. The specific steps are as follows: 1. Determine evaluation indicators: First, determine the indicator system to be evaluated, including hazard indicators, exposure indicators, vulnerability indicators, and resilience indicators.

[0085] 2. Select experts: Invite experts with relevant field knowledge and experience to participate in scoring. The experts should have an in-depth understanding of flood risk assessment.

[0086] 3. Pairwise comparison: Experts compare each pair of indicators and give a relative importance score.

[0087] 4. According to the experts' scores, construct a judgment matrix.

[0088] Consistency test is an important step in the Analytic Hierarchy Process (AHP), which is used to verify the consistency of expert scoring.

[0089] The Analytic Hierarchy Process (AHP) is a multi-criteria decision-making method based on expert scoring and consistency test, which is used to determine the weights of each indicator.

[0090] Model training: Random forest is an ensemble learning method that improves the accuracy and stability of prediction by constructing multiple decision trees.

[0091] Feature importance analysis: Random forest can provide the importance scores of each feature, and these scores reflect the contribution of the feature in the model.

[0092] The code for constructing a comprehensive risk assessment model by combining subjective weight and objective weight is: # Assume that the subjective weight and objective weight have been calculated # W_AHP = [0.3, 0.2, 0.4, 0.1] # Example subjective weight # W_RF = [0.2, 0.3, 0.4, 0.1] # Example objective weights # Combine subjective and objective weights alpha = 0.5 # Distribution coefficient of subjective weight W_combined = alpha * W_AHP + (1 - alpha) * W_RF # Build a comprehensive risk assessment model # Assume X_val is the feature data of the validation set R = np.dot(X_val, W_combined) # Evaluate the comprehensive risk print(f"Comprehensive risk R: {R}") Step S6, through real-time monitoring data combined with geographic information system technology, realize the dynamic assessment and visual display of flood risk; Among them, in step S6, the following sub-steps are also included: S6-1, through Internet of Things sensors and real-time data transmission technology, obtain real-time meteorological, hydrological and geographic information data; S6-2, use geographic information system technology to combine flood risk data with geographic information, conduct geospatial visualization, generate a risk map, and display the risk distribution. The geospatial visualization formula is: Among them, Risk Map is the risk map, Location is the geographical location, and R is the comprehensive risk, which is calculated through the comprehensive risk assessment model; S6-3, update the input of the comprehensive risk assessment model according to real-time data, recalculate the risk value, and realize dynamic monitoring and assessment. The dynamic assessment process is expressed as: Among them, is the risk value at time s, m is the number of indicators, is the weight of the g-th indicator, is the value of the g-th indicator at time s.

[0093] It should be noted that Internet of Things sensors: Sensors (rain gauges, water level gauges, weather stations) deployed at key locations collect real-time meteorological, hydrological and geographic information data.

[0094] Real-time data transmission: Transmit the data collected by the sensors to the data center in real time through wireless networks (4G / 5G, LoRa).

[0095] The Geographic Information System generates intuitive risk maps, and the specific steps are as follows: 1. Data integration: Combine real-time monitoring data with geographic information data (elevation, land use, population distribution) to form a complete risk assessment dataset.

[0096] 2. Geospatial visualization: Use GIS software (ArcGIS, QGIS) to generate risk maps and display the risk distribution.

[0097] The code for realizing dynamic monitoring and assessment includes: import numpy as np import pandas as pd import geopandas as gpd import matplotlib.pyplot as plt from sklearn.ensemble import RandomForestRegressor # Assume that the real-time data has been loaded into a DataFrame # real_time_data = pd.read_csv('real_time_data.csv') # Features and target variables X_real_time = real_time_data.drop('flood_probability', axis=1) y_real_time = real_time_data['flood_probability'] # Load geographic information data # geo_data = gpd.read_file('geo_data.shp') # Initialize the random forest model rf = RandomForestRegressor(n_estimators=100, max_depth=20, random_state=42) # Model training rf.fit(X_real_time, y_real_time) # Predict risk values y_pred_real_time = rf.predict(X_real_time) # Combine the prediction results with the geographical information data geo_data['risk'] = y_pred_real_time # Geospatial visualization fig, ax = plt.subplots(figsize=(10, 10)) geo_data.plot(column='risk', ax=ax, legend=True, cmap='viridis') plt.title('Flood Risk Map') plt.show() Step S7: Quantify and process multi-source uncertainties using Bayesian methods and Monte Carlo simulation techniques; In step S7, the following sub-steps are also included: S7-1: Identify and classify the sources of uncertainty affecting flood prediction and risk assessment. Uncertainties include data uncertainty, model uncertainty, and prediction uncertainty; S7-2: Quantify the uncertainty of model parameters through Bayes' theorem and update the posterior distribution of model parameters. The specific formula is: Where, is the posterior probability, is the likelihood function, is the prior probability, is the marginal probability of the data; S7-3: Randomly sample a large number of samples from the posterior distribution of the parameters, use Monte Carlo simulation techniques to simulate the propagation of uncertainty in the model, and evaluate the uncertainty distribution of the model output. The specific formula is: Where, U is the uncertainty distribution, N is the number of samples, is the model output; S7-4: According to the uncertainty analysis results, reduce the impact of uncertainty on the evaluation results by increasing the data volume.

[0098] It should be noted that data uncertainty: Uncertainty caused by measurement errors, data missing or incomplete.

[0099] Model uncertainty: Uncertainty caused by model assumptions, parameter estimation, or model structure selection.

[0100] Prediction Uncertainty: The uncertainty of the prediction result due to the uncertainty of future conditions (meteorological conditions, human activities).

[0101] The specific steps of the Bayesian method include: 1. Set the prior distribution: Set the prior distribution of the model parameters according to prior knowledge .

[0102] 2. Calculate the likelihood function: Calculate the likelihood function based on the observed data , which represents the probability of the data given the parameters.

[0103] 3. Update the posterior distribution: Update the posterior distribution of the parameters through Bayes' theorem .

[0104] The specific steps of the Monte Carlo simulation technique for the uncertainty distribution include: 1. Random sampling: Randomly draw a large number of samples from the posterior distribution of the parameters.

[0105] 2. Uncertainty propagation: Input these samples into the model and calculate the uncertainty distribution of the model output.

[0106] 3. Uncertainty assessment: Evaluate the uncertainty distribution of the model output through statistical analysis (calculating the mean, variance, confidence interval).

[0107] Increase the data volume: By increasing the number of observed data, reduce the randomness and uncertainty of the data.

[0108] The code for increasing the data volume to reduce uncertainty includes: import numpy as np import pandas as pd from sklearn.ensemble import RandomForestRegressor from sklearn.model_selection import train_test_split from sklearn.metrics import mean_squared_error # Assume the initial data has been loaded into a DataFrame # initial_data = pd.read_csv('initial_data.csv') # Features and target variable X_initial = initial_data.drop('flood_probability', axis=1) y_initial = initial_data['flood_probability'] # Initial model training rf = RandomForestRegressor(n_estimators=100, max_depth=20, random_state=42) rf.fit(X_initial, y_initial) # Assume new data has been loaded into a DataFrame # new_data = pd.read_csv('new_data.csv') # Add new data to the initial data X_new = new_data.drop('flood_probability', axis=1) y_new = new_data['flood_probability'] # Update the dataset X_updated = pd.concat([X_initial, X_new], ignore_index=True) y_updated = pd.concat([y_initial, y_new], ignore_index=True) # Retrain the model rf.fit(X_updated, y_updated) # Evaluate the model performance X_train, X_test, y_train, y_test = train_test_split(X_updated, y_updated, test_size=0.3, random_state=42) rf.fit(X_train, y_train) y_pred = rf.predict(X_test) mse = mean_squared_error(y_test, y_pred) print(f"Mean Squared Error after adding new data: {mse:.4f}") The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included within the protection scope of the present application.

Claims

1. A method for urban flood prediction and risk assessment based on multi-source data fusion, characterized in that, The method includes: Step S1: Collect data from multiple data sources, including meteorological data, hydrological data, geographic information data, and socioeconomic data; Step S2: Preprocess the collected data, and reduce the dimension and fuse the preprocessed data through data fusion technology; Step S3: Build a flood prediction model based on random forest to predict the flood occurrence probability; Step S4: Build a risk assessment index system, including hazard index, exposure index, vulnerability index, and resilience index; Step S5: Use the analytic hierarchy process and random forest method to determine the weights of each index, and build a comprehensive risk assessment model; Step S6: Through real-time monitoring data combined with geographic information system technology, realize the dynamic assessment and visual display of flood risk; Step S7: Adopt Bayesian method and Monte Carlo simulation technology to quantify and process multi-source uncertainties.

2. The urban flood prediction and risk assessment method based on multi-source data fusion according to claim 1, characterized in that: Wherein in step S1, the following sub-steps are further included: S1-1: Obtain meteorological data through meteorological stations, satellite remote sensing, and meteorological radars. The meteorological data includes rainfall, rainfall intensity, temperature, and humidity data; S1-2: Obtain hydrological data through hydrological monitoring stations, water level sensors, and groundwater monitoring wells. The hydrological data includes river water level, flow, and groundwater level data; S1-3: Obtain geographic information data through digital elevation models, aerial photogrammetry, and satellite remote sensing images. The geographic information data includes terrain and land use type data; S1-4: Obtain socioeconomic data through population census, mobile communication data, urban planning data, and municipal engineering data. The socioeconomic data includes population distribution, building distribution, and infrastructure distribution data.

3. The urban flood prediction and risk assessment method based on multi-source data fusion according to claim 1, characterized in that: Wherein in step S2, the following sub-steps are further included: S2-1: Clean and process the collected data, including removing duplicates, removing error records, removing irrelevant records, filling missing values, and handling outliers; S2-2: Standardize the preprocessed data, normalizing the mean of each feature to 0 and the standard deviation to 1; S2-3: Calculate the covariance matrix of the standardized data, extract eigenvalues and eigenvectors, select the top k eigenvectors with the largest eigenvalues as the principal components, and project the original data onto the principal components to achieve data dimension reduction. The specific formula of the covariance matrix is: Among them, represents the covariance matrix, is the i-th sample vector, is the sample mean vector, n is the number of samples, and T represents the transpose operation of the matrix; S2-4: Use time interpolation methods and spatial interpolation methods to synchronize data with different time resolutions and spatial resolutions to a unified time and space grid. The time interpolation methods include linear interpolation and polynomial interpolation, and the spatial interpolation methods include Kriging interpolation and nearest neighbor interpolation; S2-5: Integrate meteorological data, hydrological data, geographic information data, and socioeconomic data through data fusion technology to form a unified data set. The data fusion technology includes weighted average method, Bayesian fusion, and machine learning fusion.

4. The method for urban flood prediction and risk assessment based on multi-source data fusion according to claim 1, characterized in that: In step S3, the following sub-steps are further included: S3-1. Determine the parameters of the random forest model, and the parameters include the number of trees, the depth of the trees, and the minimum number of samples for splitting nodes; S3-2. Divide the data into a training set and a validation set, and use the training set data to train the random forest model. The training process of the random forest model is expressed as: Among them, is the predicted value, Q is the number of trees, q represents the q-th decision tree in the random forest, and a represents the feature vector input into the random forest model. is the predicted value of the q-th tree; S3-3. Use k-fold cross-validation, divide the training data into k subsets, where each subset is used as the validation set in turn, and the remaining data is used as the training set for multiple training and validation. The formula for cross-validation is: Among them, CV is the cross-validation result, k is the number of folds, is the mean squared error of the b-th fold; S3-4. According to the results of cross-validation, select the model parameter combination for the flood prediction model.

5. The method for urban flood prediction and risk assessment based on multi-source data fusion according to claim 1, characterized in that: In step S4, the following sub-steps are further included: S4-1. Determine the hazard indicators for the likelihood and intensity of flood occurrence, and the hazard indicators include flood inundation frequency, average annual rainfall during the rainy season, and average annual number of heavy rain days; S4-2. Determine the exposure indicators for evaluating the number of people and property that may be affected by floods, and the exposure indicators include population density and building distribution; S4-3. Determine the vulnerability indicators for evaluating the degree of loss that may be suffered under the action of floods, and the vulnerability indicators include the flood resistance ability of buildings and the flood prevention awareness of residents; S4-4. Determine the resilience indicators for evaluating the speed and degree of recovery after floods, and the resilience indicators include the social and economic recovery ability.

6. The method for urban flood prediction and risk assessment based on multi-source data fusion according to claim 1, characterized in that: In step S5, the following sub-steps are further included: S5-1. Through expert scoring and consistency test, use the analytic hierarchy process to construct a judgment matrix and determine the subjective weights of each indicator. The specific formula is: Where W represents the calculated weight, m is the number of indicators, is the maximum eigenvalue, represents the element in the u-th row and v-th column of the judgment matrix, indicating the importance of the u-th indicator relative to the v-th indicator, represents the sum of all elements j in the v-th column of the judgment matrix, that is, the total importance of the v-th indicator; S5-2. Through random forest method model training and feature importance analysis, determine the objective weights of each indicator. The specific formula is: Among them, is the importance of the o-th feature, Q is the number of trees, is the change in mean squared error of the o-th feature on the q-th tree; S5-3. Combine the subjective weights and objective weights to construct a comprehensive risk assessment model to evaluate the flood risk. The formula for the comprehensive risk assessment model is: Among them, R is the comprehensive risk, m is the number of indicators, is the weight of the g-th indicator, is the value of the g-th indicator.

7. The method for urban flood prediction and risk assessment based on multi-source data fusion according to claim 1, characterized in that: In step S6, the following sub-steps are further included: S6-1. Through Internet of Things sensors and real-time data transmission technology, obtain real-time meteorological, hydrological, and geographical information data; S6-2. Use geographic information system technology to combine flood risk data with geographical information for geospatial visualization, generate a risk map, and display the risk distribution. The formula for geospatial visualization is: Among them, Risk Map is the risk map, Location is the geographical location, and R is the comprehensive risk, which is calculated through the comprehensive risk assessment model; S6-3. Update the input of the comprehensive risk assessment model according to real-time data, recalculate the risk value, and achieve dynamic monitoring and assessment. The dynamic assessment process is expressed as: Among them, is the risk value at time s, m is the number of indicators, is the weight of the g-th indicator, is the value of the g-th indicator at time s.

8. The method for urban flood prediction and risk assessment based on multi-source data fusion according to claim 1, wherein: In step S7, the following sub-steps are further included: S7-1. Identify and classify the sources of uncertainty affecting flood prediction and risk assessment. The uncertainties include data uncertainty, model uncertainty, and prediction uncertainty; S7-2. Quantify the uncertainty of model parameters and update the posterior distribution of model parameters through Bayes' theorem. The specific formula is: wherein, is the posterior probability, is the likelihood function, is the prior probability, is the marginal probability of the data; S7-3. Randomly sample a large number of samples from the posterior distribution of parameters, use Monte Carlo simulation technology to simulate the propagation of uncertainty in the model, and evaluate the uncertainty distribution of the model output. The specific formula is: where U is the uncertainty distribution and N is the number of samples, is the model output; S7-4. According to the results of uncertainty analysis, reduce the impact of uncertainty on the assessment results by increasing the amount of data.

Citation Information

Patent Citations

  • Snow disaster risk assessment method

    CN115186950A

  • Urban flood risk assessment method and platform based on multi-source data fusion

    CN118886717A

  • Flood simulation and prediction system based on GIS data

    CN119128496A

  • Mine rockburst engineering risk prediction method based on random forest and analytic hierarchy process

    CN119624100A

  • Urban flood prediction method based on Bayesian convolutional neural network

    CN119648078A

Cited By

  • Basin flood real-time early warning method and system based on multi-source data fusion

    CN120974242A

  • A method and system for real-time early warning of watershed floods based on multi-source data fusion

    CN120974242B

  • Four-dimensional evaluation method for flood forecasting result of intelligent algorithm

    CN121434691A