Flood risk identification method and system based on machine learning

By combining progressive learning and incremental learning algorithms with graph neural networks, the problems of low dynamic adaptability and low data fusion efficiency in existing flood risk identification methods are solved. This enables accurate tracking of flood evolution paths and scientific classification of risk levels, thereby improving the accuracy and timeliness of flood control decisions.

CN121723282APending Publication Date: 2026-03-24GEOLOGICAL & NATURAL DISASTER PREVENTION & CONTROL INST GANSU ACADEMY OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing flood risk identification methods are ill-suited to the long-term dynamic evolution of risk patterns, lack a risk knowledge accumulation mechanism, cannot effectively track flood evolution paths, have low efficiency in multi-source data fusion and processing, and have unscientific risk level classifications, which affect the accuracy of flood control resource allocation and emergency response.

Method used

The model parameters are dynamically updated using progressive and incremental learning algorithms. A watershed connectivity analysis module is constructed by combining graph neural networks. Through multi-source data fusion and time series anomaly detection, risk characteristics are identified and classified, and flood evolution patterns and anomalous signals are captured.

Benefits of technology

It significantly improves the accuracy and timeliness of flood forecasting, enhances the dynamic adaptability to flood risks and the ability to capture signals, and provides scientific support for flood control decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121723282A_ABST
    Figure CN121723282A_ABST
Patent Text Reader

Abstract

The invention discloses a flood risk identification method and system based on machine learning, and the method comprises the steps: obtaining a time series segmentation training frame from historical flood data through progressive learning, and inputting events of different years in batches in order to obtain a model parameter fine adjustment optimization result; a structured knowledge base is obtained from key elements stored in a risk precipitation mechanism, similar risk modes are classified through a clustering algorithm to obtain a risk feature vector library, and a graph neural network is adopted to perform drainage basin connectivity analysis. Using the river tributary flood storage and detention areas as nodes and water flow directions as edges to train and recognize a flood peak transfer time attenuation rule to obtain flood routing path prediction parameters; a clustering analysis method is adopted to identify risk value natural groups according to classification response anomalies, and the classification boundary discretization continuous risk probability is adjusted to obtain risk levels. Through combination of data driving and an intelligent algorithm, the flood prediction precision and response timeliness are remarkably improved, and a scientific basis is provided for flood control decision making.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of flood risk identification and prediction, and particularly relates to a flood risk identification method and system based on machine learning. BACKGROUND

[0002] Flood risk identification is a core technology in the field of disaster prevention and mitigation, directly related to the safety of people's lives and property and the stable development of social economy. With the intensification of global climate change, extreme rainfall events occur frequently, which puts forward higher requirements for the accuracy and timeliness of flood risk identification. In recent years, machine learning technology has been widely used in flood prediction. By integrating historical flood data, real-time monitoring data, satellite remote sensing images, and social and economic statistical information, a prediction model is constructed to realize risk classification and early warning response. The existing technology mainly adopts batch training to establish a static prediction model, and combines simple threshold judgment for risk classification. Some schemes introduce incremental learning mechanism to adapt to data updates, but still lack systematic risk knowledge sedimentation and reuse mechanism, and the tracking ability of flood evolution in the basin is limited.

[0003] However, the existing flood risk identification method still has several technical problems to be solved. First, the traditional static model is difficult to adapt to the long-term dynamic evolution of risk patterns, and lags behind in response to new disaster characteristics, and cannot effectively absorb and utilize new data to realize fine adjustment of model parameters. Second, there is a lack of perfect risk sedimentation mechanism, and the historical disaster characteristics cannot be structured and saved to form a reusable knowledge base, which makes it difficult to quickly match existing knowledge when similar risk events occur repeatedly, affecting the efficiency of early warning. Third, the tracking ability of the complete travel trajectory of the flood is insufficient, and it is difficult to systematically grasp the propagation law, time decay characteristics and inundation range expansion speed of the flood peak from upstream to downstream, which restricts the timely capture ability of key early warning signals such as soil water saturation critical point and tributary confluence disaster mutation. Fourth, the fusion processing efficiency of multi-source heterogeneous data is low, and the differences in time and space scale, dimension standard of different data sources lead to complex preprocessing process, it is difficult to build a unified training sample interface, which affects the model training effect. Finally, the existing risk level division mainly relies on artificial threshold setting, and lacks scientific classification boundary adjustment mechanism based on data natural grouping, which leads to deviation between risk level division and actual risk distribution, and further affects the rationality of precise allocation of flood control resources and emergency response decision-making. These problems seriously restrict the actual application efficiency of machine learning in the field of flood risk identification. SUMMARY

[0004] To solve the above technical problems, the present application provides a flood risk identification method and system based on machine learning. Among them, a flood risk identification method based on machine learning, comprising:

[0005] According to historical flood data, a time series segmentation training framework is obtained through progressive learning, and a model parameter fine-tuning optimization result is obtained.

[0006] According to the model parameter fine-tuning optimization result, new rainfall, river level and terrain change data are absorbed through an incremental learning algorithm, and when the deviation degree is judged to exceed a preset threshold, a risk sedimentation mechanism is started, and a key element storage result is obtained.

[0007] According to the key element storage result, similar risk patterns are classified through a clustering algorithm, and a risk feature vector library is obtained.

[0008] According to the risk feature vector library, a river basin connectivity analysis module is constructed through a graph neural network, taking rivers, tributaries and flood storage areas as nodes and water flow directions as edges for training, and a flood evolution path prediction parameter is obtained.

[0009] According to the flood evolution path prediction parameter, multi-source data center real-time monitoring data, satellite remote sensing images and historical record statistical information are fused, processed through a data preprocessing pipeline, and a unified interface training sample is obtained.

[0010] According to the unified interface training sample, a normal state baseline model is established through a time series anomaly detection algorithm, the deviation degree of real-time data is calculated, and a hierarchical response anomaly result is obtained.

[0011] According to the hierarchical response anomaly result, risk value natural grouping is identified through clustering analysis method, and classification boundary is adjusted to discretize continuous risk probability, and risk grade division standard is obtained.

[0012] According to the risk grade division standard, flood risk is identified, and flood risk grade is obtained.

[0013] Preferably, the process of obtaining the model parameter fine-tuning optimization result comprises:

[0014] The historical flood data is sorted by year to obtain a time series.

[0015] The time series is divided into multiple continuous paragraphs to obtain a segmented sequence.

[0016] An initial training framework is constructed according to the segmented sequence to obtain a basic model parameter.

[0017] The current year event data is input in sequence in batches to obtain an updated data set.

[0018] The updated data set is used to perform progressive learning on the basic model parameter to obtain an intermediate model parameter.

[0019] The next year event data is processed according to the intermediate model parameter to obtain a parameter gradual change result.

[0020] Fine-tuning optimization is performed according to the parameter step change result, and final model parameters are obtained.

[0021] Preferably, the process of obtaining the key element storage result comprises:

[0022] Real-time rainfall data, water level data and topographic change information are obtained according to a plurality of monitoring points, and after cleaning processing, a structured environment monitoring data set is obtained;

[0023] According to the structured environment monitoring data set, an incremental learning method is used to gradually absorb new data for parameter updating, and an updated model prediction result is obtained;

[0024] According to the updated model prediction result, the deviation degree of rainfall data, water level data and topographic change is calculated, and if the deviation degree exceeds a preset threshold, an abnormal state judgment result is obtained;

[0025] According to the abnormal state judgment result, the distribution of high-risk areas is analyzed and combined with rainfall threshold and geological vulnerability information, and a potential risk area division result is obtained;

[0026] According to the potential risk area division result, the critical rainfall threshold and geological vulnerability point data in each area are extracted, and a risk database storage result is obtained;

[0027] According to the risk database storage result, the data acquisition frequency is dynamically adjusted and the high-risk area monitoring point density is increased, and a subsequent monitoring optimization scheme is obtained.

[0028] Preferably, the process of obtaining the risk feature vector library comprises:

[0029] According to the key element data stored by the risk sedimentation mechanism, cleaning and formatting processing is performed, and a preliminarily sorted data set is obtained;

[0030] According to the preliminarily sorted data set, core elements are extracted and layered labeled, and a structured information unit is obtained;

[0031] According to the structured information unit, a clustering method is used to group similar patterns, and if the similarity between data exceeds a preset threshold, they are classified into the same category, and a basic classification of risk patterns is obtained;

[0032] According to the basic classification of risk patterns, key features are extracted for each category and corresponding feature vectors are generated, and a risk feature vector library is obtained;

[0033] According to the risk feature vector library, data preparation is performed for subsequent matching scenarios, and if the similarity between input data and feature vectors reaches a preset standard, the corresponding risk category is output, and a matching result is obtained;

[0034] According to the matching result, the result data is archived to the knowledge base and the structured knowledge content is updated, and a dynamically updated risk information base is obtained;

[0035] According to the dynamically updated risk information base, the newly input risk data is compared in real time, and if the comparison result shows that the new data is highly related to the existing category, it is classified into the corresponding category, and the latest risk classification state is obtained.

[0036] Preferably, the process of obtaining the flood evolution path prediction parameter comprises:

[0037] The spatial distribution data of the river, tributary and flood storage area are converted into structured information defined by nodes;

[0038] According to the node-defined structured information, an edge relationship matrix is constructed using water flow direction data to generate a connection topology structure inside the basin;

[0039] According to the connection topology structure, the propagation mode of flood peak transmission is trained using a graph neural network to obtain preliminary parameters of time decay;

[0040] According to the preliminary parameters of time decay, the key path features of flood peak transmission are extracted to determine the main direction of flood evolution;

[0041] The main direction of flood evolution is compared with the historical record for deviation comparison, and the path is corrected through the spatial distribution data to obtain an abnormal node determination result;

[0042] According to the abnormal node determination result, the flood evolution path is updated in combination with the edge relationship matrix to derive the final path prediction parameter.

[0043] Preferably, the process of obtaining the unified interface training sample comprises:

[0044] According to the real-time monitoring system, the satellite remote sensing platform and the historical database, flood-related dynamic information, satellite remote sensing images and historical record statistical information are collected to obtain a preliminary data set;

[0045] According to the preliminary data set, a data cleaning tool is used to fill in the missing values, and the abnormal points exceeding the preset threshold range are corrected by the adjacent value interpolation method to obtain a cleaned data set;

[0046] According to the cleaned data set, a spatio-temporal alignment technology is used to match different source data according to a unified timestamp and spatial grid to obtain an aligned data set;

[0047] According to the aligned data set, a scale normalization method is applied to convert different dimensional data into a unified range to obtain a standardized data set;

[0048] According to the standardized data set, feature extraction is performed on key variables related to flood evolution and path prediction, irrelevant variables are removed if the correlation between features is lower than a preset threshold, and a selected feature set is obtained;

[0049] According to the selected feature set, a pre-trained model is input to perform a flood path prediction task, and a prediction result data is obtained;

[0050] According to the prediction result data, a structured training sample is generated and stored in a unified interface database.

[0051] Preferably, the process of obtaining a graded response abnormal result comprises:

[0052] According to the unified interface, multi-source historical data is collected to form a training sample;

[0053] According to the training sample, a time series anomaly detection algorithm is used to establish a normal state baseline model;

[0054] According to the baseline model, the deviation degree of the real-time data sequence is calculated, and the current state deviation value is obtained;

[0055] According to the real-time data sequence, the rainfall intensity feature and the upstream reservoir water level change rate are extracted, if the rainfall intensity feature exceeds a preset threshold and the water level change rate shows an upward trend, a signal of rapid rise of reservoir water level is marked;

[0056] According to the real-time data sequence, the soil water content change rate is extracted, if the soil water content change rate is close to the saturation state, a soil saturation critical signal is marked;

[0057] According to the detection results of the rapid rise of reservoir water level signal and the soil saturation critical signal, combined with the deviation degree amplification capture signal strength, a comprehensive abnormal score is obtained;

[0058] According to the comparison result of the comprehensive abnormal score and the threshold value of different levels, the corresponding graded response abnormal result is determined.

[0059] Preferably, the process of obtaining a risk level division standard comprises:

[0060] According to the graded response abnormal data set, a K-means clustering algorithm is used to group process the risk probability value, and a plurality of natural grouping centers are obtained;

[0061] According to the distance between adjacent centers calculated according to the natural grouping center, the grouping boundary position is determined;

[0062] According to the grouping boundary position, the continuous risk probability is discretized to obtain a preliminary discrete risk value;

[0063] According to the group range into which the preliminary discrete risk value falls, a corresponding risk level is respectively assigned, the lowest group range is assigned a low risk level, and the highest group range is assigned an extremely high risk level;

[0064] According to the assigned risk level set, a risk level division standard is formed.

[0065] The application also provides a flood risk identification system based on machine learning, comprising:

[0066] The progressive learning training module is configured to obtain a time series segmentation training framework through progressive learning according to historical flood data, and obtain a model parameter fine-tuning optimization result.

[0067] The incremental learning optimization module is configured to absorb new rainfall, river water level and topographic change data through an incremental learning algorithm according to the model parameter fine-tuning optimization result, and when it is judged that the deviation degree exceeds a preset threshold, a risk sedimentation mechanism is started, and a key element storage result is obtained.

[0068] The risk sedimentation mechanism module is configured to classify similar risk patterns through a clustering algorithm according to the key element storage result, and obtain a risk feature vector library.

[0069] The risk feature vector library construction module is configured to construct a river basin connectivity analysis module through a graph neural network according to the risk feature vector library, take rivers, tributaries and flood storage areas as nodes, and take water flow directions as edges for training, and obtain flood evolution path prediction parameters.

[0070] The flood evolution path prediction module is configured to fuse multi-source data center real-time monitoring data, satellite remote sensing images and historical record statistical information according to the flood evolution path prediction parameters, process through a data preprocessing pipeline, and obtain a unified interface training sample.

[0071] The multi-source data fusion preprocessing module is configured to establish a normal state baseline model through a time series anomaly detection algorithm according to the unified interface training sample, calculate a real-time data deviation degree, and obtain a hierarchical response abnormal result.

[0072] The risk level division module is configured to identify risk value natural grouping through a clustering analysis method according to the hierarchical response abnormal result, adjust classification boundaries to discretize continuous risk probability, and obtain a risk level division standard; and according to the risk level division standard, flood risk is identified, and a flood risk level is divided.

[0073] Compared with the prior art, the application has the following advantages and technical effects:

[0074] This invention discloses a flood risk prediction and response technology based on multi-source data fusion. Addressing the comprehensive operational challenges of predicting flood evolution paths, classifying risk levels, and capturing anomalous signals in a watershed, the technology employs progressive learning to train a model segmented from historical data. Combined with incremental learning algorithms, it dynamically absorbs new data, assesses risk deviations, and activates a risk accumulation mechanism to store key elements and build a knowledge base. Furthermore, it analyzes watershed connectivity using clustering algorithms and graph neural networks to predict flood peak transmission patterns. By integrating real-time monitoring and multi-source data, and after preprocessing, trains an anomaly detection model to capture short-term anomalous signals. Cluster analysis is then used to adjust risk level boundaries, achieving a tiered response from low to extremely high. This invention, through the combination of data-driven approaches and intelligent algorithms, significantly improves the accuracy and timeliness of flood prediction, providing a scientific basis for flood control decision-making.

[0075] This invention acquires a time-series segmented training framework from historical flood data through a progressive learning mechanism, and dynamically absorbs new data and initiates a risk accumulation mechanism using an incremental learning algorithm. This significantly improves the system's adaptability to risk pattern evolution and the structured accumulation and efficient reuse of historical knowledge. A graph neural network is used to construct a watershed connectivity analysis module, using river tributaries and flood storage areas as nodes and water flow direction as edges for training. This enables accurate tracking of flood evolution paths and the decay law of flood peak transmission time. A multi-source data fusion preprocessing pipeline performs missing value imputation, outlier detection, spatiotemporal alignment, and scale normalization, constructing a unified interface training sample, which greatly improves the efficiency of heterogeneous data processing and model training quality. A normal state baseline model is established based on a time-series anomaly detection algorithm, effectively capturing key early warning signals such as short-term heavy rainfall, rapid rise in upstream reservoir water levels, and soil moisture saturation critical points. Cluster analysis is used to identify natural groupings of risk values ​​and adjust classification boundaries, achieving a scientific classification standard for discretizing continuous risk probabilities into low, medium, high, and extremely high levels. The solution of this invention enhances the dynamic adaptability, evolution tracking capability, signal capture accuracy, and risk classification rationality of flood risk identification, providing more accurate and timely decision support for flood prevention and disaster reduction. Attached Figure Description

[0076] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:

[0077] Figure 1 This is a schematic diagram of the method flow according to an embodiment of the present invention. Detailed Implementation

[0078] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0079] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0080] Example 1

[0081] like Figure 1 As shown, this embodiment provides a flood risk identification method based on machine learning, including:

[0082] Based on historical flood data, a time series segmented training framework is obtained through progressive learning, and the model parameters are fine-tuned and optimized.

[0083] Based on the model parameter fine-tuning and optimization results, new rainfall, river water level and terrain change data are absorbed through incremental learning algorithm. When the deviation exceeds the preset threshold, the risk accumulation mechanism is activated to obtain the storage results of key elements.

[0084] Based on the storage results of key elements, similar risk patterns are classified through clustering algorithms to obtain a risk feature vector library;

[0085] Based on the risk feature vector library, a watershed connectivity analysis module is constructed using a graph neural network. Rivers, tributaries, and flood storage areas are used as nodes, and the direction of water flow is used as edges for training to obtain flood evolution path prediction parameters.

[0086] Based on the flood evolution path prediction parameters, real-time monitoring data from multiple data centers, satellite remote sensing images, and historical statistical information are integrated, and data preprocessing pipeline is used to obtain unified interface training samples.

[0087] Based on the unified interface training samples, a normal state baseline model is established through a time series anomaly detection algorithm, the degree of deviation of real-time data is calculated, and graded response anomaly results are obtained.

[0088] Based on the results of graded response anomalies, risk values ​​are naturally grouped using cluster analysis, and the classification boundaries are adjusted to discretize the probability of continuous risks, thereby obtaining the risk level classification criteria.

[0089] Flood risk is identified and classified according to the risk level classification standard.

[0090] Furthermore, the process of obtaining the model parameter fine-tuning optimization results includes:

[0091] Historical flood data are sorted by year to obtain a time series;

[0092] Divide the time series into multiple consecutive segments to obtain a segmented sequence;

[0093] An initial training framework is constructed based on the segmented sequence to obtain basic model parameters;

[0094] Input the event data for the current year in batches in sequence to obtain the updated dataset;

[0095] Perform incremental learning on the base model parameters based on the updated dataset to obtain intermediate model parameters;

[0096] The event data for the next year is further processed based on the intermediate model parameters to obtain the results of the gradual changes in parameters.

[0097] Fine-tuning and optimization are performed based on the results of gradual parameter changes to obtain the final model parameters.

[0098] Furthermore, in this embodiment, when processing historical flood data, the data is first sorted by year to form a complete time series. Assume there is flood data from 2000 to 2020, recording the number and severity of floods each year. Constructing a time series helps to intuitively understand the evolution trend of flood events; for example, the number of floods was 3 in 2000, increasing to 5 in 2005, peaking at 8 in 2010, and then gradually declining. This serialization lays the foundation for subsequent analysis and helps to capture long-term patterns.

[0099] For example, in this embodiment, when dividing the time series into multiple consecutive segments, the time series is segmented based on significant changes in flood events. Assuming the first segment is from 2000 to 2005, with a gradually increasing flood frequency; the second segment is from 2006 to 2010, with high-frequency fluctuations; and the third segment is from 2011 to 2020, with a gradually decreasing frequency. Such segmented sequences can reflect the characteristics of different stages, facilitating targeted modeling. The advantage of segmentation is that it avoids overfitting a single model to complex trends and improves the model's adaptability to local features.

[0100] For example, in constructing the initial training framework in this embodiment, a basic model, such as a linear regression model, is selected based on the segmented sequences to initially fit the flood frequency trend of each segment. By analyzing data from 2000 to 2005, initial parameters, such as slope and intercept, are determined to reflect the increasing trend of flood frequency over time. This initial framework provides a starting point for subsequent optimization, ensuring that the model has a certain predictive ability while reducing the complexity of subsequent training.

[0101] For example, if an event in the current year is earlier than a previously processed year, it is skipped to avoid duplicate calculations. Suppose the current processing is up to 2015 data, and the new input is an event from 2010; in this case, it is skipped, ensuring the temporal order and efficiency of data processing. This mechanism avoids resource waste and improves system efficiency.

[0102] For example, when inputting event data for the current year sequentially in batches, flood data from 2016 can be added to the updated dataset in batches. Assuming there were four flood events in 2016, the dataset would be updated after each input, containing the latest frequency and severity information. This progressive input method ensures that the model can respond promptly to new data, maintaining the real-time nature of predictions.

[0103] For example, in this embodiment, when incremental learning is performed on the basic model parameters by updating the dataset, changes in the intermediate model parameters can be observed. Assuming the initial model slope parameter is 0.5, it is adjusted to 0.6 after adding 2016 data, reflecting the increasing trend in flood frequency. The advantage of incremental learning is that the model can dynamically adapt to new data, avoiding the risk of becoming outdated due to one-time training.

[0104] For example, when processing event data for the following year, the model parameters are continuously adjusted based on intermediate data. For instance, after inputting data for 2017, the parameters are fine-tuned from 0.6 to 0.65. This gradual adjustment ensures the model's sensitivity to new trends while maintaining stability and reducing prediction errors caused by drastic parameter fluctuations.

[0105] For example, when performing fine-tuning to obtain the final model parameters, the parameters can be adjusted through multiple iterations to minimize the model's prediction error from 2017 to 2020. Assuming a final slope of 0.7 and a slight adjustment to the intercept, the model's prediction of future flood trends becomes more accurate. The benefit of fine-tuning is that it improves the model's generalization ability, ensuring that it maintains high accuracy when facing unknown data.

[0106] Through the above methods, from time series construction to final model parameter optimization, the entire process forms a rigorous logical chain, with each link supporting the others, ensuring the scientific validity and practicality of the flood prediction model, while significantly improving prediction accuracy and adaptability.

[0107] Furthermore, the process of obtaining the key element storage results includes:

[0108] Based on real-time rainfall data, water level data, and topographic change information obtained from multiple monitoring points, a structured environmental monitoring dataset was obtained after cleaning and processing.

[0109] Based on the structured environmental monitoring dataset, an incremental learning method is used to gradually absorb new data to update the parameters and obtain the updated model prediction results.

[0110] Based on the updated model prediction results, the deviation of rainfall data, water level data and topographic changes is calculated. If the deviation exceeds the preset threshold, the abnormal state judgment result is obtained.

[0111] Based on the results of the abnormal status assessment, the distribution of high-risk areas is analyzed, and combined with rainfall thresholds and geological vulnerability information, the results of the potential risk area classification are obtained.

[0112] Based on the results of the potential risk area division, the critical rainfall threshold and geological vulnerability point data of each area are extracted to obtain the risk database storage results;

[0113] Based on the results stored in the risk database, the data collection frequency is dynamically adjusted and the density of monitoring points in high-risk areas is increased to obtain subsequent monitoring optimization plans.

[0114] Furthermore, this embodiment addresses the acquisition of rainfall data, water level data, and topographic change information by deploying multiple sensor nodes upstream and downstream of the river to collect hourly rainfall and water level data in real time. For example, a monitoring point might record 80 millimeters of rainfall and a 0.5-meter rise in water level within 24 hours. Simultaneously, satellite imagery is used to analyze topographic subsidence. After preliminary cleaning to remove outliers, this data is stored in a database to form a structured dataset for subsequent analysis.

[0115] For example, when using incremental learning to update model parameters, consider this scenario: newly collected data shows a continuous increase in rainfall in a certain area. The system will input this data into the pre-trained model in batches, gradually adjusting the model weights to ensure that the prediction results better reflect the current environmental changes. For instance, the model might adjust the predicted probability of flood risk for a certain river based on the new data, increasing it from 30% to 45%, thereby improving the timeliness of early warnings.

[0116] For example, this embodiment calculates the degree of deviation and determines abnormal states by setting a threshold. Assuming the normal water level fluctuation range is 0.2 meters, and the water level at a certain monitoring point rises by 0.8 meters in a short period, exceeding the threshold by four times, the system will automatically mark it as an abnormal state and lock in data within a 10-kilometer radius of that area for in-depth analysis. This method helps to quickly locate the problem area.

[0117] For example, in the risk assessment mechanism of this embodiment, the system combines rainfall thresholds and geological vulnerability information to analyze the distribution of high-risk areas. Suppose that historical data shows a region is prone to landslides when rainfall exceeds 100 mm, and topographic data indicates a steep slope, the system will classify this region as a potential risk zone. This classification provides a basis for subsequent resource allocation.

[0118] For example, when extracting critical rainfall thresholds and geological vulnerability point data in this embodiment, key indicators can be recorded separately for each risk area. For instance, if the critical rainfall threshold for a certain area is 90 mm and the geological vulnerability points are concentrated in the middle of the hillside, the system will store this data in the risk database and sort it according to the risk level, prioritizing the areas most likely to be affected by disasters.

[0119] For example, regarding dynamically adjusting the data collection frequency, this embodiment increases the monitoring frequency from once per hour to once every 30 minutes for high-risk areas, while also adding temporary monitoring points in the area to ensure higher data coverage. For instance, after adding three monitoring points in a high-risk area, the system can more accurately capture water level change trends. This optimization scheme can significantly improve the precision of monitoring and provide reliable support for subsequent decision-making.

[0120] Furthermore, the process of obtaining the risk feature vector library includes:

[0121] Based on the key element data stored by the risk accumulation mechanism, the data is cleaned and formatted to obtain a preliminary organized dataset.

[0122] Based on the initially compiled dataset, core elements are extracted and layered annotations are performed to obtain structured information units;

[0123] Based on structured information units, clustering methods are used to group similar patterns. If the similarity between data exceeds a preset threshold, they are classified into the same category to obtain the basic classification of risk patterns.

[0124] Based on the basic classification of risk patterns, key features are extracted for each category and corresponding feature vectors are generated to obtain a risk feature vector library.

[0125] Based on the risk feature vector library, data preparation is carried out for subsequent matching scenarios. If the similarity between the input data and the feature vector reaches the preset standard, the corresponding risk category is output and the matching result is obtained.

[0126] Based on the matching results, the result data is archived into the knowledge base and the structured knowledge content is updated to obtain a dynamically updated risk information database.

[0127] Based on a dynamically updated risk information database, newly input risk data is compared in real time. If the comparison results show that the new data is highly correlated with existing categories, it is classified into the corresponding category, and the latest risk classification status is obtained.

[0128] Furthermore, in processing the mechanistic data in the risk accumulation, this embodiment first extracts key information related to rainfall and water levels from historical monitoring records. For example, if past records show that areas frequently experience flooding when rainfall exceeds 50 mm, the automated tool, using this data, cleans the original records of fields such as time, location, and rainfall, removing invalid or duplicate entries to form a standardized dataset. This process ensures the accuracy of subsequent analysis.

[0129] For example, this embodiment focuses on core element extraction and hierarchical labeling. Based on preset rules, the data is divided into three categories: high, medium, and low risk levels. For instance, if rainfall in a certain area exceeds 40 mm for three consecutive days, and the terrain shows an increase in slope, it is labeled as high-risk. This hierarchical labeling helps to quickly identify potentially threatening areas and provides clear guidance for subsequent processing.

[0130] For example, in this embodiment, the similarity threshold is set to 80% when performing clustering. If the rainfall patterns and water level fluctuation trends of two sets of data reach this standard, they are classified into the same category. For instance, if data from two monitoring points both show that the water level rose by more than 1 meter within 24 hours after rainfall, they are classified into the same risk pattern. This classification method helps to identify common problems and improves analysis efficiency.

[0131] For example, in this embodiment, when extracting key features and generating feature vectors, peak rainfall, water level rise rate, and geological stability are used as core indicators. Assuming a region has a peak rainfall of 60 mm and a water level rise rate of 0.5 meters per hour, a corresponding feature vector is generated. This vector library lays the foundation for subsequent matching.

[0132] For example, in a matching scenario, if new input data shows rainfall of 55 millimeters and a water level rise rate of 0.4 meters per hour, and the similarity to a high-risk category in the feature vector library reaches 90%, then the output is a high-risk category. This matching mechanism can quickly respond to new data and provide timely risk assessments.

[0133] For example, in this embodiment, when archiving data to the knowledge base, matching results can be stored by time and region. For instance, high-risk data for a specific region in a particular month can be archived separately for easy subsequent retrieval. This dynamic updating method ensures the timeliness of the knowledge base.

[0134] For example, in real-time comparison of new input data, if the new data shows rainfall of 58 mm and has an 85% similarity to an existing high-risk category, it will be classified into that category. This real-time classification mechanism can continuously optimize the risk information database and maintain the dynamic adaptability of the data.

[0135] Through the coordinated processing of these multiple stages, the accuracy of risk identification and the speed of response have been significantly improved, providing strong support for flood monitoring and early warning.

[0136] Furthermore, the process of obtaining flood path prediction parameters includes:

[0137] Transform the spatial distribution data of rivers, tributaries, and flood storage areas into structured information defined by nodes;

[0138] Based on the structured information defined by the nodes, an edge relationship matrix is ​​constructed using water flow direction data to generate the connection topology within the watershed;

[0139] Based on the connection topology, a graph neural network is used to train the propagation mode of the flood peak to obtain preliminary parameters of time decay.

[0140] Based on the preliminary parameters of time decay, the key path features of flood peak transmission are extracted to determine the main direction of flood evolution;

[0141] By comparing the main direction of flood evolution with historical records and correcting the path using spatial distribution data, the results of abnormal node identification are obtained.

[0142] Based on the results of the abnormal node determination, the flood evolution path is updated by combining the edge relationship matrix, and the final path prediction parameters are derived.

[0143] Furthermore, in the construction of the watershed connectivity analysis module in this embodiment, the spatial distribution data of river tributaries and flood storage areas are first acquired. This data comes from vector layers in a Geographic Information System (GIS), including river centerlines and flood zone polygon boundaries. When converting these geographic entities into structured information defined by nodes, the starting and ending points of each river segment can be defined as nodes, while flood storage areas are included as special nodes, ensuring that all hydrological units are abstracted into computable point elements.

[0144] Specifically, in one possible implementation, node definitions include node ID, coordinates, and elevation. For example, a node on the main river channel might be labeled ID:001, with coordinates (115.32, 34.18) and an elevation of 45.6 meters. Nodes in adjacent flood storage and detention areas would have additional capacity attributes, such as a volume of 250 million cubic meters. This structured transformation facilitates subsequent unified processing and avoids the heterogeneity issues of the original spatial data.

[0145] It should be noted that when constructing the edge relationship matrix using water flow direction data, the water flow direction is derived from the digital elevation model (DEM). If there is a downstream relationship between node i and node j in the matrix, the corresponding element is 1; otherwise, it is 0.

[0146] For example, in a small to medium-sized watershed with 50 nodes, the edge relation matrix may exhibit a sparse upper triangular structure, reflecting the unidirectional nature of water flow. This matrix not only captures topological connectivity but also provides the basic propagation path for flood peak transmission.

[0147] In one embodiment, when training the propagation pattern of flood peak transmission using a graph neural network, the input is an edge relationship matrix combined with historical flood event data. For example, in 2020, a flood peak took 4.2 hours to propagate from an upstream node to the downstream. The network learns the time decay parameters between nodes through multi-layer graph convolution, initially obtaining an average decay coefficient of 0.85 per kilometer. These parameters reflect the energy dissipation of the flood peak during propagation, helping to more accurately simulate real water flow resistance.

[0148] For example, in this embodiment, when extracting critical path features of flood peak transmission from preliminary parameters of time decay, the main channel is determined as the primary direction by using the shortest path algorithm or feature importance ranking, while tributaries only account for 15% of the propagation contribution. If it is found that the main direction of flood evolution deviates significantly from historical records, such as the simulated path deviating from the actual flood trail by more than 10 kilometers, path correction is performed by comparing spatial distribution data. For example, the accuracy of the DEM is re-examined or the obstruction factor of dikes is added to identify potential abnormal nodes, such as flow interruption points caused by human-induced diversion.

[0149] Specifically, after obtaining the correction results of abnormal nodes, this embodiment recalculates the flood evolution path in conjunction with the updated edge relation matrix, and finally derives the path prediction parameters. For example, the predicted delay time for a tributary to flow into the main channel is adjusted from 3.8 hours to 5.1 hours. This correction significantly improves the reliability of the path and avoids the bias of relying solely on historical data.

[0150] In one possible implementation, this embodiment generates dynamic simulation results of flood propagation within the basin using path prediction parameters. It can output hourly water level diffusion maps, clearly identifying the distribution of key areas affected by the flood. For example, after upstream flood storage and detention areas are diverted, the water depth in downstream urban areas decreases by 0.7 meters. This dynamic simulation not only supports real-time early warning but also optimizes flood control scheduling decisions, reduces potential economic losses, and enhances the overall resilience of the basin.

[0151] Furthermore, the process of obtaining training samples for the unified interface includes:

[0152] Based on real-time monitoring systems, satellite remote sensing platforms, and historical databases, we collect flood-related dynamic information, satellite remote sensing images, and historical statistical information to obtain a preliminary dataset.

[0153] Based on the initial dataset, data cleaning tools were used to fill in missing values, and outliers exceeding the preset threshold range were corrected using nearest neighbor interpolation to obtain the cleaned dataset.

[0154] Based on the cleaned dataset, spatiotemporal alignment technology is used to match data from different sources according to a unified timestamp and spatial grid to obtain an aligned dataset;

[0155] Based on the aligned dataset, the scale normalization method is applied to convert data of different scales into a uniform range to obtain a standardized dataset.

[0156] Based on the standardized dataset, features are extracted from key variables related to flood evolution and path prediction. If the correlation between features is lower than a preset threshold, irrelevant variables are removed to obtain a carefully selected feature set.

[0157] Based on a carefully selected feature set, the pre-trained model is input to perform a flood path prediction task, and prediction results are obtained.

[0158] Based on the prediction results, structured training samples are generated and stored in a unified interface database.

[0159] Furthermore, this embodiment acquires dynamic flood-related information, such as hourly water level and flow data reported by upstream hydrological stations, downloads the latest rainfall distribution images from a satellite remote sensing platform, and retrieves historical flood record statistics to form a preliminary dataset. This data acquisition method can comprehensively cover the natural driving factors and social influencing factors of floods, providing a rich foundation for subsequent forecasting.

[0160] In one possible implementation, real-time monitoring data provides dynamic data including the main river level reaching 8.5 meters and tributary flow exceeding 1,500 cubic meters per second, while satellite remote sensing captures cloud images of 24-hour cumulative rainfall exceeding 200 millimeters in the upstream mountainous areas. Historical records supplement the arrival time of past flood peaks under similar rainfall conditions. These data together constitute a preliminary dataset, effectively improving the timeliness and accuracy of predictions.

[0161] Specifically, in this embodiment, when using data cleaning tools to fill in missing values, if a monitoring station experiences a three-hour gap in water level data due to equipment failure, it can be supplemented through linear interpolation of preceding and following time periods or spatial interpolation of adjacent stations. For outlier detection, when a flow rate suddenly jumps to 5000 cubic meters per second, exceeding twice the threshold range of the station's historical maximum value, the system automatically marks it and corrects it to a reasonable range using nearest-neighbor interpolation, thus obtaining a cleaned dataset. This processing significantly reduces noise interference and improves data reliability.

[0162] In one possible implementation, after cleaning, the missing rate in the original dataset is reduced from 15% to below 2%, and outlier correction makes the overall data fluctuations more in line with physical laws, laying a high-quality foundation for subsequent alignment.

[0163] It should be noted that, in this embodiment, when matching data from different sources using spatiotemporal alignment technology, a unified timestamp, such as the hourly mark, is used as the reference. The rainfall grid of satellite imagery and the point data of hydrological stations are projected onto the same spatial grid, for example, using a grid system with a 1-kilometer resolution, to ensure that all data are consistent in the spatiotemporal dimension. This alignment helps to capture the complete process of flooding from upstream rainfall to downstream water level response.

[0164] For example, in a heavy rainfall event, the upstream remote sensing rainfall data showed 250 mm in grid A, while the downstream hydrological station showed a water level rise 6 hours later. By aligning the data, the causal relationship between the two can be clearly linked, improving the causal traceability capability of evolution path prediction.

[0165] Specifically, in this embodiment, when applying the scale normalization method, data of different dimensions such as water level, flow rate, and rainfall are converted into a range of 0 to 1. For example, the highest historical water level of 15 meters is normalized to 1, and the lowest of 3 meters is normalized to 0. This standardization facilitates the comparison of the importance of different variables in the model and avoids bias caused by differences in dimensions.

[0166] In one embodiment, the correlation coefficient between normalized flow and rainfall is improved from the original 0.62 to a more balanced feature weight distribution, making subsequent feature extraction fairer and more effective.

[0167] For example, when constructing the feature extraction module, key variables related to flood evolution, such as upstream rainfall accumulation, river slope, and soil saturation, are screened. If the correlation between a variable and the flood peak arrival time is lower than a preset threshold of 0.3, it is discarded, thus obtaining a carefully selected feature set. This screening reduces redundancy and improves model efficiency.

[0168] Specifically, this embodiment inputs a carefully selected feature set into a pre-trained random forest model to perform a prediction task, outputting the predicted arrival time and peak flow of downstream stations. For example, it predicts that the main river flood peak will arrive in 12 hours, with a peak flow of 3200 cubic meters per second. This prediction result is converted into structured samples and stored in a database to support the connectivity training of subsequent graph neural network modules, forming a complete data preparation loop and effectively improving the robustness and accuracy of the entire flood evolution path prediction system.

[0169] Furthermore, the process of obtaining graded response anomaly results includes:

[0170] Training samples are generated by collecting historical data from multiple sources using a unified interface.

[0171] Based on the training samples, a normal state baseline model is established using a time series anomaly detection algorithm;

[0172] The deviation of the real-time data sequence is calculated based on the baseline model to obtain the current state deviation value;

[0173] Based on the real-time data sequence, the rainfall intensity characteristics and the rate of change of the upstream reservoir water level are extracted. If the rainfall intensity characteristics exceed the preset threshold and the rate of change of the water level shows an upward trend, the signal of rapid rise in the reservoir water level is marked.

[0174] The rate of change of soil moisture content is extracted from the real-time data sequence. If the rate of change of soil moisture content is close to saturation, the critical signal of soil saturation is marked.

[0175] Based on the detection results of the rapid rise in reservoir water level and the critical soil saturation signal, and combined with the deviation amplification to capture signal intensity, a comprehensive anomaly score is obtained.

[0176] Based on the comparison results of the comprehensive anomaly score and the thresholds of different levels, the corresponding graded response anomaly results are determined.

[0177] Furthermore, this embodiment collects multi-source historical data from a unified interface to form training samples. It is understood that this data includes records of multi-year rainfall, reservoir water levels, soil moisture content, and river flow, which are efficiently aggregated through a unified interface, avoiding the retrieval delays caused by previous decentralized storage.

[0178] In one possible implementation, a time series anomaly detection algorithm is used to process the training samples to establish a baseline model of the normal state.

[0179] Specifically, this embodiment uses isolated forest or autoencoder to train historical sequences and learn normal fluctuation patterns. For example, historical data shows that the average daily rainfall in summer is usually between 20-80 mm and the daily water level change rate does not exceed 0.5 meters, thereby constructing a baseline that reflects seasonal and regional patterns.

[0180] It should be noted that in this embodiment, the current state deviation value is obtained by calculating the deviation of the real-time data sequence through the baseline model.

[0181] For example, when real-time monitoring shows that rainfall reaches 150 mm for three consecutive days, the model calculates the Mahalanobis distance from the baseline mean and obtains a deviation value of 0.85. This quantitative deviation helps to detect potential flood signs early.

[0182] For example, in this embodiment, rainfall intensity characteristics and upstream reservoir water level change rate are extracted based on real-time data sequences.

[0183] In one embodiment, rainfall intensity is calculated in millimeters per hour and water level change rate is expressed in meters per hour. If the rainfall intensity exceeds 60 millimeters and the water level change rate is greater than 0.8 meters for two consecutive hours, a signal of rapid rise in reservoir water level is marked. This dual-condition judgment can effectively filter out false alarms caused by isolated heavy rainfall.

[0184] Specifically, the rate of change in soil moisture content is extracted from real-time data sequences. If the rate of change rises from 30% to over 85% within a short period, approaching saturation, a critical soil saturation signal is marked. This signal indicates a sharp increase in surface runoff, significantly accelerating the downstream flood accumulation rate.

[0185] For example, if a signal of rapid rise in reservoir water level or critical soil saturation is detected, the signal intensity is amplified by combining the deviation to obtain a comprehensive anomaly score.

[0186] In one possible implementation, the deviation value is multiplied by the signal strength as a weight. For example, the deviation value of 0.85 is multiplied by the reservoir rise signal strength of 1.2 to obtain a local score, which is then added to the soil signal to form a comprehensive score of 2.4. This fusion method makes anomaly detection more sensitive.

[0187] It should be noted that if the overall anomaly score exceeds the threshold for different levels, the corresponding level of response anomaly will be determined.

[0188] For example, a score of 1.5-2.5 triggers a yellow alert, necessitating enhanced monitoring; a score above 2.5 triggers an orange alert, initiating preparations for evacuation; and a score above 3.5 triggers a red alert, immediately initiating emergency response. This tiered mechanism allows for the rational allocation of resources based on the severity of the anomaly, avoiding waste and improving response timeliness. It ensures targeted measures are taken early in the flood's evolution, significantly reducing disaster losses.

[0189] Furthermore, the process of obtaining risk level classification criteria includes:

[0190] Based on the hierarchical response anomaly dataset, the risk probability values ​​were grouped using the K-means clustering algorithm to obtain multiple natural grouping centers.

[0191] Calculate the distance between adjacent centers based on the natural grouping centers to determine the location of the grouping boundaries;

[0192] The continuous risk probability is discretized based on the group boundary position to obtain the preliminary discrete risk value;

[0193] Based on the grouping range into which the initial discrete risk value falls, the corresponding risk level is assigned respectively, with the lowest grouping range assigned a low risk level and the highest grouping range assigned an extremely high risk level.

[0194] Based on the allocated set of risk levels, a risk level classification standard is formed.

[0195] For example, in this embodiment, the graded response anomaly data set is obtained by using anomaly events marked as different response levels in historical monitoring records, including rapid rise in reservoir water level, soil moisture saturation, and situations where the comprehensive anomaly score exceeds the threshold, forming a sample set containing multiple risk probability values.

[0196] Specifically, this embodiment uses the K-means clustering algorithm to group risk probability values, which can automatically discover the inherent natural structure of the data. Through iterative optimization, K-means groups samples with similar risk probability values ​​into the same cluster, ultimately obtaining multiple natural grouping centers. These centers represent typical clustering points of risk levels, helping to avoid subjective bias from manually setting thresholds, thus more objectively reflecting the actual risk distribution.

[0197] In one embodiment, assuming the risk probability values ​​range from 0 to 1, clustering yields four natural clustering centers: 0.15, 0.35, 0.60, and 0.85. This indicates that the data naturally forms four potential risk clusters: low, medium, high, and very high. Next, the distances between adjacent centers are calculated based on these natural clustering centers. For example, the distance between 0.35 and 0.15 is 0.20, the distance between 0.60 and 0.35 is 0.25, and the distance between 0.85 and 0.60 is 0.25. These distances reflect the strength of risk level transitions and can be used to determine the location of grouping boundaries. Typically, the median value of adjacent centers is taken as the boundary, such as boundaries set at 0.25, 0.475, and 0.725.

[0198] It should be noted that this embodiment discretizes the continuous risk probability through these grouping boundaries, which can transform the originally continuous probability values ​​into discrete categories, facilitating rapid decision-making in the future.

[0199] For example, a risk probability of 0.42 calculated in real time will fall within the range of 0.25 to 0.475, thus yielding a preliminary discrete risk value corresponding to a medium-risk category. This discretization process significantly improves the interpretability and response efficiency of the system, especially in scenarios where water levels change rapidly due to short-term heavy rainfall, enabling the timely conversion of continuous signals into clearly defined levels.

[0200] For example, if the initial discrete risk value falls within the lowest group range, such as a probability less than 0.25, it is assigned a low-risk level. This means that the current rise in upstream reservoir water level and changes in soil moisture content are still within normal fluctuations, and there is no need to immediately initiate advanced response measures. Conversely, if it falls within the highest group range, such as a probability greater than 0.725, it is assigned an extremely high-risk level. At this point, both the rapid rise in reservoir water level and the critical soil saturation signal may have been detected simultaneously, and the system must immediately trigger the highest-level warning to prevent the downstream flood disaster from escalating.

[0201] In one possible implementation, this embodiment uses a set of risk levels generated from multiple historical events to ultimately form a risk level classification standard. This standard can be solidified into a rule table, for example, low risk corresponds to a green alert, medium risk to yellow, high risk to orange, and extremely high risk to red. This standard is directly applied to the comprehensive anomaly score calculated from the deviation of real-time data, enabling automated triggering of graded responses, greatly improving the timeliness and accuracy of early warnings, while reducing the burden of manual judgment, and ensuring more reliable disaster prevention and control under the combined influence of upstream reservoirs and soil during heavy rainfall.

[0202] Example 2

[0203] Based on the same inventive concept, this embodiment also provides a flood risk identification system based on machine learning, including:

[0204] The progressive learning training module is used to obtain a time series segmented training framework based on historical flood data through progressive learning, and to obtain the results of fine-tuning and optimizing the model parameters.

[0205] The incremental learning optimization module is used to fine-tune the optimization results based on the model parameters. It absorbs new rainfall, river water level and terrain change data through incremental learning algorithm. When the deviation exceeds the preset threshold, the risk accumulation mechanism is activated to obtain the storage results of key elements.

[0206] The risk accumulation mechanism module is used to classify similar risk patterns based on the storage results of key elements and obtain a risk feature vector library through clustering algorithms.

[0207] The risk feature vector library construction module is used to construct a watershed connectivity analysis module based on the risk feature vector library and a graph neural network. The module uses rivers, tributaries and flood storage areas as nodes and the direction of water flow as edges for training to obtain flood evolution path prediction parameters.

[0208] The flood path prediction module is used to integrate real-time monitoring data from multiple data centers, satellite remote sensing images, and historical statistical information based on flood path prediction parameters, and process them through a data preprocessing pipeline to obtain unified interface training samples.

[0209] The multi-source data fusion preprocessing module is used to train samples based on a unified interface, establish a normal state baseline model through a time series anomaly detection algorithm, calculate the degree of deviation of real-time data, and obtain graded response anomaly results.

[0210] The risk level classification module is used to identify natural groupings of risk values ​​based on the results of graded response anomalies through cluster analysis, adjust the classification boundaries to discretize the probability of continuous risks, and obtain risk level classification standards; based on the risk level classification standards, flood risk is identified and flood risk levels are obtained.

[0211] The flood risk identification system based on machine learning provided in this embodiment has all the advantages of the flood risk identification method based on machine learning provided in Embodiment 1.

[0212] Example 3

[0213] This embodiment also discloses a computer device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method described in Embodiment 1.

[0214] Example 4

[0215] This embodiment also discloses a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the method described in Embodiment 1.

[0216] Example 5

[0217] This embodiment also discloses a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described in Embodiment 1.

[0218] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A flood risk identification method based on machine learning, characterized in that, include: Based on historical flood data, a time series segmented training framework is obtained through progressive learning, and the model parameters are fine-tuned and optimized. Based on the model parameter fine-tuning optimization results, new rainfall, river water level and terrain change data are absorbed through incremental learning algorithm. When the deviation exceeds the preset threshold, the risk sedimentation mechanism is activated to obtain the key element storage results. Based on the storage results of the key elements, similar risk patterns are classified using a clustering algorithm to obtain a risk feature vector library; Based on the risk feature vector library, a watershed connectivity analysis module is constructed using a graph neural network. Rivers, tributaries, and flood storage areas are used as nodes, and the direction of water flow is used as an edge for training to obtain flood evolution path prediction parameters. Based on the flood evolution path prediction parameters, real-time monitoring data from multiple data centers, satellite remote sensing images, and historical statistical information are integrated and processed through a data preprocessing pipeline to obtain unified interface training samples. Based on the training samples of the unified interface, a normal state baseline model is established through a time series anomaly detection algorithm, the degree of deviation of real-time data is calculated, and a graded response anomaly result is obtained. Based on the results of graded response anomalies, risk values ​​are naturally grouped using cluster analysis, and the classification boundaries are adjusted to discretize the probability of continuous risks, thereby obtaining the risk level classification criteria. Flood risk is identified and classified according to the aforementioned risk level classification criteria.

2. The method according to claim 1, characterized in that, The process of obtaining the model parameter fine-tuning optimization results includes: Historical flood data are sorted by year to obtain a time series; The time series is divided into multiple consecutive segments to obtain a segmented sequence; An initial training framework is constructed based on the segmented sequence to obtain basic model parameters; Input the event data for the current year in batches in sequence to obtain the updated dataset; Based on the updated dataset, progressive learning is performed on the base model parameters to obtain intermediate model parameters; Based on the intermediate model parameters, the event data for the next year is further processed to obtain the results of gradual parameter changes. Fine-tuning and optimization are performed based on the results of the gradual changes in the parameters to obtain the final model parameters.

3. The method according to claim 1, characterized in that, The process of obtaining key element storage results includes: Based on real-time rainfall data, water level data, and topographic change information obtained from multiple monitoring points, a structured environmental monitoring dataset was obtained after cleaning and processing. Based on the structured environmental monitoring dataset, an incremental learning method is used to gradually absorb new data to update the parameters and obtain the updated model prediction results. Based on the updated model prediction results, the deviation of rainfall data, water level data and topographic changes is calculated. If the deviation exceeds the preset threshold, the abnormal state judgment result is obtained. Based on the abnormal state determination results, the distribution of high-risk areas is analyzed and combined with rainfall thresholds and geological vulnerability information to obtain the results of potential risk area delineation; Based on the potential risk area division results, the critical rainfall threshold and geological vulnerability point data of each area are extracted to obtain the risk database storage results; Based on the results stored in the risk database, the data collection frequency is dynamically adjusted and the density of monitoring points in high-risk areas is increased to obtain subsequent monitoring optimization plans.

4. The method according to claim 1, characterized in that, The process of obtaining a risk feature vector library includes: Based on the key element data stored by the risk accumulation mechanism, the data is cleaned and formatted to obtain a preliminary organized dataset. Based on the initially compiled dataset, core elements are extracted and layered annotations are performed to obtain structured information units; Based on the structured information units, clustering methods are used to group similar patterns. If the similarity between data exceeds a preset threshold, they are classified into the same category to obtain the basic classification of risk patterns. Based on the basic classification of the risk patterns, key features are extracted for each category and corresponding feature vectors are generated to obtain a risk feature vector library. Based on the risk feature vector library, data preparation is carried out for subsequent matching scenarios. If the similarity between the input data and the feature vector reaches a preset standard, the corresponding risk category is output to obtain the matching result. Based on the matching results, the result data is archived into the knowledge base and the structured knowledge content is updated to obtain a dynamically updated risk information database. Based on the dynamically updated risk information database, newly input risk data is compared in real time. If the comparison results show that the new data is highly correlated with existing categories, it is classified into the corresponding category to obtain the latest risk classification status.

5. The method according to claim 1, characterized in that, The process of obtaining flood path prediction parameters includes: Transform the spatial distribution data of rivers, tributaries, and flood storage areas into structured information defined by nodes; Based on the structured information defined by the nodes, an edge relationship matrix is ​​constructed using water flow direction data to generate the connection topology within the watershed; Based on the aforementioned connection topology, a graph neural network is used to train the propagation mode of the flood peak transmission to obtain preliminary parameters of time decay. Based on the preliminary parameters of the time decay, the key path features of the flood peak transmission are extracted to determine the main direction of flood evolution; By comparing the main direction of flood evolution with historical records and correcting the path using spatial distribution data, the results of abnormal node identification are obtained. Based on the abnormal node determination results, the flood evolution path is updated in conjunction with the edge relationship matrix, and the final path prediction parameters are derived.

6. The method according to claim 1, characterized in that, The process of obtaining training samples for a unified interface includes: Based on real-time monitoring systems, satellite remote sensing platforms, and historical databases, we collect flood-related dynamic information, satellite remote sensing images, and historical statistical information to obtain a preliminary dataset. Based on the preliminary dataset, data cleaning tools are used to fill in missing values, and outliers exceeding a preset threshold range are corrected using nearest neighbor interpolation to obtain the cleaned dataset. Based on the cleaned dataset, spatiotemporal alignment technology is used to match data from different sources according to a unified timestamp and spatial grid to obtain an aligned dataset; Based on the aligned dataset, a scale normalization method is applied to convert data of different scales into a uniform range to obtain a standardized dataset. Based on the standardized dataset, features are extracted from key variables related to flood evolution and path prediction. If the correlation between features is lower than a preset threshold, irrelevant variables are removed to obtain a selected feature set. Based on the selected feature set, the pre-trained model is input to perform a flood path prediction task to obtain prediction result data; Based on the predicted data, structured training samples are generated and stored in a unified interface database.

7. The method according to claim 1, characterized in that, The process of obtaining graded response anomaly results includes: Training samples are generated by collecting historical data from multiple sources using a unified interface. Based on the training samples, a normal state baseline model is established using a time series anomaly detection algorithm; The deviation of the real-time data sequence is calculated based on the baseline model to obtain the current state deviation value; Based on the real-time data sequence, the rainfall intensity characteristics and the rate of change of the upstream reservoir water level are extracted. If the rainfall intensity characteristics exceed the preset threshold and the rate of change of the water level shows an upward trend, the signal of rapid rise in the reservoir water level is marked. The rate of change of soil moisture content is extracted from the real-time data sequence. If the rate of change of soil moisture content is close to saturation, the critical signal of soil saturation is marked. Based on the detection results of the rapid rise in reservoir water level and the critical soil saturation signal, and combined with the deviation amplification to capture signal intensity, a comprehensive anomaly score is obtained. Based on the comparison results between the comprehensive anomaly score and the thresholds of different levels, the corresponding graded response anomaly results are determined.

8. The method according to claim 1, characterized in that, The process of obtaining risk level classification criteria includes: Based on the hierarchical response anomaly dataset, the risk probability values ​​were grouped using the K-means clustering algorithm to obtain multiple natural grouping centers. Calculate the distance between adjacent centers based on the natural grouping centers to determine the grouping boundary positions; The continuous risk probability is discretized based on the grouping boundary position to obtain a preliminary discrete risk value; Based on the grouping range into which the preliminary discrete risk value falls, the corresponding risk level is assigned respectively, with the lowest grouping range assigned a low risk level and the highest grouping range assigned an extremely high risk level. Based on the allocated set of risk levels, a risk level classification standard is formed.

9. A flood risk identification system based on machine learning, characterized in that, include: The progressive learning training module is used to obtain a time series segmented training framework based on historical flood data through progressive learning, and to obtain the results of fine-tuning and optimizing the model parameters. The incremental learning optimization module is used to fine-tune the optimization results according to the model parameters. It absorbs new rainfall, river water level and terrain change data through incremental learning algorithm. When the deviation exceeds the preset threshold, the risk accumulation mechanism is activated to obtain the key element storage results. The risk accumulation mechanism module is used to classify similar risk patterns based on the storage results of the key elements and obtain a risk feature vector library through clustering algorithms. The risk feature vector library construction module is used to construct a watershed connectivity analysis module based on the risk feature vector library through a graph neural network. The module uses rivers, tributaries and flood storage areas as nodes and water flow direction as edges for training to obtain flood evolution path prediction parameters. The flood evolution path prediction module is used to integrate real-time monitoring data from multiple data centers, satellite remote sensing images, and historical statistical information based on the flood evolution path prediction parameters, and process them through a data preprocessing pipeline to obtain unified interface training samples. The multi-source data fusion preprocessing module is used to train samples based on the unified interface, establish a normal state baseline model through a time series anomaly detection algorithm, calculate the degree of deviation of real-time data, and obtain graded response anomaly results. The risk level classification module is used to identify natural groupings of risk values ​​based on the results of graded response anomalies through cluster analysis, adjust the classification boundaries to discretize the probability of continuous risks, and obtain risk level classification criteria; and to identify flood risks based on the risk level classification criteria to obtain flood risk levels.