A flood intelligent prediction method, system, device and medium based on concept-data driven coupling
A concept-data-driven flood forecasting model was constructed by combining a cloud-based clustering algorithm and the TFTformer model with an improved star-sparrow optimization algorithm. Error correction was performed using a spatiotemporal graph convolutional network, which solved the problem of insufficient parameter estimation in the Xin'anjiang model and enabled accurate and rapid flood forecasting and early warning.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- POWERCHINA HUADONG ENG CORP LTD
- Filing Date
- 2023-12-07
- Publication Date
- 2026-07-21
AI Technical Summary
The existing conceptual Xin'anjiang model lacks sufficient research on parameter estimation methods, resulting in low flood forecast accuracy. Furthermore, the challenges of spatiotemporal variability of hydrological models under changing environments have not been effectively addressed.
A cloud-based clustering algorithm is used to classify flood types. A concept-data-driven coupled flood forecasting model is constructed by combining the TFTformer model and the improved star-sparrow optimization algorithm. Error correction is performed through a spatiotemporal graph convolutional network, and the model parameters are dynamically adjusted to improve forecast accuracy.
It enables accurate and rapid flood forecasting, allowing for early flood warnings, reducing casualties and property losses, and improving the model's simulation accuracy and robustness.
Smart Images

Figure CN117874485B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of flood forecasting technology, specifically relating to a flood intelligent forecasting method, system, device, and medium based on concept-data driven coupling. Background Technology
[0002] Flood forecasting is a crucial research area in hydrology, playing a vital role in water resource management. By predicting rainfall and runoff, accurate water quantity information can be provided for agriculture, industry, and urban water supply. It also enables reservoir scheduling to rationally store water and allocate water resources downstream to meet urban water demand. Accurate and rapid flood forecasts allow decision-makers to take timely measures to allocate water resources. In responding to flood disasters, predicting peak flow and the evolution of flood events, as well as the potentially affected areas, allows for early flood warnings, emergency evacuations, reinforcement of temporary flood control facilities, and the implementation of emergency rescue measures. Timely and accurate flood warnings can minimize casualties and property damage.
[0003] Therefore, the spatiotemporal variability of meteorological and hydrological elements under changing environments poses new challenges to hydrological models. Flood forecasting methods are gradually gaining attention, and currently, conceptual watershed hydrological models are the most studied, influential, rapidly developing, and practically applied in the field of hydrology. A conceptual watershed hydrological model involves identifying actual hydrological patterns, specific research objects and objectives, searching for various factors influencing these patterns, distinguishing between primary and secondary factors, and then proposing hypotheses and generalizations. The goal is to establish a model that conforms to hydrological realities, with its structure and parameters having as clear a physical meaning as possible. Among these, the Xin'anjiang model, with its high forecasting accuracy and strong applicability, is widely used in rainfall-runoff simulation and forecasting in humid and semi-humid regions. The Xin'anjiang conceptual model divides runoff into surface runoff, interflow, and groundwater runoff, playing a crucial role in flood prediction. However, the parameter values in the Xin'anjiang conceptual model significantly impact simulation results. Research on parameter estimation methods for conceptual Xin'anjiang models is limited, while conceptual Xin'anjiang models are more commonly used in production units. Therefore, studying parameter estimation methods for conceptual Xin'anjiang models is of great significance. Based on the theory of the conceptual Xin'anjiang model and combined with optimization algorithms, this approach avoids the problem of unreasonable parameter values. This combination of concept and data exhibits high robustness, effectively avoiding interference from local extrema and improving the simulation accuracy of the model. Summary of the Invention
[0004] The first objective of this invention is to address the problems mentioned in the background art by proposing a concept-data driven coupling-based intelligent flood forecasting method to achieve accurate and rapid flood forecasting.
[0005] Therefore, the above-mentioned objective of the present invention is achieved through the following technical solution:
[0006] A concept-data-driven intelligent flood forecasting method includes the following steps:
[0007] S1: Collect historical datasets of the causes of floods, including flood duration, peak flow, runoff depth, surface rainfall, and surface evaporation.
[0008] S2: Using the collected historical datasets of floods, classify and categorize flood types using a cloud model-based clustering algorithm;
[0009] S3: Considering the characteristics of the decentralized Xin'anjiang conceptual model's sub-units and the information processing capabilities of the TFTformer model, the linear reservoir and the Muskingen River network confluence of the Xin'anjiang conceptual model are replaced by the TFTformer model to establish a concept-data driven coupled flood forecasting model.
[0010] S4: Concept-data driven coupled flood forecasting model parameter calibration. Given a fixed flood model structure, the model's flood process simulation performance depends on the model's parameter calibration. This is achieved by introducing an improved daemon optimization algorithm into the concept-data driven coupled flood forecasting model to dynamically adjust its hyperparameters.
[0011] S5: Use the historical dataset from step S1 as the model sample dataset, and divide it into training and test sets and feed it into the concept-data-driven coupled flood forecasting model to predict peak flow and runoff depth.
[0012] S6: Use a Spatiotemporal Graph Convolutional Network (STGCN) to correct the errors in the above prediction results and output the flood type for intelligent prediction.
[0013] While adopting the above technical solutions, the present invention may also adopt or combine the following technical solutions:
[0014] As a preferred embodiment of the present invention: in step S1, the data includes the duration of the flood. Peak flow runoff depth Surface rainfall evaporation from water surface .
[0015] As a preferred technical solution of the present invention: In step S2, the clustering algorithm based on the cloud model is used to classify flood types as follows:
[0016] A cloud-based clustering algorithm embeds clusters with stochastic uncertainty and analyzes the changes in concepts within the clusters. The cloud-based clustering algorithm consists of three parts:
[0017] (1) Initial concept generation: Initialize a concept space:
[0018] Duration of the flood Peak flow runoff depth Constructing includes A conceptual space , where each concept Described as a triple ( , , ).expect Representing the initial concept Entropy is a fundamental measure of determinism. It represents a measure of uncertainty in the initial concept, which is determined by the randomness and fuzziness of the concept.
[0019] hyperentropy Represents entropy Uncertainty measures. Unlike knowledge discovery, which extracts concepts from complex data, concepts effectively integrate randomness and fuzziness. Furthermore, concepts can make objects in the same group more similar than those in other clusters, and this can be achieved through three numerical features (…). , , Describes the overall quantitative properties of a concept.
[0020] Cloud-based clustering algorithms, through datasets including flood duration, Peak flow runoff depth Embedding random uncertainty in the data, grouping the dataset into... Under these concepts, a stable conceptual space is formed. .
[0021] (2) Concept-based refined aggregation involves repeatedly dividing data embedded with random uncertainty into new concepts and searching for the optimal concept space:
[0022] To partition data with a more precise concept, cloud-based clustering algorithms incorporate randomness into data partitioning, expanding the range of data distribution. This addresses the uncertainty of the clustering concept. It is a curve that defines how each point in the input space maps to an uncertain value between 0 and 1. Cloud-based clustering algorithms calculate the randomness contained in the data through the uncertainty of the clustering concept; randomness is determined by the value in the concept space. and Generate, represented as:
[0023]
[0024] in, It is randomly generated from (0,1). Then, the calculation is performed. The ambiguity of each data point in the dataset is defined in the concept, and randomness is embedded within it.
[0025] Uncertainty of the clustering concept:
[0026]
[0027] in, The time series length of the dataset. Representing concepts The m-th concept center.
[0028] In data partitioning, clustering concept uncertainty can be effectively evaluated. The concept uncertainty per data point in the dataset is considered. Each data point has a different concept uncertainty. The uncertainty is then normalized to find a better data partition. The result of the normalization is called the concept importance, and the formula for importance is defined as follows:
[0029]
[0030] According to the principle of maximum membership, data points The larger, the more conceptual Most likely to include Under the concept of conceptual space, data is divided into...
[0031] (3) Evaluation of conceptual uncertainty: The uncertainty of concepts during the iteration process is evaluated.
[0032] The goal of cloud-based clustering algorithms is to make objects within a concept have high similarity, while objects between concepts have low similarity.
[0033] Based on the conceptual importance of cluster center data points, concept centers... The calculation is as follows:
[0034]
[0035] in, For the concept Lower point The importance of the concept.
[0036]
[0037] Conceptual selection of data points The calculation is as follows:
[0038]
[0039] Conceptual parameters The calculation is as follows:
[0040]
[0041] in, and They are respectively Mean and variance.
[0042] By calculating clusters of , and Generate concepts . Reflects the concept Discreteness of the midpoint Reflects the concept The degree of clustering at the midpoint.
[0043] Concept-based refined aggregation leverages the fuzziness and randomness of objects to retain uncertain information, resulting in more refined concepts in clustering. The function can be composed of a dataset and concepts C as follows:
[0044]
[0045] This indicates that the dataset belongs to the concept. The probability of.
[0046]
[0047] in, Indicates that the concept center belongs to the concept The probability of. In a concept space, a point represents a concept. The probability of.
[0048] Based on the dataset, flood duration Peak flow runoff depth In concept The data is divided into four flood types based on probability. The data is then fed into a concept-data-driven coupled flood forecasting model optimized using an improved sparrow optimization algorithm. The model outputs four flood forecast types. The optimal forecast is determined by calculating the reference distance between each of the four forecast results and its respective concept center. The formula for the reference distance is as follows:
[0049]
[0050] in, Time series length, To predict the duration of the flood using the model, For the duration of the flood at the concept center, To predict peak flood flow using the model, For the peak flow of the concept center, To predict runoff depth using the model, The runoff depth at the concept center.
[0051] As a preferred technical solution of the present invention: In step S3, the process of constructing the concept-data driven coupled flood forecasting model is as follows:
[0052] The Xin'anjiang conceptual model is a typical decentralized conceptual model, characterized by its three-part division: unit-based, water source-based, and phase-based. The unit-based division considers the uneven distribution of rainfall, dividing the entire basin into several unit basins based on the river network structure, and performing separate runoff generation and collection calculations for each unit basin. The water source-based division categorizes runoff into surface runoff, interflow, and groundwater runoff based on their collection velocities, with surface runoff being the fastest and groundwater runoff the slowest.
[0053] The concept-data-driven coupled model computation mainly includes the following three modules:
[0054] (1) Evapotranspiration calculation: Soil evapotranspiration is divided into upper layer, lower layer and deep layer. The evapotranspiration rate is calculated using a three-layer evapotranspiration model. The parameters include the average tension water capacity of the watershed. upper tension water capacity Lower layer tension water capacity Deep tension water capacity Evapotranspiration conversion factor and deep evaporation diffusion coefficient The calculation formula is as follows:
[0055]
[0056]
[0057]
[0058]
[0059] in, This represents the total tension water storage capacity; This refers to the upper tension water storage capacity; This refers to the storage capacity of tension water in the lower layer; For deep tensile water storage; This represents the total evaporation rate; This refers to the evaporation rate of the upper layer. This refers to the evaporation rate of the lower layer. This refers to the deep evaporation rate; P represents evapotranspiration capacity, and P represents surface rainfall.
[0060] like ,but:
[0061]
[0062] like ,but:
[0063]
[0064] like ,but:
[0065]
[0066] like but:
[0067]
[0068] like but:
[0069]
[0070] The parameters to be optimized are and .
[0071] (2) Calculation of runoff: Full runoff mode, the formula is as follows:
[0072]
[0073] In the formula: The area of runoff generation; The drainage area; This refers to the water storage capacity at a single point within the watershed. This represents the maximum water storage capacity at a single point in the basin. The index of the water storage capacity-area distribution curve reflects the uneven distribution of water storage capacity in the vadose zone of the watershed. A larger value indicates greater unevenness. The smaller the value, the more uniform it is. When it is 0, it means that the water storage capacity of the vadose zone in the watershed is uniform and unchanged. The parameter to be optimized is B.
[0074] (3) Water source division: The water source division of the free water reservoir is matched, and the water source is divided into three types of runoff: surface runoff, interflow, and groundwater.
[0075]
[0076] In the formula: The free water storage capacity at a single point in the basin; It represents the maximum free water storage capacity at a single point in the basin. The power of the watershed free water storage capacity-area distribution curve.
[0077] Surface runoff RS:
[0078]
[0079] Rangzhongliu RI:
[0080]
[0081] Subsurface runoff RG:
[0082]
[0083] In the formula: and These represent the ratio of runoff area at the beginning and end of the time period, respectively. Mean free water storage capacity; The flow rate is the generated flow velocity. The parameters to be optimized in the water source calculation include: , , and .
[0084] The surface rainfall in step 1 evaporation from water surface As input to the concept-data driven coupled model, the runoff sequence on permeable and impermeable areas is calculated. The runoff from permeable areas is divided into non-runoff areas and runoff areas. The non-runoff areas include: tension water, upper layer, lower layer and deep layer, and the evaporation of these three layers is calculated to form interflow. The runoff areas form surface free water, which constitutes groundwater runoff. The runoff from impermeable areas and tension water constitute surface runoff.
[0085] The datasets of surface runoff, interflow, and total groundwater inflow obtained through the conceptual model are fed into the TFTformer model below for training and prediction. The specific details are as follows:
[0086] The TFTformer model is a novel attention-based architecture that combines high-performance multi-view prediction with interpretable insights into temporal dynamics. To learn temporal relationships at different scales, it employs recurrent layers for local processing, interpretable self-attention layers for long-term dependencies, and leverages specialized components to select relevant features and a series of gating layers to suppress unwanted components, thereby achieving high performance across a wide range of scenarios.
[0087] The TFTformer model mainly has the following structure:
[0088] (1) Gating mechanism: It can skip any unused components in the architecture, providing adaptive depth and network complexity to adapt to a wide range of datasets and scenarios.
[0089] (2) Variable selection network layer: Select relevant input variables at each time step.
[0090] (3) Static covariate encoder layer: integrates static features into the network and adjusts the temporal dynamics by encoding the context vector.
[0091] (4) Temporal processing layer: Learns long-term and short-term temporal relationships from observed and known time-varying inputs. Sequence layers are used for local processing, while long-term dependencies are captured using a novel interpretable multi-head attention block.
[0092] (5) Output layer: The prediction interval of the possible target value range of each prediction layer is determined by quantile prediction.
[0093] As a preferred technical solution of the present invention: In step S4, the parameter calibration of the concept-data driven coupled flood forecasting model is specifically as follows:
[0094] Given a fixed structure for a smart flood forecasting model, the model's simulation performance during flood events depends on its parameter calibration. By introducing an improved sparrow optimization algorithm into the concept-data driven coupled model to dynamically adjust its hyperparameters, the reliability of the simulation results can be effectively improved.
[0095] An improved star-sparrow optimization algorithm, NOA, is introduced to initialize the population. The following objective function is defined as the fitness function of the population:
[0096]
[0097] in, , and These are hyperparameters that control the weights of each term in the fitness function and can be adjusted during training. This represents the relative error of the flood peak; a result close to 0 indicates a more accurate flood peak forecast. For predictive evaluation indicators, the closer the result is to 0, the higher the predictive indicator. The formula is:
[0098]
[0099] in, This represents the simulated value of the concept-data driven coupled model after inputting relevant data. Represents the actual observed value. This represents the average of the actual observed values;
[0100] It mainly includes two strategies:
[0101] (1) Simulate the food collection and storage strategy of the starbird in summer and autumn, introduce the golden sine algorithm to traverse all values of the sine function, and at the same time, update the location of the food to greatly improve the search speed. The specific formula is as follows:
[0102] Exploration phase:
[0103]
[0104] in, For the current generation The new position of the star sparrow; For the current generation The first star sparrow One location; and These are the upper and lower bounds in an optimization problem; It is a random number; , and These are three different indicators randomly selected from the population to explore high-quality food sources; , A random real number in the range [0,1]. It is the j-th dimension mean of all solutions in the current population during the t-th iteration; µ is a number randomly generated between 0 and 1 based on a normal distribution.
[0105] The development phase of introducing the golden sine algorithm:
[0106]
[0107]
[0108] in, The current iteration The new location of Zhongxingque's storage area , A random real number in the range [0,1]. It is a factor that linearly decreases in development behavior from 1 to 0; It is the optimal solution; and It is a random number. This determines the distance the starbird will move in the next iteration. ; Determine the direction of position update for the next iteration. ; and The coefficients are obtained through the golden ratio.
[0109] (2) Simulate the search and retrieval strategy of the star sparrow in spring and winter, i.e., cache search and retrieval strategy. Improve the cache search and retrieval strategy by using Levy flight to enhance the star sparrow's global search capability and avoid the algorithm getting trapped in local optima. The specific formula is as follows:
[0110] Introducing Levy Flight cache search:
[0111]
[0112]
[0113]
[0114]
[0115] in, The search behavior mechanism of the star sparrow. The ratio of the current iteration number to the total number of iterations. Power; The step size is random. Dot product; To comply with Constraints under distributed random search paths; and It follows a standard normal distribution; =1.5; This is the search magnification factor; It is a gamma function;
[0116] Introducing the Levy flight recovery strategy:
[0117]
[0118]
[0119] in, Restore the behavior mechanism of the starbird.
[0120] in, To restore the behavior mechanism of the star-sparrow. This optimization algorithm is used to optimize parameters in the above concept-data-driven coupling, such as the parameters to be optimized in evaporation calculations. and The parameter to be optimized in the runoff calculation is B; the parameters to be optimized in the water source calculation include: , , and The learning rate and number of hidden layers in the TFTformer model and the following spatiotemporal graph convolutional network STGCN are used to obtain the parameters of the concept-data driven coupling model suitable for this watershed.
[0121] As a preferred technical solution of the present invention: In step S6, a spatiotemporal graph convolutional network (STGCN) is used to correct the error of the above prediction results, as follows:
[0122] Due to limitations imposed by factors such as dataset size, model structure parameters, algorithm convergence, and robustness during the forecasting process, the forecasting accuracy of concept-data-driven coupled flood forecasting models needs further improvement. Therefore, real-time correction of flood forecasting errors is necessary to enhance the accuracy and practicality of the forecasting model. Real-time correction, based on real-time information during the forecasting process, applies modern information technology theories and methods to dynamically adjust model parameters, model inputs, and forecast results, constructing a real-time feedback mechanism between the forecasting model and the correction model to reduce flood forecasting errors. This is a crucial component of real-time flood forecasting.
[0123] Spatiotemporal graph convolutional networks contain the following structure:
[0124] (1) Input layer: based on concept-data driven coupled flood forecasting with surface rainfall evaporation from water surface Flood forecasting is performed using forecasting factors; therefore, the input layer of this model uses concept-data driven methods to couple the predicted values of the flood forecasting model. and observed values Error sequence between The set of rainfall sequences at control stations in a sub-basin is defined as follows: As input to the spatiotemporal graph convolutional network, the flow lengths from relevant stations within the watershed to the target station are used to construct the STGCN adjacency matrix. The construction, training set, and test set partitioning. The calculation formula is as follows:
[0125]
[0126] in, This indicates the runoff length between two stations.
[0127] (2) Network Layers: Training a mapping function that reflects the features of the error sequence. First, two ST-Conv blocks are constructed, consisting of a gated sequence convolutional layer and a spatial sequence convolutional layer, to extract the temporal and local spatial features of the error sequence. The relationship between the simulated error values output by the STGCN layer and the input layer data is as follows:
[0128]
[0129] in, express The concept of time - the simulated error value obtained from the data-driven coupled model; yes Time error sequence; , ... yes time Rainfall at several relevant rain gauge stations; That is, the mapping function determined by the concept-data driven coupling model.
[0130] (63) Output layer: Outputs error correction results. Calculates the error. The simulation error obtained in step (2) Difference between and initialize its minimum value. ;Compare , Size and ( and )and( and The results are compared using common evaluation metrics such as root mean square error, coefficient of determination, and mean absolute error. If the evaluation metrics show improvement, the simulated error values are inversely normalized before being output; finally, the original predicted values from the forecast model are used. and final simulation error The model correction results were calculated. The relationship is as follows:
[0131]
[0132] The second objective of this invention is to provide a smart flood forecasting system based on concept-data driven coupling.
[0133] Therefore, the above-mentioned objective of the present invention is achieved through the following technical solution:
[0134] A concept-data driven intelligent flood forecasting system includes the following modules:
[0135] - Historical data collection module, which is used to collect historical datasets that lead to floods, including flood duration, peak flow, runoff depth, surface rainfall and water surface evaporation;
[0136] - Flood type classification and grading module, which is used to classify and grade flood types using a cloud model-based clustering algorithm based on the historical data collection module that leads to floods;
[0137] - Concept-data driven coupled flood forecasting model establishment module: The concept-data driven coupled flood forecasting model establishment module is used to consider the characteristics of the decentralized Xin'anjiang concept model sub-units and the information processing capabilities of the TFTformer model, and to replace the linear reservoirs and the Muskingen river network confluence of the Xin'anjiang concept model with the TFTformer model to establish a concept-data driven coupled flood forecasting model;
[0138] - Concept-data driven coupled flood forecasting model parameter calibration module, which is used to dynamically adjust the hyperparameters of the concept-data driven coupled flood forecasting model by introducing an improved star-sparrow optimization algorithm;
[0139] - Data prediction module, which uses the historical dataset obtained by the historical data collection module as the model sample dataset, and divides it into training set and test set and feeds it into the concept-data driven coupled flood forecasting model to predict peak flow and runoff depth;
[0140] - Flood type intelligent prediction module, which uses a spatiotemporal graph convolutional network (STGCN) to correct errors in the above prediction results and outputs the flood type for intelligent prediction.
[0141] The third objective of this invention is to provide a smart flood forecasting device based on concept-data driven coupling.
[0142] Therefore, the above-mentioned objective of the present invention is achieved through the following technical solution:
[0143] A concept-data-driven intelligent flood forecasting device includes:
[0144] - At least one processor;
[0145] - At least one memory for storing at least one computer program;
[0146] The processor executes a computer program in memory to implement the steps of the concept-data-driven coupling-based intelligent flood forecasting method as described above.
[0147] Another objective of this invention is to provide a computer storage medium.
[0148] Therefore, the above-mentioned objective of the present invention is achieved through the following technical solution:
[0149] A computer storage medium storing a computer program, the computer program being executed by a computer to implement the steps of the concept-data-driven coupling-based intelligent flood forecasting method as described above.
[0150] This invention provides a concept-data driven coupled intelligent flood forecasting method, system, device, and medium. First, a cloud model-based clustering algorithm is used to cluster historical flood events to obtain flood types. Addressing the sensitivity of the conceptual hydrological model's fitting accuracy to confluence parameters, the TFTformer model replaces the linear reservoir and Muskingen river network confluence of the Xin'anjiang conceptual model, constructing a concept-data driven coupled flood forecasting model to predict peak flow and runoff depth. An evolutionary algorithm is used to calibrate the model's parameters. Furthermore, to address the inherent weakness of optimization algorithms in getting trapped in local optima, the Lévy fly-through algorithm and the golden sine algorithm are introduced for improvement. Finally, a spatiotemporal graph convolutional network is used to correct the errors in the prediction results and output the flood type for intelligent forecasting. This invention can predict peak flow and flood events, issue early flood warnings, urgently evacuate people, reinforce temporary flood control facilities, and take emergency rescue measures to minimize casualties and property losses.
[0151] Specifically, compared with the prior art, the present invention has the following beneficial effects:
[0152] 1) Considering the characteristics of the Xin'anjiang conceptual model's sub-units and the ability of the TFTformer model to process time series, the TFTformer model is used to replace the linear reservoir and the Muskingen river network confluence of the Xin'anjiang conceptual model.
[0153] 2) In view of the complexity of flood forecasting, a method is proposed to use cloud model cluster analysis of historical flood events to classify floods, accurately identify flood types after model prediction, and take corresponding countermeasures.
[0154] 3) Introducing an improved star-sparrow optimization algorithm to optimize the hidden layer dimension and learning rate of the concept-data driven coupled model can accelerate the model convergence speed, improve the model training efficiency, and further improve the prediction accuracy.
[0155] 4) An error correction strategy based on spatiotemporal graph convolutional networks is proposed to adjust the model prediction results. This strategy can correct possible biases and provide more reliable prediction results. Attached Figure Description
[0156] Figure 1 This is an overall flowchart of the intelligent flood forecasting method based on concept-data driven coupling provided by the present invention.
[0157] Figure 2 This is a flowchart for classifying flood types based on cloud model clustering.
[0158] Figure 3 The flowchart for a concept-data driven coupled flood forecasting model.
[0159] Figure 4 A flowchart for optimizing the concept of the Star Sparrow algorithm - data-driven coupling.
[0160] Figure 5 This is a flowchart for error correction in the spatiotemporal graph convolutional network STGCN. Detailed Implementation
[0161] The present invention will now be described in further detail with reference to the accompanying drawings.
[0162] A concept-data driven intelligent flood forecasting method includes the following steps:
[0163] S1: Collect historical datasets of the causes of floods, including flood duration, peak flow, runoff depth, surface rainfall, and surface evaporation.
[0164] S2: As Figure 2 As shown, the cloud model-based clustering algorithm embeds clusters with random uncertainty and analyzes the changes in concepts within the clusters. The cloud model-based clustering algorithm consists of three parts:
[0165] (1) Initial concept generation: Initialize a concept space:
[0166] Duration of the flood Peak flow runoff depth Constructing includes A conceptual space , where each concept Described as a triple ( , , ).expect Representing the initial concept Entropy is a fundamental measure of determinism. It represents a measure of uncertainty in the initial concept, which is determined by the randomness and fuzziness of the concept.
[0167] hyperentropy Represents entropy Uncertainty measures. Unlike knowledge discovery, which extracts concepts from complex data, concepts effectively integrate randomness and fuzziness. Furthermore, concepts can make objects in the same group more similar than those in other clusters, and this can be achieved through three numerical features (…). , , Describes the overall quantitative properties of a concept.
[0168] Cloud-based clustering algorithms, through datasets including flood duration, Peak flow runoff depth Embedding random uncertainty in the data, grouping the dataset into... Under these concepts, a stable conceptual space is formed. .
[0169] (2) Concept-based refined aggregation involves repeatedly dividing data embedded with random uncertainty into new concepts and searching for the optimal concept space:
[0170] To partition data with a more precise concept, cloud-based clustering algorithms incorporate randomness into data partitioning, expanding the range of data distribution and reducing the uncertainty of the clustering concept. It is a curve that defines how each point in the input space maps to an uncertain value between 0 and 1. Cloud-based clustering algorithms calculate the randomness contained in the data through the uncertainty of the clustering concept. Randomness is determined by the number of points in the concept space. and Generate, represented as:
[0171]
[0172] in, It is randomly generated from (0,1). Then, the calculation is performed. The ambiguity of each data point in the dataset is defined in the concept, and randomness is embedded within it.
[0173] Uncertainty of the clustering concept:
[0174]
[0175] in, The time series length of the dataset. Representing concepts The m-th concept center.
[0176] In data partitioning, clustering concept uncertainty can be effectively evaluated. The concept uncertainty per data point in the dataset is considered. Each data point has a different concept uncertainty. The uncertainty is then normalized to find a better data partition. The result of the normalization is called the concept importance, and the formula for importance is defined as follows:
[0177]
[0178] According to the principle of maximum membership, data points The larger, the more conceptual Most likely to include Under the concept of conceptual space, data is divided into...
[0179] (3) Evaluation of conceptual uncertainty: The uncertainty of concepts during the iteration process is evaluated.
[0180] The goal of cloud-based clustering algorithms is to make objects within a concept have high similarity, while objects between concepts have low similarity.
[0181] Based on the conceptual importance of cluster center data points, concept centers... The calculation is as follows:
[0182]
[0183] in, For the concept Lower point The importance of the concept.
[0184]
[0185] Conceptual selection of data points The calculation is as follows:
[0186]
[0187] Conceptual parameters The calculation is as follows:
[0188]
[0189] in, and They are respectively Mean and variance.
[0190] By calculating clusters of , and Generate concepts . Reflects the concept Discreteness of the midpoint Reflects the concept The degree of clustering at the midpoint.
[0191] Concept-based refined aggregation leverages the fuzziness and randomness of objects to retain uncertain information, resulting in more refined concepts in clustering. The function can be composed of a dataset and concepts C as follows:
[0192]
[0193] This indicates that the dataset belongs to the concept. The probability of.
[0194]
[0195] in, Indicates that the concept center belongs to the concept The probability of. In a concept space, a point represents a concept. The probability of.
[0196] Based on the dataset, flood duration Peak flow runoff depth In concept The data is divided into four flood types based on probability. The data is then fed into a concept-data-driven coupled flood forecasting model optimized using an improved sparrow optimization algorithm. The model outputs four flood forecast types. The optimal forecast is determined by calculating the reference distance between each of the four forecast results and its respective concept center. The formula for the reference distance is as follows:
[0197]
[0198] in, Time series length, To predict the duration of the flood using the model, For the duration of the flood at the concept center, To predict peak flood flow using the model, For the peak flow of the concept center, To predict runoff depth using the model, The runoff depth at the concept center.
[0199] S3: As Figure 3 As shown, a concept-data driven coupled flood forecasting model is constructed, as follows:
[0200] The Xin'anjiang conceptual model is a typical decentralized conceptual model, characterized by its three-part division: unit-based, water source-based, and phase-based. The unit-based division considers the uneven distribution of rainfall, dividing the entire basin into several unit basins based on the river network structure, and performing separate runoff generation and collection calculations for each unit basin. The water source-based division categorizes runoff into surface runoff, interflow, and groundwater runoff based on their collection velocities, with surface runoff being the fastest and groundwater runoff the slowest.
[0201] The concept-data-driven coupled model computation mainly includes the following three modules:
[0202] (1) Evapotranspiration calculation: Soil evapotranspiration is divided into upper layer, lower layer and deep layer. The evapotranspiration rate is calculated using a three-layer evapotranspiration model. The parameters include the average tension water capacity of the watershed. upper tension water capacity Lower layer tension water capacity Deep tension water capacity Evapotranspiration conversion factor and deep evaporation diffusion coefficient The calculation formula is as follows:
[0203]
[0204]
[0205]
[0206]
[0207] in, This represents the total tension water storage capacity; This refers to the upper tension water storage capacity; This refers to the storage capacity of tension water in the lower layer; For deep tensile water storage; This represents the total evaporation rate; This refers to the evaporation rate of the upper layer. This refers to the evaporation rate of the lower layer. This refers to the deep evaporation rate; P represents evapotranspiration capacity, and P represents surface rainfall.
[0208] like ,but:
[0209]
[0210] like ,but:
[0211]
[0212] like ,but:
[0213]
[0214] like but:
[0215]
[0216] like but:
[0217]
[0218] The parameters to be optimized are and .
[0219] (2) Calculation of runoff: Full runoff mode, the formula is as follows:
[0220]
[0221] In the formula: The area of runoff generation; The drainage area; This refers to the water storage capacity at a single point within the watershed. This represents the maximum water storage capacity at a single point in the basin. The index of the water storage capacity-area distribution curve reflects the uneven distribution of water storage capacity in the vadose zone of the watershed. A larger value indicates greater unevenness. The smaller the value, the more uniform it is. When it is 0, it means that the water storage capacity of the vadose zone in the watershed is uniform and unchanged. The parameter to be optimized is B.
[0222] (3) Water source division: The water source division of the free water reservoir is matched, and the water source is divided into three types of runoff: surface runoff, interflow, and groundwater.
[0223]
[0224] In the formula: The free water storage capacity at a single point in the basin; It represents the maximum free water storage capacity at a single point in the basin. The power of the watershed free water storage capacity-area distribution curve.
[0225] Surface runoff RS:
[0226]
[0227] Rangzhongliu RI:
[0228]
[0229] Subsurface runoff RG:
[0230]
[0231] In the formula: and These represent the ratio of runoff area at the beginning and end of the time period, respectively. Mean free water storage capacity; The flow rate is the generated flow velocity. The parameters to be optimized in the water source calculation include: , , and .
[0232] The surface rainfall in step S1 evaporation from water surface As input to the concept-data driven coupled model, the runoff sequence on permeable and impermeable areas is calculated. The runoff from permeable areas is divided into non-runoff areas and runoff areas. The non-runoff areas include: tension water, upper layer, lower layer, and deep layer, and the evaporation of these three layers is calculated to constitute interflow. The runoff areas form surface free water, which constitutes groundwater runoff. The runoff from impermeable areas and tension water constitute surface runoff.
[0233] The datasets of surface runoff, interflow, and total groundwater inflow obtained through the conceptual model are fed into the TFTformer model below for training and prediction. The specific details are as follows:
[0234] The TFTformer model is a novel attention-based architecture that combines high-performance multi-view prediction with interpretable insights into temporal dynamics. To learn temporal relationships at different scales, it employs recurrent layers for local processing, interpretable self-attention layers for long-term dependencies, and leverages specialized components to select relevant features and a series of gating layers to suppress unwanted components, thereby achieving high performance across a wide range of scenarios.
[0235] The TFTformer model mainly has the following structure:
[0236] (1) Gating mechanism: It can skip any unused components in the architecture, providing adaptive depth and network complexity to adapt to a wide range of datasets and scenarios.
[0237] (2) Variable selection network layer: Select relevant input variables at each time step.
[0238] (3) Static covariate encoder layer: integrates static features into the network and adjusts the temporal dynamics by encoding the context vector.
[0239] (4) Temporal processing layer: Learns long-term and short-term temporal relationships from observed and known time-varying inputs. Sequence layers are used for local processing, while long-term dependencies are captured using a novel interpretable multi-head attention block.
[0240] (5) Output layer: The prediction interval of the possible target value range of each prediction layer is determined by quantile prediction.
[0241] S4: Concept-data driven coupled flood forecasting model parameter calibration. Given a fixed flood model structure, the model's flood process simulation performance depends on the model's parameter calibration. This is achieved by introducing an improved daemon optimization algorithm into the concept-data driven coupled flood forecasting model to dynamically adjust its hyperparameters.
[0242] like Figure 4 As shown, the details are as follows:
[0243] Given a fixed structure for a smart flood forecasting model, the model's simulation performance during flood events depends on its parameter calibration. By introducing an improved sparrow optimization algorithm into the concept-data driven coupled model to dynamically adjust its hyperparameters, the reliability of the simulation results can be effectively improved.
[0244] An improved star-sparrow optimization algorithm, NOA, is introduced to initialize the population. The following objective function is defined as the fitness function of the population:
[0245]
[0246] in, , and These are hyperparameters that control the weights of each term in the fitness function and can be adjusted during training. This represents the relative error of the flood peak; a result close to 0 indicates a more accurate flood peak forecast. For predictive evaluation indicators, the closer the result is to 0, the higher the predictive indicator. The formula is:
[0247]
[0248] in, This represents the simulated value of the concept-data driven coupled model after inputting relevant data. Represents the actual observed value. This represents the average of the actual observed values;
[0249] It mainly includes two strategies:
[0250] (1) Simulate the food collection and storage strategy of the starbird in summer and autumn, introduce the golden sine algorithm to traverse all values of the sine function, and at the same time, update the location of the food to greatly improve the search speed. The specific formula is as follows:
[0251] Exploration phase:
[0252]
[0253] in, For the current generation The new position of the star sparrow; For the current generation The first star sparrow One location; and These are the upper and lower bounds in an optimization problem; It is a random number; , and These are three different indicators randomly selected from the population to explore high-quality food sources; , A random real number in the range [0,1]. It is the j-th dimension mean of all solutions in the current population during the t-th iteration; µ is a number randomly generated between 0 and 1 based on a normal distribution.
[0254] The development phase of introducing the golden sine algorithm:
[0255]
[0256]
[0257] in, The current iteration The new location of Zhongxingque's storage area , A random real number in the range [0,1]. It is a factor that linearly decreases in development behavior from 1 to 0; It is the optimal solution; and It is a random number. This determines the distance the starbird will move in the next iteration. ; Determine the direction of position update for the next iteration. ; and The coefficients are obtained through the golden ratio.
[0258] (2) Simulate the search and retrieval strategy of the starbird in spring and winter to search for and retrieve storage locations, i.e., cache search and retrieval strategy. Improve the cache search and retrieval strategy by using Levy flight to enhance the starbird's global search capability and avoid the algorithm getting trapped in local optima. The specific formula is as follows:
[0259] Introducing Levy flight cache search:
[0260]
[0261]
[0262]
[0263]
[0264] in, The search behavior mechanism of the star sparrow. The ratio of the current iteration number to the total number of iterations. Power; The step size is random. Dot product; To comply with Constraints under distributed random search paths; and It follows a standard normal distribution; =1.5; This is the search magnification factor; It is a gamma function;
[0265] Introducing the Levy flight recovery strategy:
[0266]
[0267]
[0268] in, To restore the behavior mechanism of the star-sparrow. This optimization algorithm is used to optimize parameters in the above concept-data-driven coupling, such as the parameters to be optimized in evaporation calculations. and The parameter to be optimized in the runoff calculation is B; the parameters to be optimized in the water source calculation include: , , and The learning rate and number of hidden layers in the TFTformer model and the following spatiotemporal graph convolutional network STGCN are used to obtain the parameters of the concept-data driven coupling model suitable for this watershed.
[0269] S5: Use the historical dataset from step S1 as the model sample dataset, and divide it into training and test sets and feed it into the concept-data-driven coupled flood forecasting model to predict peak flow and runoff depth.
[0270] S6: A Spatiotemporal Graph Convolutional Network (STGCN) is used to correct errors in the above prediction results and outputs the flood type for intelligent prediction. For example... Figure 5 As shown, the details are as follows:
[0271] Due to limitations imposed by factors such as dataset size, model structure parameters, algorithm convergence, and robustness during the forecasting process, the forecasting accuracy of concept-data-driven coupled flood forecasting models needs further improvement. Therefore, real-time correction of flood forecasting errors is necessary to enhance the accuracy and practicality of the forecasting model. Real-time correction, based on real-time information during the forecasting process, applies modern information technology theories and methods to dynamically adjust model parameters, model inputs, and forecast results, constructing a real-time feedback mechanism between the forecasting model and the correction model to reduce flood forecasting errors. This is a crucial component of real-time flood forecasting.
[0272] Spatiotemporal graph convolutional networks contain the following structure:
[0273] (1) Input layer: based on concept-data driven coupled flood forecasting with surface rainfall evaporation from water surface Flood forecasting is performed using forecasting factors; therefore, the input layer of this model uses concept-data driven methods to couple the predicted values of the flood forecasting model. and observed values Error sequence between The set of rainfall sequences at control stations in a sub-basin is defined as follows: As input to the spatiotemporal graph convolutional network, the flow lengths from relevant stations within the watershed to the target station are used to construct the STGCN adjacency matrix. The construction, training set, and test set partitioning. The calculation formula is as follows:
[0274]
[0275] in, This indicates the runoff length between two stations.
[0276] (2) Network Layers: Training a mapping function that reflects the features of the error sequence. First, two ST-Conv blocks are constructed, consisting of a gated sequence convolutional layer and a spatial sequence convolutional layer, to extract the temporal and local spatial features of the error sequence. The relationship between the simulated error value output by the STGCN layer and the input layer data is as follows:
[0277]
[0278] in, express The concept of time - the simulated error value obtained from the data-driven coupled model; yes Time error sequence; , ... yes time Rainfall at several relevant rain gauge stations; That is, the mapping function determined by the concept-data driven coupling model.
[0279] (3) Output layer: Outputs error correction results. Calculates the error. The simulation error obtained in step (2) Difference between and initialize its minimum value. ;Compare , Size and ( and )and( and The results are compared using common evaluation metrics such as root mean square error, coefficient of determination, and mean absolute error. If the evaluation metrics show improvement, the simulated error values are inversely normalized before being output; finally, the original predicted values from the forecast model are used. and final simulation error The model correction results were calculated. The relationship is as follows:
[0280]
[0281] This invention also provides a concept-data driven intelligent flood forecasting system, comprising the following modules:
[0282] - Historical data collection module, which is used to collect historical datasets that lead to floods, including flood duration, peak flow, runoff depth, surface rainfall and water surface evaporation;
[0283] - Flood type classification and grading module, which is used to classify and grade flood types using a cloud model-based clustering algorithm based on the historical data collection module that leads to floods;
[0284] - Concept-data driven coupled flood forecasting model establishment module: The concept-data driven coupled flood forecasting model establishment module is used to consider the characteristics of the decentralized Xin'anjiang concept model sub-units and the information processing capabilities of the TFTformer model, and to replace the linear reservoirs and the Muskingen river network confluence of the Xin'anjiang concept model with the TFTformer model to establish a concept-data driven coupled flood forecasting model;
[0285] - Concept-data driven coupled flood forecasting model parameter calibration module, which is used to dynamically adjust the hyperparameters of the concept-data driven coupled flood forecasting model by introducing an improved star-sparrow optimization algorithm;
[0286] - Data prediction module, which uses the historical dataset obtained by the historical data collection module as the model sample dataset, and divides it into training set and test set and feeds it into the concept-data driven coupled flood forecasting model to predict peak flow and runoff depth;
[0287] - Flood type intelligent prediction module, which uses a spatiotemporal graph convolutional network (STGCN) to correct errors in the above prediction results and outputs the flood type for intelligent prediction.
[0288] This invention also provides a smart flood forecasting device based on concept-data driven coupling, comprising:
[0289] - At least one processor;
[0290] - At least one memory for storing at least one computer program;
[0291] The processor executes a computer program in memory to implement the steps of the concept-data-driven coupling-based intelligent flood forecasting method as described above.
[0292] The present invention also provides a computer storage medium storing a computer program, which is executed by a computer to implement the steps of the concept-data driven coupling-based intelligent flood forecasting method described above.
[0293] The above specific embodiments are used to explain and illustrate the present invention, and are only preferred embodiments of the present invention, not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made to the present invention within the spirit and scope of the claims shall fall within the protection scope of the present invention.
Claims
1. A smart flood forecasting method based on concept-data driven coupling, characterized in that, The method includes the following steps: S1: Collect historical datasets of the causes of floods, including flood duration, peak flow, runoff depth, surface rainfall, and surface evaporation. S2: Using the collected historical datasets of floods, classify and categorize flood types using a cloud model-based clustering algorithm; S3: Considering the characteristics of the decentralized Xin'anjiang conceptual model's sub-units and the information processing capabilities of the TFTformer model, the linear reservoir and the Muskingen River network confluence of the Xin'anjiang conceptual model are replaced by the TFTformer model to establish a concept-data driven coupled flood forecasting model. S4: Concept-data driven coupled flood forecasting model parameter calibration. Given a fixed flood model structure, the model's flood process simulation performance depends on the model's parameter calibration. This is achieved by introducing an improved daemon optimization algorithm into the concept-data driven coupled flood forecasting model to dynamically adjust its hyperparameters. S5: Use the historical dataset from step S1 as the model sample dataset, and divide it into training and test sets and feed it into the concept-data-driven coupled flood forecasting model to predict peak flow and runoff depth. S6: The Spatiotemporal Graph Convolutional Network (STGCN) is used to correct the errors in the above prediction results and output the flood type for intelligent prediction. In step S3, the process of constructing the concept-data driven coupled flood forecasting model is as follows: The Xin'anjiang conceptual model is a typical decentralized conceptual model, characterized by three parts: unit division, water source division, and stage division. Unit division takes into account the impact of uneven rainfall distribution, thus dividing the entire watershed into several unit watersheds based on the river network structure, and performing separate runoff generation and runoff calculations for each unit watershed. Water source division classifies runoff into surface runoff, interflow, and groundwater runoff based on runoff velocity. The runoff velocities of the three water sources are different, with surface runoff being the fastest and groundwater runoff being the slowest. The concept-data-driven coupled model computation comprises the following three modules: (1) Evapotranspiration calculation: Soil evapotranspiration is divided into upper layer, lower layer and deep layer. The evapotranspiration rate is calculated using a three-layer evapotranspiration model. The parameters include the average tension water capacity of the watershed. upper tension water capacity Lower layer tension water capacity Deep tension water capacity Evapotranspiration conversion factor and deep evaporation diffusion coefficient The calculation formula is as follows: in, This represents the total tension water storage capacity; This refers to the upper tension water storage capacity; This refers to the storage capacity of tension water in the lower layer; For deep tensile water storage; This represents the total evaporation rate; This refers to the evaporation rate of the upper layer. This refers to the evaporation rate of the lower layer. This refers to the deep evaporation rate; P represents evapotranspiration capacity, and P represents surface rainfall. like ,but: like ,but: like ,but: like but: like but: The parameters to be optimized are and ; (2) Calculation of runoff: Full runoff mode, the formula is as follows: In the formula: The area of runoff generation; The drainage area; This refers to the water storage capacity at a single point within the watershed. This represents the maximum water storage capacity at a single point in the basin. The index of the water storage capacity-area distribution curve reflects the uneven distribution of water storage capacity in the vadose zone of the watershed. A larger value indicates greater unevenness. The smaller the value, the more uniform it is. When it is 0, it means that the water storage capacity of the vadose zone in the watershed is uniform and unchanged. The parameter to be optimized is B. (3) Water source division: The water source division of the free water reservoir is matched, and the water source is divided into three types of runoff: surface runoff, interflow, and groundwater. In the formula: The free water storage capacity at a single point in the basin; It represents the maximum free water storage capacity at a single point in the basin. The power of the watershed free water storage capacity-area distribution curve; Surface runoff RS: Rangzhongliu RI: Subsurface runoff RG: In the formula: and These represent the ratio of runoff area at the beginning and end of the time period, respectively. Mean free water storage capacity; The flow velocity is the generated flow rate; the parameters to be optimized in the water source calculation include: , , and ; The surface rainfall in step S1 evaporation from water surface As input to the concept-data driven coupled model, the runoff sequence on permeable and impermeable areas is calculated. The runoff from permeable areas is divided into non-runoff areas and runoff areas. The non-runoff areas include: tension water, upper layer, lower layer and deep layer, and the evaporation of these three layers is calculated to form interflow. The runoff areas form surface free water, which constitutes groundwater runoff. The runoff from impermeable areas and tension water constitute surface runoff. The datasets of surface runoff, interflow, and total groundwater inflow obtained through the conceptual model are fed into the TFTformer model for training and prediction.
2. The intelligent flood forecasting method based on concept-data driven coupling according to claim 1, characterized in that, In step S1, the data includes the duration of the flood. Peak flow runoff depth Surface rainfall evaporation from water surface .
3. The intelligent flood forecasting method based on concept-data driven coupling according to claim 1, characterized in that, In step S2, the clustering algorithm based on the cloud model is used to classify flood types as follows: The cloud model-based clustering algorithm embeds clusters with random uncertainty and analyzes the changes in concepts within the clusters. The cloud model-based clustering algorithm consists of three parts: (1) Initial concept generation: Initialize a concept space: Duration of the flood Peak flow runoff depth Constructing includes A conceptual space , where each concept Described as a triple ( , , );expect Representing the initial concept A fundamental measure of determinism; entropy The term "hyperentropy" represents a measure of uncertainty in the initial concept, determined by the randomness and fuzziness of the concept. Represents entropy Uncertainty measure; Unlike knowledge discovery, which extracts concepts from complex data, concepts effectively integrate randomness and fuzziness; furthermore, concepts make objects in the same group more similar than those in other clusters, and through three numerical features ( , , Describe the overall quantitative properties of the concept; Cloud-based clustering algorithms, through datasets including flood duration, Peak flow runoff depth Embedding random uncertainty in the data, grouping the dataset into... Under these concepts, a stable conceptual space is formed. ; (2) Concept-based refined aggregation involves repeatedly dividing data embedded with random uncertainty into new concepts and searching for the optimal concept space: To partition data with a more precise concept, cloud-based clustering algorithms incorporate randomness into data partitioning, expanding the range of data distribution; the uncertainty of the clustering concept. It is a curve that defines how each point in the input space maps to an uncertain value between 0 and 1. Cloud-based clustering algorithms calculate the randomness contained in the data through the uncertainty of the clustering concept; randomness is determined by the value in the concept space. and Generate, represented as: in, Randomly generated from (0,1); Then, calculate The ambiguity of each data point in the dataset within a concept, and the embedding of randomness; Uncertainty of the clustering concept: in, The time series length of the dataset. Representing concepts The m-th concept center; In data partitioning, clustering concept uncertainty is effectively evaluated. The concept uncertainty is determined by the ambiguity of each data point in the dataset. Each data point has a different concept uncertainty. Then, the uncertainty is normalized to find a better data partition. The result of the normalization is called the concept importance, and the formula for the degree of importance is defined as follows: According to the principle of maximum membership, data points The larger, the more conceptual Most likely to include Under the concept of conceptual space, data is divided into... ; (3) Evaluation of conceptual uncertainty: The uncertainty of concepts during the iteration process is evaluated. The goal of cloud-based clustering algorithms is to make objects within a concept have high similarity, while objects between concepts have low similarity. Based on the conceptual importance of cluster center data points, concept centers... The calculation is as follows: in, For the concept Lower point Conceptual importance; Conceptual selection of data points The calculation is as follows: Conceptual parameters The calculation is as follows: in, and They are respectively Mean and variance; By calculating clusters of , and Generate concepts ; Reflection concept Discreteness of the midpoint Reflection concept The degree of clustering at the midpoint; Concept-based refined aggregation utilizes the fuzziness and randomness of objects to retain uncertain information, giving clusters more refined concepts. The function consists of a dataset and a concept C, as follows: This indicates that the dataset belongs to the concept. The probability of: in, Indicates that the concept center belongs to the concept The probability of; In a concept space, a point represents a concept. The probability of; Based on the dataset, flood duration Peak flow runoff depth In concept The data is divided into four flood types based on probability. The data is then fed into a concept-data-driven coupled flood forecasting model optimized using an improved sparrow optimization algorithm. The model outputs four flood forecast types. The optimal forecast is determined by calculating the reference distance between each of the four forecast results and its respective concept center. The formula for the reference distance is as follows: in, Time series length, To predict the duration of the flood using the model, For the duration of the flood at the concept center, To predict peak flood flow using the model, For the peak flow of the concept center, To predict runoff depth using the model, The runoff depth at the concept center.
4. The intelligent flood forecasting method based on concept-data driven coupling according to claim 1, characterized in that, In step S3, the specific content of the TFTformer model is as follows: The TFTformer model is a novel attention-based architecture that combines high-performance multi-view prediction with interpretable insights into temporal dynamics. To learn temporal relationships at different scales, it uses recurrent layers for local processing, interpretable self-attention layers for long-term dependencies, and leverages specialized components to select relevant features and a series of gating layers to suppress unwanted components, thereby achieving high performance in a wide range of scenarios. The TFTformer model mainly has the following structure: (1) Gating mechanism: Skip any unused components in the architecture, provide adaptive depth and network complexity to adapt to a wide range of datasets and scenarios; (2) Variable selection network layer: Select relevant input variables at each time step; (3) Static covariate encoder layer: integrates static features into the network and adjusts the temporal dynamics by encoding the context vector; (4) Temporal processing layer: learns long-term and short-term temporal relationships from observed and known time-varying inputs; the sequence layer is used for local processing, while long-term dependencies are captured using a novel interpretable multi-head attention block; (5) Output layer: The prediction interval of the possible target value range of each prediction layer is determined by quantile prediction.
5. The intelligent flood forecasting method based on concept-data driven coupling according to claim 1, characterized in that, In step S4, the parameters of the concept-data driven coupled flood forecasting model are calibrated, as follows: Given a fixed structure for a smart flood forecasting model, the simulation performance of the flood process depends on the model's parameter calibration. By introducing an improved sparrow optimization algorithm into the concept-data driven coupled model to dynamically adjust its hyperparameters, the reliability of the simulation results can be effectively improved. An improved star-sparrow optimization algorithm, NOA, is introduced to initialize the population. The following objective function is defined as the fitness function of the population: in, , and These are hyperparameters that control the weights of each term in the fitness function and are adjusted during training. This represents the relative error of the flood peak; a result close to 0 indicates a more accurate flood peak forecast. For predictive evaluation indicators, the closer the result is to 0, the higher the predictive indicator. The formula is: in, This represents the simulated value of the concept-data driven coupled model after inputting relevant data. Represents the actual observed value. This represents the average of the actual observed values; It mainly includes two strategies: (1) Simulate the food collection and storage strategy of the starbird in summer and autumn, introduce the golden sine algorithm to traverse all values of the sine function, and at the same time, update the location of the food to greatly improve the search speed. The specific formula is as follows: Exploration phase: in, For the current generation The new position of the star sparrow; For the current generation The first star sparrow One location; and These are the upper and lower bounds in an optimization problem; It is a random number; , and These are three different indicators randomly selected from the population to explore high-quality food sources; , A random real number in the range [0,1]. It is the j-th dimension mean of all solutions in the current population during the t-th iteration; µ is a number randomly generated between 0 and 1 based on a normal distribution. The development phase of introducing the golden sine algorithm: in, The current iteration The new location of Zhongxingque's storage area , A random real number in the range [0,1]. It is a factor that linearly decreases in development behavior from 1 to 0; It is the optimal solution; and It is a random number. This determines the distance the starbird will move in the next iteration. ; Determine the direction of position update for the next iteration. ; and These are coefficients obtained through the golden ratio; (2) Simulate the search and retrieval strategy of the star sparrow in spring and winter, i.e., cache search and retrieval strategy. Improve the cache search and retrieval strategy by using Levy flight to enhance the star sparrow's global search capability and avoid the algorithm getting trapped in local optima. The specific formula is as follows: Introducing Levy flight cache search: in, The search behavior mechanism of the star sparrow. The ratio of the current iteration number to the total number of iterations. Power; The step size is random. Dot product; To comply with Constraints under distributed random search paths; and It follows a standard normal distribution; =1.5; This is the search magnification factor; It is a gamma function; Introducing the Levy flight recovery strategy: in, Restore the behavior mechanism of the starbird; This optimization algorithm is used to optimize parameters in the above concept-data driven coupling, such as the parameters to be optimized in evaporation calculations. and The parameter to be optimized in the runoff calculation is B; the parameters to be optimized in the water source calculation include: , , and The learning rate and hidden layer number in TFTformer and the following spatiotemporal graph convolutional network STGCN are used to obtain the parameters of the concept-data driven coupling model suitable for this watershed.
6. The intelligent flood forecasting method based on concept-data driven coupling according to claim 1, characterized in that, In step S6, the spatiotemporal graph convolutional network STGCN is used to correct the errors in the above prediction results, as follows: Due to limitations imposed by factors such as dataset size, model structure parameters, algorithm convergence, and robustness during the forecasting process, the forecasting accuracy of concept-data-driven coupled flood forecasting models needs further improvement. Therefore, real-time correction of flood forecasting errors is necessary to enhance the accuracy and practicality of the forecasting model. Real-time correction, based on real-time information during the forecasting process, applies modern information technology theories and methods to dynamically adjust model parameters, model inputs, and forecasting results, constructing a real-time feedback mechanism between the forecasting model and the correction model to reduce flood forecasting errors. This is a crucial component of real-time flood forecasting. Spatiotemporal graph convolutional networks contain the following structure: (1) Input layer: based on concept-data driven coupled flood forecasting with surface rainfall evaporation from water surface Flood forecasting is performed using forecasting factors; Therefore, the input layer of this model uses concept-data driven predictions from a coupled flood forecasting model. and observed values Error sequence between The set of rainfall sequences at control stations in a sub-basin is defined as follows: As an input variable to the spatiotemporal graph convolutional network; The flow lengths from relevant stations within the watershed to the target station are used to construct the STGCN adjacency matrix. Construction, training set and test set partitioning, The calculation formula is as follows: in, Indicates the runoff length between two stations; (2) Network layer: Training a mapping function that reflects the features of the error sequence; first, construct two ST-Conv blocks consisting of gated sequence convolutional layers and spatial sequence convolutional layers to extract the temporal and local spatial features of the error sequence; the relationship between the simulated error value output by the STGCN layer and the input layer data is as follows: in, express The concept of time - the simulated error value obtained from the data-driven coupled model; yes Time error sequence; , ... yes time Rainfall at several relevant rain gauge stations; That is, the mapping function determined by the concept-data driven coupling model; (3) Output layer: Output error correction results; calculate error The simulation error obtained in step (2) Difference between and initialize its minimum value. ;Compare , Size and ( and )and( and The root mean square error, coefficient of determination, and mean absolute error of the two results are commonly used evaluation indicators. If the evaluation indicator values improve, the simulated error values are output after inverse normalization. Finally, the original predicted values from the forecast model are used. and final simulation error The model correction results were calculated. The relationship is as follows:
7. A smart flood forecasting system based on concept-data driven coupling, characterized in that, The system is based on the concept-data driven coupling-based intelligent flood forecasting method as described in claim 1, and includes the following modules: - Historical data collection module, which is used to collect historical datasets that lead to floods, including flood duration, peak flow, runoff depth, surface rainfall and water surface evaporation; - Flood type classification and grading module, which is used to classify and grade flood types using a cloud model-based clustering algorithm based on the historical data collection module that leads to floods; - Concept-data driven coupled flood forecasting model establishment module: The concept-data driven coupled flood forecasting model establishment module is used to consider the characteristics of the decentralized Xin'anjiang concept model sub-units and the information processing capabilities of the TFTformer model, and to replace the linear reservoirs and the Muskingen river network confluence of the Xin'anjiang concept model with the TFTformer model to establish a concept-data driven coupled flood forecasting model; - Concept-data driven coupled flood forecasting model parameter calibration module, which is used to dynamically adjust the hyperparameters of the concept-data driven coupled flood forecasting model by introducing an improved star-sparrow optimization algorithm; - Data prediction module, which uses the historical dataset obtained by the historical data collection module as the model sample dataset, and divides it into training set and test set and feeds it into the concept-data driven coupled flood forecasting model to predict peak flow and runoff depth; - Flood type intelligent prediction module, which uses a spatiotemporal graph convolutional network STGCN to correct errors in the above prediction results and outputs the flood type for intelligent prediction.
8. A smart flood forecasting device based on concept-data driven coupling, characterized in that, The device includes: - At least one processor; - At least one memory for storing at least one computer program; The processor executes a computer program in memory to implement the steps of the concept-data driven coupling-based intelligent flood forecasting method as described in any one of claims 1-6.
9. A computer storage medium, characterized in that, The computer storage medium stores a computer program, which is executed by a computer to implement the steps of the concept-data driven coupling-based intelligent flood forecasting method as described in any one of claims 1-6.