A knowledge and data fusion driven online optimization method for flood forecasting model parameters
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HOHAI UNIV
- Filing Date
- 2023-06-13
- Publication Date
- 2026-08-07
AI Technical Summary
[0073]有益效果:与现有技术相比,本发明的有益效果:本发明利用知识图谱组织领域知识和规则,通过分析流域情势在知识图谱中匹配需优化参数和方向;通过模式库构建多种情势数据,并以数据驱动的形式实现优化参数和范围的匹配,这种方法可以更加全面地考虑参数的影响范围,从而修订由知识图谱中规则匹配的参数范围,达到更加准确的匹配效果;通过知识和数据融合驱动,实现洪水预报模型参数智能在线优化,提高业务处理的效率;此外,在自优化的基础上,具有一定自动化模拟能力能够执行类似人类的智能活动,不断自学习实现知识的自增长。
Smart Images

Figure CN116776581B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent optimization technology driven by the fusion of digital twin watershed data and knowledge and flood forecasting, and particularly relates to an online optimization method for flood forecasting model parameters driven by knowledge and data fusion. Background Technology
[0002] Accurate and timely flood forecasts are crucial for basin-wide flood control and disaster relief decision-making. The structure, parameters, and states of hydrological models must be flexible enough to adapt to the complex and ever-changing underlying surface environment of the basin. This is not only essential for improving flood forecasting capabilities but also a requirement for the construction of digital twin basins. With the successful experimentation of modular, flexible hydrological models, developing parameter optimization methods to adapt to standardized model interfaces has become imperative. Online optimization of model parameters during the operational forecasting phase restores the true nature of the time-varying characteristics of sensitive parameters in the basin's hydrological models, representing a new and feasible path to improve the accuracy of operational flood forecasts.
[0003] The fusion of knowledge and data has become a crucial technology in current scientific research and practice, possessing significant application prospects and development potential. Data and knowledge mutually reinforce and support each other, enabling more accurate, reliable, and in-depth information analysis and decision-making. In this invention, knowledge graphs are used to organize various types of knowledge, forming a more comprehensive, accurate, and in-depth knowledge system. This system includes not only traditional rule-based knowledge but also knowledge in various forms such as experience and experimental data. Analysis of watershed situations is used to optimize parameters and ranges based on rule matching within the knowledge graph. Furthermore, pattern library technology can construct various situational data to effectively compensate for insufficient data volume, achieving optimized parameter and range matching in a data-driven manner. A more comprehensive consideration of the parameter's influence range allows for revision of the parameter range matched by rules in the knowledge graph, achieving a more accurate matching effect.
[0004] The two methods complement and support each other. Knowledge graphs provide rich knowledge resources, enabling the rapid retrieval of knowledge relevant to specific situations through rule matching. Pattern libraries, on the other hand, offer substantial data support, allowing for a more comprehensive consideration of the parameters' influence and enabling data-driven revision of matching results. The combination of these two methods forms a more complete, accurate, and reliable knowledge- and data-driven optimization scheme, possessing significant practical application value. Summary of the Invention
[0005] Purpose of the invention: The purpose of this invention is to provide a knowledge and data fusion-driven online optimization method for flood forecast model parameters, thereby achieving intelligent online optimization of flood forecast model parameters.
[0006] Technical Solution: This invention provides a knowledge and data fusion-driven online optimization method for flood forecasting model parameters, comprising the following steps:
[0007] (1) Data preprocessing: Analyze the data samples during the preheating period of the watershed, and divide the time series data into flood events and extract features;
[0008] (2) Analyze and determine the water conservancy calculation objects and models: spatial contribution analysis determines the water conservancy objects to be adjusted, the characteristic quantities of the adjusted flood are determined according to the characteristics of the basin, and the corresponding model identifiers of the water conservancy calculation objects are queried;
[0009] (3) Knowledge graph retrieval optimization parameters and directions: Based on the knowledge graph matching model, the sensitive parameters and adjustment directions under flood adjustment characteristics, and the optimization range of matching parameters;
[0010] (4) Revise the parameter value range based on the pattern library: Extract features online to match similar patterns, and compare them with the parameter direction and range obtained in step (3). If they are inconsistent, revise the parameter value range.
[0011] (5) Update the parameter value range to the parameter calibration file, run the parameter calibration algorithm, and record the decision-making process of determining the parameter optimization range;
[0012] (6) Evaluation and feedback of optimization effect: The forecast accuracy is evaluated based on the calculation results of the parameter calibration method, and a new evaluation model is extracted based on the parameter optimization decision-making process. The model library is dynamically updated to achieve self-growth of the model library.
[0013] Furthermore, the implementation process of step (1) is as follows:
[0014] (11) Flood events are divided into distinct time series, and the complex changes in floods are analyzed. Long-term flood time series data are processed into short-term time series data.
[0015] The argrelmax function is used to find local extreme points in the time series data, and extreme points with flow values greater than 100 and time intervals greater than 24 hours are selected. The data is divided into segments with a window value of 10 according to the actual situation of the watershed, and process labels are applied, with 0 indicating that it is not a flood process and 1 indicating that it is a flood process. The flood process segmentation is manually corrected where it is unclear.
[0016] (12) Rainstorm and flood feature extraction: extract rainfall, flow and watershed features that affect parameter optimization during the preheating period flow process, realize the transformation of scene data to feature data, and describe the different roles of rainfall, flow and watershed in parameter optimization.
[0017] Furthermore, the implementation process of step (2) is as follows:
[0018] (22) Spatial contribution analysis determines the water conservancy targets for adjustment. By analyzing the regional composition and rainfall situation of flood events during the preheating period, the parameters adjustment sections or intervals are determined:
[0019] Extract flood features for each object to form a feature matrix X; calculate weights based on the extracted features and their weights:
[0020]
[0021] Where θ is the weight matrix, θ i Let X represent the weight values of each feature, and let X be the feature matrix. i This represents the value of each feature; based on the weight calculation results, the global contribution is sorted, and the node or region with the largest contribution is determined as the water conservancy object to be adjusted, denoted as UnitNo;
[0022] (22) Analyze and determine the adjustment characteristic quantities, and identify the sensitive parameters to be optimized based on the deviation of flood characteristic quantities:
[0023] Set thresholds for peak time difference, flood peak height difference, and their priorities; extract the peak time difference and flood peak difference features, compare them with the thresholds respectively, and record whether adjustments are needed; if both are greater than the thresholds, calculate the weights of the two priorities, and the one with the larger weight is the feature that needs adjustment, denoted as FeatureNo;
[0024] (23) Determine the adjustment object model. Based on the UnitNo calculated in step (21) and the FeatureNo calculated in step (22), use them as keys to query the predefined parameter optimization rules in the knowledge graph and return the key value, which is the identifier of the water conservancy object model.
[0025] Furthermore, the implementation process of step (3) is as follows:
[0026] (31) Assemble query statements based on adjusting the water conservancy object model identifier and adjusting the flood deviation situation of characteristic quantities;
[0027] (32) Initiate a query request to the knowledge graph: Initiate a query based on the query address provided by the knowledge graph to query the predefined rules for parameter optimization organized by the knowledge graph. The predefined rules include three types of models: SMS_3, LAG_3, and MSK models, and five types of optimization sensitive parameters: SM, CS, LAG, X, and MP. Among them, SM represents the soil moisture sensitive parameter, which is used to describe the degree of soil response to changes in water content; CS represents the river channel storage sensitive parameter, which is used to describe the river channel's capacity to store and release runoff; LAG represents the watershed confluence delay sensitive parameter, which is used to describe the delay time between rainfall arrival and runoff response; X represents the flow proportion coefficient, which is used to adjust the proportional relationship of each link in the model; MP represents the number of sub-river segments, reflecting the degree of flood process shift.
[0028] (33) Return query results: Adjust the optimization sensitivity parameters and directions of water conservancy objects.
[0029] Furthermore, the implementation process of step (4) is as follows:
[0030] (41) Online scene feature extraction: extract rainfall and flow features that significantly affect parameter optimization during the preheating period of the watershed, and realize the transformation of scene data into feature data;
[0031] (42) The model library matching optimization parameter range is based on the matching of the spatiotemporal characteristics of the watershed with the multiple types of models in the offline model library, and the most similar model is output as the optimization parameter range.
[0032] (43) Update parameter range: Compare with the parameter optimization range obtained in step (3). If they are inconsistent, update the parameter optimization range to the pattern library matching result and update the parameter range file.
[0033] Furthermore, the implementation process of step (5) is as follows:
[0034] (51) Execute the hydrological parameter calibration algorithm: Start the calibration service, call the function multiValueMapParam, and output the optimal parameter values in the parameter calibration process;
[0035] (52) Update computation time: Change the computation start time in the runfile and input files of each model from the warm-up period to the full duration;
[0036] (53) Calculate the forecast segment flow: Call the function SingleStep to output the calculated flow of the forecast segment.
[0037] Furthermore, the implementation process of step (6) is as follows:
[0038] (61) Evaluate forecast accuracy: After optimizing the parameters of the operation forecast, extract the peak time, rise time, receding time, total flood volume, peak flow, and rainstorm center point sequence features, calculate the difference between these features and the measured features, and the difference is considered to be successful if it is within the threshold range;
[0039] (62) Extracting new patterns: Based on the parameter optimization range decision process record of the calibration algorithm, analyze historical flood samples, extract the features of the preheating period samples, including peak flow, peak time, rise time, receding time, lag time, and rainstorm center point sequence, and map the features, parameter optimization range, and forecast accuracy evaluation results to a pattern; the formal expression of the features and parameter optimization range is as follows:
[0040] U = {the set of all characteristic factors of the flood} (11)
[0041] X={S(i) k |0≤i≤n,n=|U} (12)
[0042] Y = R(k) (13)
[0043] F(X,Y) (14)
[0044] D = {F(X,Y)} (15)
[0045] Where, S(i) k Let X represent the feature set of the k-th flood, and R(k) represent the parameter optimization range and forecast accuracy assessment result of the k-th flood. The formula indicates that there is a mapping relationship F between the feature set X and the parameter model Y. The parameter optimization model is defined as F(X,Y). The set of multiple parameter optimization models forms the model library D. The generated new models are filled into the model library to gradually form a parameter optimization model library D adapted to the characteristics of the current watershed.
[0046] (64) Update the model library: After optimizing the parameters of the operation forecast, promptly evaluate the forecast accuracy, record the decision-making process of the parameter optimization range, generate a new mapping sample of "warm-up period conditions - parameter optimization range - forecast accuracy evaluation", provide feedback to supplement the parameter optimization rules, and gradually form a parameter optimization model library that adapts to the current characteristics of the watershed, thereby improving the success rate of parameter correction.
[0047] Furthermore, the implementation process of step (12) is as follows:
[0048] Extract offline flow characteristics, including: peak flow, the maximum instantaneous flow during a flood event, i.e., the flow at the highest point on the flood process line; peak time, the time when the flood peak occurs; lag time, the difference between the measured peak time and the calculated peak time; peak difference, the difference between the measured peak flow and the calculated peak flow; flood initiation time, the time when the flood process begins; and receding time, the time when the flood process ends.
[0049] Extracting offline rainfall features, including: the rainstorm center point sequence, which is a sequence B in tuple form. d ={d1,d2,...d i ,...,d n ], where d i =[x i ,y i ,t i ], x i ,y i This represents the coordinates of the storm's center point, which is the location of the grid point with the maximum rainfall. t i This represents the time at the center of the rainstorm; the trajectory of the rainstorm center is a sequence of tuples. Transform the rainstorm center sequence into edge form, sequence B. e =
[0050] e1, e2, ... e i,...,e n ], where e i =[(x i-1 ,y i-1 ,t i-1 ),(x i ,y i ,t i )];
[0051] Extract basic features of the offline watershed, including watershed area and the area enclosed by the watershed divide.
[0052] Furthermore, the implementation process of step (41) is as follows:
[0053] (411) Online feature selection: including total flood volume, peak flow, rise time, receding time, duration, peak water level, rise duration, receding duration, and rainfall feature storm center point sequence;
[0054] (412) Online feature extraction: Total flood volume, the total amount of water flowing out of the outlet section of the basin within a certain duration; Peak flow, the flow rate at the moment the flood basin reaches its peak; Onset time, for the time nodes of the flood process, the time when the flood process begins is taken as the onset time of the flood; Receding time, the time when the flood process ends is taken as the receding time of the flood; Duration, the flood receding time minus the start time is taken as the duration of the flood; Peak water level, the highest water level during the flood process is taken as the peak water level; Onset duration, the time from the start time of the flood to the peak; Receding duration, the time from the last flood peak to the end time of the flood; Basin rainfall characteristics, including the sequence of rainstorm center points: a sequence B in the form of a tuple. d ={d1,d2,...d i ,...,d n}, where d i =[x i ,y i ,t i ], x i ,y i This represents the coordinates of the storm's center point, which is the location of the grid point with the maximum rainfall. t i The time of the storm's center point is represented; finally, the extracted features are used to form a feature matrix F, realizing the transformation from scene data to feature data.
[0055] Furthermore, the implementation process of step (42) is as follows:
[0056] For non-time-series data (i.e., single numerical values) in the feature matrix F, the values at the first t-1 time points are automatically padded with zeros; for the extracted multiple hydrological features, ensemble weights are used to assign weights to different features:
[0057]
[0058] Where, θ . r represents the degree of trust that expert knowledge has in the k-th weight vector. j The negative ideal solution of the index; construct the Lagrangian function using formula (2) and solve for θ. . Then by The integrated weight U is obtained;
[0059] Given two multivariate time series, the real values of each column are used as coordinate parameters of points in an S-dimensional space. Point-to-point comparisons are performed by calculating the Euclidean distance between the two points. Then, DTW (Digital Time-Warping) is used to calculate the similarity of the hydrological multivariate time series. For the weighting part involved in the similarity calculation, an integrated weight calibration method suitable for hydrological multivariate time series is used to determine the weights. The multivariate fusion algorithm is as follows:
[0060]
[0061]
[0062] Let x p =(x p1 ,x p2 ,…,x pm ), y p =(y p1 ,y p2 ,…,y pn If X and Y are denoted as X, then the DTW distance between them is:
[0063]
[0064] Among them, DTW p Representing sequence x p and y p DTW distance, γ p Let γ be the weight for each dimension, and ∑γ = 1;
[0065] Let x i =(x 1i ,x 2i ,…,x si ), y j =(y 1j ,y 2j ,…,y sj ), then the MDTW distance between X and Y is:
[0066] DTW(X,Y)=d(n,m)(6)
[0067] d(i,j)=f(x i 6 j )+d best (i,j)(7)
[0068]
[0069]
[0070] d(0,0)=0; d(i,0)=0; d(0,j)=∞(10)
[0071] (i=1,2,…,n; j=1,2,…,m)
[0072] The calculated distance is the similarity. By sorting the matching results, the process line with the highest similarity is selected. The model and parameter values obtained in this process are the results returned by the pattern library matching.
[0073] Beneficial Effects: Compared with existing technologies, the present invention offers the following advantages: It utilizes a knowledge graph to organize domain knowledge and rules, and matches the parameters and directions to be optimized within the knowledge graph by analyzing the watershed situation; it constructs various situational data through a pattern library and achieves the matching of optimization parameters and ranges in a data-driven manner. This method can more comprehensively consider the influence range of parameters, thereby revising the parameter range matched by rules in the knowledge graph to achieve a more accurate matching effect; through knowledge and data fusion, it realizes intelligent online optimization of flood forecast model parameters, improving the efficiency of business processing; furthermore, based on self-optimization, it has a certain degree of automated simulation capability, enabling it to perform intelligent activities similar to humans, and continuously learns to achieve self-growth of knowledge. Attached Figure Description
[0074] Figure 1 This is a flowchart of the present invention;
[0075] Figure 2 This is a schematic diagram of the process for analyzing and determining the adjustment of water conservancy objects proposed in this invention;
[0076] Figure 3 This is a schematic diagram illustrating the range of revision parameters for the pattern library proposed in this invention;
[0077] Figure 4 This is a schematic diagram of the execution parameter calibration algorithm proposed in this invention;
[0078] Figure 5 This is a schematic diagram of the optimized evaluation and update mode library proposed in this invention. Detailed Implementation
[0079] The present invention will now be described in further detail with reference to the accompanying drawings.
[0080] This invention provides an online optimization method for flood forecasting model parameters based on knowledge and data fusion, such as... Figure 1 As shown, it includes the following steps:
[0081] Step 1: Data preprocessing, analyzing data samples during the watershed warm-up period, dividing the time series data into flood events and extracting features.
[0082] Flood events are divided into different time series, and the complex changes in floods are analyzed. Long-term flood time series data are processed into short-term data that can be processed by computers.
[0083] The argrelmax function is used to find the local extreme points of the time series data, and extreme points with a value greater than 100 and a distance greater than 24 hours are selected. The data is divided into segments with a window value of 10 according to the actual situation of the watershed, and process labels are applied, with 0 indicating that it is not a flood process and 1 indicating that it is a flood process. The flood process segmentation is manually corrected where it is unclear.
[0084] Rainfall and flood feature extraction: Based on expert knowledge, this method extracts rainfall and flow features that significantly affect parameter optimization during the preheating period, transforming scenario data into feature data and describing the different roles of rainfall and flow in parameter optimization.
[0085] Extract flow characteristics, including peak flow: the maximum instantaneous flow during a flood event, i.e., the flow at the highest point on the flood process line; peak time: the time when the flood peak occurs; lag time: the difference between the measured peak time and the calculated peak time; peak difference: the difference between the measured peak flow and the calculated peak flow; flood initiation time: the time when the flood process begins; and flood recession time: the time when the flood process ends.
[0086] Extracting rainfall features, including the storm center sequence: The storm center sequence is a sequence B in tuple form. d ={d1,d2,...d i ,...,d n}, where d i =[x i ,y i ,t i ], x i ,y i This represents the coordinates of the storm's center point, which is the location of the grid point with the maximum rainfall. t i Represents the time of the storm's epicenter; the trajectory of the storm's epicenter is a sequence of tuples. Transform the storm's epicenter sequence into an edge-like form, sequence B. e ={e1,e2,...e i ,...,e n}, where e i =[(x i-1 ,y i-1 ,t i-1 ),(x i ,y i ,ti )).
[0087] Extract the basic characteristics of the watershed, including the watershed area: the area enclosed by the watershed divide, calculated in square kilometers.
[0088] Step 2: Analyze and determine the water conservancy calculation objects and models; spatial contribution analysis determines the water conservancy objects to be adjusted; based on watershed characteristics, determine the characteristic quantities of adjusted floods; and query the corresponding models for the water conservancy calculation objects. For example... Figure 2 As shown, the specific process is as follows:
[0089] (2.1) Spatial contribution analysis determines the adjustment target. By analyzing the regional composition and rainfall situation of floods during the preheating period, the parameter adjustment section or interval is determined and denoted as UnitNo.
[0090] Flood characteristics of each object are extracted, such as peak flow, peak time, and peak time difference, to form a feature matrix X. Based on the extracted features and their weights, weights are calculated. In the experiment, the feature weight matrix θ is set to [0.2, 0.3, 0.5]. This weight can be fine-tuned according to different watershed conditions. Then, formula (1) is used to calculate the weights:
[0091]
[0092] Where θ is the weight matrix, θ i Let X represent the weight values of each feature, and let X be the feature matrix. i This represents the value of each feature. Based on the weight calculation results, the global contributions are sorted, and the node or region with the largest contribution is identified as the adjustment target, denoted as UnitNo; the adjustment model is then determined based on the adjustment target.
[0093] (2.2) Analyze and determine the flood adjustment characteristic quantity. In order to imitate the forecaster's decision-making process of "identifying sensitive parameters - determining the parameter optimization range", based on the deviation of the flood characteristic quantity, identify the sensitive parameters to be optimized and record the adjustment characteristic quantity as FeatureNo.
[0094] Threshold and priority settings: Based on domain knowledge, set the thresholds for peak time difference, peak height difference, and their priorities. The thresholds vary depending on the watershed; for example, in the relatively small Tunxi watershed, the peak time difference threshold should be set to 1 hour.
[0095] Threshold determination: Extract the peak time difference and flood peak difference features, compare them with the thresholds respectively, and record whether adjustment is needed. If both are greater than the threshold, determine the feature quantity to be adjusted; calculate the priority of the two and the weight of the one with the larger weight is the feature quantity that needs to be adjusted in subsequent steps.
[0096] (2.3) Determine the adjustment object model, query the predefined parameter optimization rules in the knowledge graph based on UnitNo and FeatureNo as keys, and return the key value, i.e. the water conservancy object model identifier.
[0097] Step 3: Optimize parameters and directions for knowledge graph retrieval. Based on the knowledge graph matching model, optimize the sensitive parameters and adjustment directions of the corresponding feature quantities, and the range of matching parameter optimization.
[0098] (3.1) Assemble query statements based on adjusting the water conservancy object model identifier and adjusting the flood deviation situation of characteristic quantities.
[0099] (3.2) Initiating a query request to the knowledge graph: A query is initiated based on the query address provided by the knowledge graph to retrieve the predefined rules for parameter optimization organized by the knowledge graph. These predefined rules include three types of models: SMS_3, LAG_3, and MSK models, and five types of optimization-sensitive parameters: SM, CS, LAG, X, and MP. SM represents the soil moisture sensitivity parameter, used to describe the degree of soil response to changes in moisture; a higher SM value indicates that soil moisture has a greater impact on the model output. CS represents the river channel storage sensitivity parameter, used to describe the river's capacity to store and release runoff; a higher CS value indicates that river channel storage has a greater impact on the model output. LAG represents the watershed confluence delay sensitivity parameter, used to describe the delay time between rainfall arrival and runoff response; a higher LAG value indicates that watershed confluence delay has a greater impact on the model output. X represents the flow proportion coefficient, used to adjust the proportional relationship between various components in the model; a higher X value indicates that this component has a greater impact on the model output. MP represents the number of sub-segments, reflecting the degree of shift in the flood process.
[0100] (3.3) Return query results: optimize sensitive parameters and directions.
[0101] Step 4: Based on the pattern library, revise the parameter value range, extract features, match the corresponding patterns in the pattern library, and compare them with the parameter direction and range obtained in Step 3. If they are inconsistent, revise the parameter value range. Specifically, as follows... Figure 3 As shown, the implementation process is as follows:
[0102] (4.1) Online scene feature extraction: extract rainfall and flow features that significantly affect parameter optimization during the preheating period of the watershed, and realize the transformation of scene data into feature data.
[0103] Feature selection: A runoff process can be characterized by many features. Selecting as few sub-features as possible will not significantly reduce model performance, and the resulting class distribution should closely approximate the true class distribution. Based on expert knowledge, the features to be extracted include the characteristics of large and small flood events used in the calibration scheme, the location of the storm center, and the combination of inflow from main streams and tributaries. Ultimately, the extracted features include total flood volume, peak flow, flood onset time, flood recession time, duration, peak flood level, flood onset duration, flood recession duration, and the rainfall characteristic storm center sequence.
[0104] Feature extraction: First, basic features of the watershed runoff are extracted, including total flood volume: the total amount of water flowing out of the watershed outlet section within a certain duration; peak flow: the flow rate at the moment the flood reaches its peak, usually marked by a peak indicator; if not, the maximum point of the flood process curve is used as the peak indicator; flood onset time: the time when the flood process begins is taken as the flood onset time; receding time: the time when the flood process ends is taken as the flood receding time; duration: the flood receding time minus the start time is taken as the flood duration; peak water level: the highest water level during the flood process is taken as the peak water level; flood onset duration is the time from the start time of the flood to the peak, calculated in hours; receding duration is the time from the last flood peak to the end of the flood, calculated in hours. Second, watershed rainfall features are extracted, including the storm center sequence: a sequence B in tuple form. d ={d1,d2,...d i ,...,d n}, where d i =[x i ,y i ,t i ], x i ,y i This represents the coordinates of the storm's center point, which is the location of the grid point with the maximum rainfall. t i The time of the storm's center point is represented; finally, the extracted features are used to form a feature matrix F, realizing the transformation from scene data to feature data.
[0105] (4.2) Matching the model library to optimize the parameter range: Based on the spatiotemporal characteristics of the watershed and the matching of multiple types of models in the offline model library, the most similar model is output as the optimized parameter range.
[0106] Feature matrix processing: To facilitate matching with the subsequent pattern library, the non-time-series data (i.e., single values) in the feature matrix F obtained in the above steps are automatically padded with zeros for the first t-1 time steps.
[0107] S422: Feature Weighting: For the multiple hydrological features extracted in the above steps, ensemble weights are used to assign weights to different features. The weights are calculated using the optimization model shown in formula (2), where θ... kr represents the degree of trust that expert knowledge has in the k-th weight vector. j The negative ideal solution for the index is given. Using this optimization model, the Lagrangian function is constructed and θ is solved. . Then by The integrated weight U is obtained.
[0108]
[0109] S423: Multivariate Fusion Algorithm Matching. Hydrological time series is a type of multivariate time series. For similarity measurement of multivariate time series, a multivariate fusion algorithm is used to measure the similarity of hydrological feature time series. Given two multivariate time series, the real values of each column are used as coordinate parameters of points in an S-dimensional space. Point comparisons are performed by calculating the Euclidean distance between the two points. Then, the similarity of the hydrological multivariate time series is calculated using the DTW (Multivariate Time Series Design) approach. Based on the MDTW algorithm for multivariate time series, for the weighting part involved in the similarity calculation, an integrated weight calibration method suitable for hydrological multivariate time series is used in step S422 to determine the weights. The multivariate fusion algorithm is as follows:
[0110]
[0111]
[0112] Let x p =(x p1 ,x p2 ,…,x pm ), y p =(y p1 ,y p2 ,…,y pn If X and Y are 0, then the DTW distance between X and Y can be expressed by the following formula:
[0113]
[0114] Among them DTW p Representing sequence x p and y p DTW distance, γ p Let γ be the weight of each dimension, and ∑γ = 1.
[0115] Let x i =(x 1i ,x 2i ,…,x si ), y j =(y 1j ,y 2j ,…,y sj Then, the MDTW distance between X and Y can be expressed by the following formula:
[0116] DTW(X,Y)=d(n,m)(6)
[0117] d(i,j)=f(x i 6 j )+d best (i,j)(7)
[0118]
[0119]
[0120] d(0,0)=0; d(i,0)=0; d(0,j)=∞(10)
[0121] (i=1,2,…,n; j=1,2,…,m)
[0122] The calculated distance is the similarity. By sorting the matching results, the process line with the highest similarity is selected. The modified model and parameter values of this process are the results returned by the pattern library matching.
[0123] (4.3) Update parameter range: Compare with the parameter optimization range obtained in step 3. If they are inconsistent, update the parameter optimization range to the pattern library matching result and update the parameter range file.
[0124] Step 5: Update the parameter range obtained in the above steps to the parameter calibration file, run the parameter calibration algorithm, and record the decision-making process for determining the parameter optimization range. For example... Figure 4 As shown, the implementation process is as follows:
[0125] Execute the hydrological parameter calibration algorithm: Start the calibration service, call the function multiValueMapParam, and output the optimal parameter values in the parameter calibration process.
[0126] Update computation time: Change the computation start time in the runfile and input files of each model from the warm-up period to the full-time duration;
[0127] Forecast segment flow calculation: Call the function SingleStep to output the calculated flow of the forecast segment.
[0128] Step 6: Post-optimization evaluation and feedback. Based on the parameter calibration method, assess forecast accuracy. Develop a decision-making process based on parameter optimization, extract new evaluation models, and dynamically update the model library to achieve self-growth. For example... Figure 5 As shown, the implementation process is as follows:
[0129] Assessing forecast accuracy: The calculation process after optimizing the parameters of the operational forecast involves extracting the peak time, rise time, receding time, total flood volume, peak flow, and rainstorm center point sequence features. The difference between these features and the measured features is calculated, and the adjustment is considered successful if the difference is within the threshold range.
[0130] Extracting a new pattern: Based on the parameter optimization range decision-making process record of the calibration algorithm, historical flood samples are analyzed to extract features of the pre-flood period samples, including peak flow, peak time, rise time, receding time, lag time, and storm center sequence. These features, along with the parameter optimization range and forecast accuracy evaluation results, are mapped to a new pattern. The formal expression of the features and parameter optimization range is as follows:
[0131] U = {the set of all characteristic factors of the flood} (11)
[0132] X={S(i) k |0≤i≤n,n=|U} (12)
[0133] Y = R(k) (13)
[0134] F(X,Y) (14)
[0135] D = {F(X,Y)} (15)
[0136] Where, S(i) k Let X represent the feature set of the k-th flood, and R(k) represent the parameter optimization range and forecast accuracy assessment result of the k-th flood. The formula indicates a mapping relationship F between the feature set X and the parameter model Y, and the parameter optimization model is defined as F(X,Y). A collection of multiple parameter optimization models forms a model library D. New models are added to the model library, gradually forming a parameter optimization model library D adapted to the characteristics of the current watershed.
[0137] Update the model library: After optimizing the parameters of the operation forecast, promptly evaluate the forecast accuracy, record the decision-making process of the parameter optimization range, generate a new mapping sample of "warm-up period conditions - parameter optimization range - forecast accuracy evaluation", provide feedback to supplement the parameter optimization rules, and gradually form a parameter optimization model library that adapts to the current characteristics of the watershed, thereby improving the success rate of parameter correction.
[0138] The technical solution of the present invention has been described in conjunction with the specific experimental procedures shown in the accompanying drawings. However, the scope of protection of the present invention is not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions resulting from such changes or substitutions will all fall within the scope of protection of the present invention.
Claims
1. A knowledge and data fusion-driven online optimization method for flood forecasting model parameters, characterized in that, Includes the following steps: (1) Data preprocessing: Analyze the data samples during the preheating period of the watershed, and divide the time series data into flood events and extract features; (2) Analyze and determine the water conservancy calculation objects and models: spatial contribution analysis determines the water conservancy objects to be adjusted, the characteristic quantities of the adjusted flood are determined according to the characteristics of the basin, and the corresponding model identifiers of the water conservancy calculation objects are queried; (3) Knowledge graph retrieval optimization parameters and directions: Based on the knowledge graph matching model, the sensitive parameters and adjustment directions under flood adjustment characteristics, and the optimization range of matching parameters; (4) Revise the parameter value range based on the pattern library: extract features online to match similar patterns, and compare them with the parameter direction and range obtained in step (3). If they are inconsistent, revise the parameter value range. (5) Update the parameter value range to the parameter calibration file, run the parameter calibration algorithm, and record the decision-making process of determining the parameter optimization range; (6) Evaluation and feedback of optimization effect: The forecast accuracy is evaluated based on the calculation results of the parameter calibration method, and a new evaluation model is extracted based on the parameter optimization decision-making process. The model library is dynamically updated to realize the self-growth of the model library. The implementation process of step (3) is as follows: (31) Assemble query statements based on the adjusted water conservancy object model identifier and the adjusted characteristic quantity flood deviation situation; (32) Initiate a query request to the knowledge graph: Initiate a query based on the query address provided by the knowledge graph to query the predefined rules for parameter optimization organized by the knowledge graph. The predefined rules include three types of models: SMS_3, LAG_3, and MSK models, and five types of optimization sensitive parameters: SM, CS, LAG, X, and MP. Among them, SM represents the soil moisture sensitive parameter, which is used to describe the degree of soil response to changes in water content; CS represents the river channel storage sensitive parameter, which is used to describe the river channel's capacity to store and release runoff; LAG represents the watershed confluence delay sensitive parameter, which is used to describe the delay time between rainfall arrival and runoff response; X represents the flow proportion coefficient, which is used to adjust the proportional relationship of each link in the model; MP represents the number of sub-river segments, which reflects the degree of shift of the flood process. (33) Return query results: Adjust the optimization sensitivity parameters and directions of the water conservancy objects; The implementation process of step (4) is as follows: (41) Online scene feature extraction: extract rainfall and flow features that significantly affect parameter optimization during the preheating period of the watershed, and realize the transformation of scene data into feature data; (42) Matching the model library to optimize the parameter range: Based on the spatiotemporal characteristics of the watershed and the matching of multiple types of models in the offline model library, the most similar model is output as the optimized parameter range; (43) Update parameter range: Compare with the parameter optimization range obtained in step (3). If they are inconsistent, update the parameter optimization range to the pattern library matching result and update the parameter range file. The implementation process of step (6) is as follows: (61) Evaluate forecast accuracy: After optimizing the parameters of the operation forecast, extract the peak time, rise time, receding time, total flood volume, peak flow, and rainstorm center point sequence features, calculate the difference between these features and the measured features, and the difference is considered to be successful if it is within the threshold range; (62) Extracting new patterns: Based on the parameter optimization range decision process record of the calibration algorithm, analyze historical flood samples, extract the features of the preheating period samples, including peak flow, peak time, rise time, receding time, lag time, and rainstorm center point sequence, and map the features, parameter optimization range, and forecast accuracy evaluation results to the pattern; the formal expression of the features and parameter optimization range is as follows: in, Let the set of characteristics of the k-th flood be represented. The parameter optimization range and forecast accuracy assessment results for the k-th flood are represented by the formula F, which indicates the mapping relationship F between the feature set X and the parameter model Y. The parameter optimization model is defined as F(X,Y). The set of multiple parameter optimization models forms the model library D. The generated new models are filled into the model library to gradually form a parameter optimization model library D adapted to the characteristics of the current watershed. (63) Update the model library: After optimizing the parameters of the operation forecast, evaluate the forecast accuracy in a timely manner, record the decision-making process of the parameter optimization range, generate a new "preheating period conditions-parameter optimization range-forecast accuracy evaluation" mapping sample, provide feedback to supplement the parameter optimization rules, and gradually form a parameter optimization model library that adapts to the current characteristics of the watershed, thereby improving the success rate of parameter correction.
2. The online optimization method for flood forecasting model parameters driven by knowledge and data fusion according to claim 1, characterized in that, The implementation process of step (1) is as follows: (11) Flood events are divided into different time series, and the complex changes in floods are analyzed. Long-term flood time series data are processed into short-term flood time series data: The argrelmax function is used to find local extreme points in the time series data, and extreme points with flow values greater than 100 and time intervals greater than 24 hours are selected. The data is divided into segments with a window value of 10 according to the actual situation of the watershed, and process labels are applied, with 0 indicating that it is not a flood process and 1 indicating that it is a flood process. The flood process segmentation is manually corrected where it is unclear. (12) Extraction of rainstorm and flood features: extract rainfall, flow and watershed features that affect parameter optimization during the preheating period flow process, realize the transformation of scene data to feature data, and describe the different roles of rainfall, flow and watershed in parameter optimization.
3. The online optimization method for flood forecasting model parameters driven by knowledge and data fusion according to claim 1, characterized in that, The implementation process of step (2) is as follows: (21) Spatial contribution analysis determines the water conservancy targets for adjustment. By analyzing the regional composition and rainfall situation of flood events during the preheating period, the parameters adjustment sections or intervals are determined: Extract flood features for each object to form a feature matrix X; calculate weights based on the extracted features and their weights: in, This is the weight matrix. This represents the weight value of each feature. The characteristic matrix, This represents the value of each feature; based on the weight calculation results, the global contribution is sorted, and the node or region with the largest contribution is determined as the water conservancy object to be adjusted, denoted as UnitNo; (22) Analyze and determine the adjustment characteristic quantities, and identify the sensitive parameters to be optimized based on the deviation of flood characteristic quantities: Set thresholds for peak time difference, flood peak height difference, and their priorities; extract the peak time difference and flood peak difference features, compare them with the thresholds respectively, and record whether adjustments are needed; if both are greater than the thresholds, calculate the weights of the two priorities, and the one with the larger weight is the feature that needs adjustment, denoted as FeatureNo; (23) Determine the adjustment object model. Based on the UnitNo calculated in step (21) and the FeatureNo calculated in step (22), use them as keys to query the predefined parameter optimization rules in the knowledge graph and return the key value, which is the identifier of the water conservancy object model.
4. The online optimization method for flood forecasting model parameters driven by knowledge and data fusion according to claim 1, characterized in that, The implementation process of step (5) is as follows: (51) Execute the hydrological parameter calibration algorithm: Start the calibration service, call the function multiValueMapParam, and output the optimal parameter values in the parameter calibration process; (52) Update computation time: Change the computation start time in the runfile and input files of each model from the warm-up period to the full duration; (53) Calculate the forecast segment flow: Call the function SingleStep to output the calculated flow of the forecast segment.
5. The online optimization method for flood forecasting model parameters driven by knowledge and data fusion according to claim 2, characterized in that, The implementation process of step (12) is as follows: Extract offline flow characteristics, including: peak flow, the maximum instantaneous flow during a flood event, i.e., the flow at the highest point on the flood process line; peak time, the time when the flood peak occurs; lag time, the difference between the measured peak time and the calculated peak time; peak difference, the difference between the measured peak flow and the calculated peak flow; flood initiation time, the time when the flood process begins; and receding time, the time when the flood process ends. Extracting offline rainfall features, including: the storm center point sequence, which is a sequence in tuple form. ,in This indicates the coordinates of the center point of the rainstorm, which is the location of the grid point with the maximum rainfall. This represents the time at the center of the rainstorm; the trajectory of the rainstorm center point is a sequence of tuples. Converting this sequence of rainstorm center points into edge form, the sequence... ,in ; Extract basic features of the offline watershed, including watershed area and the area enclosed by the watershed divide.
6. The online optimization method for flood forecasting model parameters driven by knowledge and data fusion according to claim 1, characterized in that, The implementation process of step (41) is as follows: (411) Online feature selection: including total flood volume, peak flow, rise time, receding time, duration, peak water level, rise duration, receding duration, and rainfall feature storm center point sequence; (412) Online feature extraction: Total flood volume, the total amount of water flowing out of the outlet section of the basin within a certain duration; Peak flow, the flow rate at the moment the flood basin reaches its peak; Onset time, for the time nodes of the flood process, the time when the flood process begins is taken as the onset time of the flood; Receding time, the time when the flood process ends is taken as the receding time of the flood; Duration, the flood receding time minus the start time is taken as the duration of the flood; Peak water level, the highest point of water level during the flood process is taken as the peak water level; Onset duration, the time from the start time of the flood to the peak; Receding duration, the time from the last flood peak to the end time of the flood; Basin rainfall characteristics, including the sequence of rainstorm center points: a sequence in tuple form. ,in This indicates the coordinates of the center point of the rainstorm, which is the location of the grid point with the maximum rainfall. The time of the storm's center point is represented; finally, the extracted features are used to form a feature matrix F, realizing the transformation from scene data to feature data.
7. The online optimization method for flood forecasting model parameters driven by knowledge and data fusion according to claim 1, characterized in that, The implementation process of step (42) is as follows: For non-time-series data (i.e., single numerical values) in the feature matrix F, the values at the first t-1 time points are automatically padded with zeros; for the extracted multiple hydrological features, ensemble weights are used to assign weights to different features: in, This represents the degree of trust that experts have in the k-th weight vector. The negative ideal solution of the index; construct the Lagrangian function using formula (2) and solve for it. Then by The integrated weight U is obtained; Given two multivariate time series, the real values of each column are used as coordinate parameters of points in an S-dimensional space. Point-to-point comparisons are performed by calculating the Euclidean distance between the two points. Then, DTW (Digital Time-Warping) is used to calculate the similarity of the hydrological multivariate time series. For the weighting part involved in the similarity calculation, an integrated weight calibration method suitable for hydrological multivariate time series is used to determine the weights. The multivariate fusion algorithm is as follows: set up , Then the DTW distance between X and Y is: in, Representative sequence and DTW distance, Let each dimension have a weight, and ; set up , Then the MDTW distance between X and Y is: The calculated distance is the similarity. By sorting the matching results, the process line with the highest similarity is selected. The model and parameter values obtained in this process are the results returned by the pattern library matching.