Hydrological flood information data intelligent detection method and system and medium
The intelligent detection system for hydrological flood data constructed through deep learning algorithms solves the problem of long-time series hydrological data processing on multiple sites, realizes high-precision outlier detection and prediction, and improves the real-time and accuracy of water intelligence flood data.
Patent Information
- Application Number
- CN202510255338.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art is difficult to effectively process hydrological flood data from multiple sites and long time series, especially in the frequent occurrence of flood and drought disasters, which cannot meet the requirements of real-time and accuracy, and insufficient manual experience judgment.
Using intelligent algorithms based on deep learning, through data preprocessing, feature extraction, feature fusion and prediction model training, an intelligent detection system for hydrological flood reporting data is built, and combined with the upstream and downstream water flow relationships, outlier value detection is realized.
It has improved the detection accuracy and real-time processing capabilities of hydrological flood data, scientifically predicted future water level flow changes in sections, and supported the quality management of water conditions data in the basin.
Smart Images

Figure CN120296336A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of hydrology and water resources, and particularly to an intelligent detection method, system and medium for hydrological flood reporting data. Background Technique
[0002] Accurate hydrological flood reporting data is an important basis for exploring the propagation law of river flood runoff, making water regime and rainfall forecasts, and scheduling reservoir groups, and is an important data basis for supporting basin flood and drought disaster prevention and water resources management. However, with the development of automatic reporting equipment and the continuous optimization and improvement of the hydrological station network, the real-time performance, efficiency of hydrological flood reporting data, and the continuous increase in data volume have put forward higher requirements and challenges for the management of hydrological flood reporting data.
[0003] In the past, the operation and maintenance of hydrological flood reporting data mostly relied on manual experience and was judged based on the water regime characteristics of the station and the relationship between upstream and downstream; however, this technical idea is difficult to meet the monitoring requirements of multiple stations and long time series. Especially in the context of frequent flood and drought disasters, there is an urgent need to find new solutions and technologies to meet the new requirements of flood reporting. Therefore, how to integrate intelligent algorithms to replace manual experience, realize intelligent detection of hydrological flood reporting data, and real-time processing of abnormal data is the key problem faced by the current sound quality management system of hydrological flood reporting data. Summary of the Invention
[0004] The purpose of the embodiments of this application is to overcome the deficiencies of the prior art, and provide an intelligent detection method, system and medium for hydrological flood reporting data, which are fully based on the time series characteristics of hydrological flood reporting data of river cross-sections, and consider the water flow evolution process of the upstream and downstream relationships, strengthen its prediction characteristics, so as to replace manual experience judgment and improve the detection accuracy.
[0005] To achieve the above purpose, this application provides the following technical solutions:
[0006] In the first aspect, the embodiments of this application provide an intelligent detection method for hydrological flood reporting data, including the following steps:
[0007] S1. Data preprocessing. To ensure that the data effectively covers the main hydrological characteristics of the basin, the flood reporting data of the water regime stations at the main control sections of the main stream of the basin are used as the model training data set;
[0008] S2. Feature extraction. Extract hydrological time series features, utilize the time series characteristics of hydrological flood reporting data, consider the periodic patterns and trend changes in the data, and effectively extract the time series features of the time series; consider the spatial distribution, extract spatial features, mark them according to the upstream and downstream relationships, and at the same time extract the water volume propagation time according to the peak change situation, and extract the corresponding upstream and downstream relationship influence features;
[0009] S3. Feature fusion, making full use of the time series characteristics and introducing the attention mechanism to incorporate the upstream and downstream relationships of spatial feature extraction, considering the input delay and prediction step size, using a fully connected network for dimension alignment, and forming a feature set after normalization processing;
[0010] S4. Prediction model training, designing the input layer, training layer, and output layer of the model, using the dataset information divided into training set and validation set as training data, through multiple iterations and cross-validation, and evaluating the model accuracy through evaluation metrics to form a hydrological sequence prediction model;
[0011] S5. Threshold interval calculation, based on the prediction results, evaluating the accuracy according to the model output effect, and calculating the confidence level as the effective threshold interval for the station water regime data;
[0012] S6. Outlier judgment, comparing the real-time flood reporting data with the data in the predicted threshold interval to check whether it exceeds the effective threshold interval;
[0013] If it exceeds the effective threshold interval, it is determined as an outlier and handed over to manual intervention;
[0014] If it does not exceed the effective threshold interval, it is determined as a normal value and can be directly stored in the database.
[0015] In the above S1, data preprocessing is a process of cleaning and processing the dataset to ensure the integrity and effectiveness of the data quality for model training, so as to achieve better simulation calculations. The main steps include:
[0016] S11. Analyze the physical structure of the upstream and downstream of the flood reporting stations in the basin, quantitatively describe the topological composition characteristics of the flood reporting stations in the basin, and determine the equation as:
[0017] R = {r1, r2, r3,..., r i ,..., r N};
[0018] Among them, R is the composition characteristic of the flood reporting stations in the basin; r i is the i-th basic component unit of the system, that is, the flood reporting station in the topological structure, which is a multi-dimensional vector representing the location and the number of upstream and downstream units, and the dimension is determined by the specific number of characterization factors; i is the serial number of the basic component unit of the system, i = 1, 2,..., N; N is the number of basic component units, that is, the dataset should include the characteristics and structures of the flood reporting data in the time scale and space scale;
[0019] S12. Null value detection, due to different section times of different flood reporting stations, there are cases where some data are missing in a unified manner, and marks are made for the detected null values;
[0020] S13. Linear interpolation of time series data. To ensure data continuity, data filling is performed on the attribute columns with a small amount of missing data. For continuous data changes, adjacent data is selected for linear filling, as shown in formula (1);
[0021]
[0022] In the formula, it is the data of a certain attribute with a current vacancy, x i+1 is the data adjacent to the right of the missing data, x i-1 is the data adjacent to the left of the missing data;
[0023] S14. Filling of discrete data with the mode. For discrete data with a mode, the mode is selected for filling;
[0024] x filled = mode(X) (2)
[0025] In the formula, mode(X) represents the mode in the X attribute;
[0026] S15. Remove the entire column of attributes with a large amount of missing data.
[0027] In S2, feature extraction is based on the flood reporting indicators collected by the basin stations, including historical flood reporting data such as water level, flow rate, and water regime. They are sorted in the same time order, and the time series features of the flood reporting stations are extracted in the time order. At the same time, based on the upstream and downstream relationship of the spatial distribution, their upstream and downstream relationship and water flow propagation characteristics are extracted, effectively increasing the training features. At the same time, the extracted features are differentiated and normalized to form an effective feature set. The specific steps are as follows:
[0028] S21. Time series feature extraction. Based on the flood reporting indicators collected by the basin stations, including water level, flow rate, and water regime, the historical flood reporting data is sorted in the same time order. Multiple branches are constructed to process features of different time scales. The convolutional layer is used to extract short-term features, the fully connected layer is used to extract long-term features, and the embedding template is used to capture time interactions; the time series features of the flood reporting stations are extracted in the time order;
[0029] S22. Utilize the relationship between upstream and downstream stations. Add the features of the upstream stations to the attributes of the downstream stations to increase their prediction features;
[0030] S23. Convert the time series features of the flood reporting elements of the upstream stations into lag features, considering their impact on the downstream stations. The main method is to use the water level and flow rate of the upstream stations the day before as the input features of the downstream reservoir according to the flood propagation time.
[0031] In S3, feature fusion is to form a feature set through fusion and processing based on the extracted feature set, providing a standardized feature interval for the next model training. The specific steps are as follows:
[0032] S31 uses a fully connected network for dimension alignment;
[0033] S32 designs two factorization modules to capture dependencies. Among them, the time factorization module decomposes the time series to identify and extract time patterns; the channel factorization module processes the correlations between different variables to reduce redundant information;
[0034] S33 converts the time series features of flood reporting elements at upstream stations into lag features, considering their impacts on downstream stations. The main method is to use the water level and flow of the upstream station the previous day as the input features of the downstream reservoir according to the flood propagation time;
[0035] S34 uses 24 time steps as input features to better reflect the dynamic characteristics of water level changes, thereby enhancing the learning ability of the model, that is, the lag feature X logged ;
[0036] X logged ={X(t - n), X(t - n + 1), …, X(t - 1)} (3)
[0037] In the formula, n is the historical time step, and X(t - 1) is the target variable at the past time point;
[0038] S35 performs output fusion, incorporates the attention mechanism, and uses the methods of concatenation and weighted average to integrate the information in the time and channel dimensions to form a comprehensive feature representation;
[0039] S36 performs differential processing on the feature data, performs differential processing on the original time series to eliminate the trend component and enhance the stationarity of the data;
[0040] S37 performs data normalization processing, scales the data to the [-1, 1] interval to improve the stability and convergence speed of model training,
[0041]
[0042] where X norm is the normalized data.
[0043] In S4, the prediction model training is a process of training a station prediction model based on the flood reporting feature values of water information. The specific steps are as follows:
[0044] S41 designs the input layer to receive the preprocessed multivariate time series data;
[0045] S42 divides the dataset. The data is divided according to 7:3, that is, 70% is used as the training set and 30% is used as the test set;
[0046] S43. Output fusion, using concatenation and weighted averaging methods, integrates information in the temporal and channel dimensions to form a comprehensive feature representation;
[0047] S44. Use a fully connected layer to output the predicted value and perform denormalization processing on the predicted value;
[0048] S45. Use cross-validation to verify the model accuracy;
[0049] S46. Calculate the mean squared error (MSE) and mean absolute error (MAE) metrics to evaluate and optimize the model.
[0050] In the above S5, the threshold interval calculation is based on the prediction result, and the accuracy is evaluated according to the model output effect. By calculating the confidence level as the effective threshold interval of the station water regime data, the specific steps are as follows:
[0051] S51. Output the prediction result for the next 24 hours according to the input actual data, and calculate the standard deviation of the prediction error;
[0052] S52. Use the standard deviation and the quantiles of the normal distribution to calculate the upper and lower limits of the confidence interval, and determine the interval distribution of the prediction result.
[0053] In the second aspect, the embodiments of the present application provide a hydrological flood reporting data intelligent detection system, including
[0054] A data preprocessing module, using the flood reporting data of the water regime stations at the main control sections of the main stream of the basin as the model training data set;
[0055] A feature extraction module, utilizing the time series characteristics of the hydrological flood reporting data, considering the periodic patterns and trend changes in the data, effectively extracts the time series temporal features; considering the spatial distribution, extracts spatial features, marks them according to the upstream and downstream relationships, and at the same time extracts the water volume propagation time according to the peak change situation, and extracts the corresponding upstream and downstream relationship influence features;
[0056] A feature fusion module, making full use of the time series characteristics and introducing the attention mechanism to introduce the upstream and downstream relationships extracted from the spatial features, considering the input delay and prediction step length, uses a fully connected network for dimension alignment, and after normalization processing, forms a feature set;
[0057] A prediction model training module, designs the model input layer, training layer, and output layer, divides the data set information into training sets and validation sets as training data, through multiple iterations and cross-validation, and evaluates the model accuracy through evaluation metrics to form a hydrological sequence prediction model;
[0058] A threshold interval calculation module, based on the prediction result, evaluates the accuracy according to the model output effect, and calculates the confidence level as the effective threshold interval of the station water regime data;
[0059] An outlier judgment module compares the real-time flood reporting data with the data in the predicted threshold interval to determine whether it exceeds the effective threshold interval.
[0060] In a third aspect, an embodiment of the present application provides an intelligent detection device for hydrological flood reporting data, characterized in that the intelligent detection device for hydrological flood reporting data includes: a memory and at least one processor, and instructions are stored in the memory; at least one of the processors calls the instructions in the memory to enable the intelligent detection device for hydrological flood reporting data to execute each step of the above-mentioned intelligent detection method for hydrological flood reporting data.
[0061] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, and the computer-readable storage medium stores program codes, and when the program codes are executed by a processor, the steps of the above-mentioned intelligent detection method for hydrological flood reporting data are realized.
[0062] Compared with the prior art, the beneficial effects of the present invention are:
[0063] Fully referring to the manual judgment experience, according to the hydraulic connection between upstream and downstream and the characteristics of river channel composition, extracting the time series characteristics and water flow propagation characteristics of hydrological flood reporting stations, and based on intelligent algorithms such as deep learning, constructing an effective prediction model through processes such as data preprocessing, feature extraction, and model training. After evaluating the prediction results, an effective threshold interval is formed as the standard interval for data comparison, and the outlier detection of basin flood reporting data is completed by comparing the flood reporting data with the threshold. Through this technology, the future multi-step water level and flow time series change process of the cross-section can be scientifically predicted, the outlier detection of hydrological flood reporting data can be realized, and important technical support is provided for the quality management of basin water regime data. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0065] Figure 1 It is the flowchart of the method of the present application;
[0066] Figure 2 It is the schematic diagram of the prediction model structure;
[0067] Figure 3 It is the schematic diagram of the single-station prediction result;
[0068] Figure 4 It is the system block diagram of the present application;
[0069] Figure 5 This is the device block diagram of the present application. Specific embodiments
[0070] Next, the technical solutions in the embodiments of the present application will be described in conjunction with the accompanying drawings in the embodiments of the present application. It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0071] The term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "including a..." does not exclude the presence of additional identical elements in the process, method, article or device including the element.
[0072] The terms "first", "second", etc. are only used to distinguish one entity or operation from another entity or operation, and cannot be construed as indicating or implying relative importance, nor can it be construed as requiring or implying any actual relationship or order between these entities or operations.
[0073] Please refer to Figure 1 , a method for detecting anomalies in hydrological flood reporting data based on intelligent algorithms, including:
[0074] S1. Data preprocessing. To ensure that the data effectively covers the main hydrological characteristics of the basin, the flood reporting data of the main control section hydrological stations of the main stream of the basin is used as the model training data set.
[0075] S2. Feature extraction. On the one hand, hydrological time series feature extraction is carried out, making full use of the time series characteristics of the hydrological flood reporting data, considering the periodic patterns and trend changes in the data, and effectively extracting the time series features; on the other hand, considering the spatial distribution, spatial features are extracted, marked according to the upstream and downstream relationship, and at the same time, the water volume propagation time is extracted according to the peak change situation, and the corresponding upstream and downstream relationship influence features are extracted.
[0076] S3. Feature fusion. Making full use of the time series characteristics, introducing the upstream and downstream relationship extracted by the spatial feature through the fusion and attention mechanism, considering the input delay and prediction step, using a fully connected network for dimension alignment, and after normalization processing, forming a feature set to provide a standardized feature interval for the next model training.
[0077] S4. Prediction model training. Design the input layer, training layer, and output layer of the model. After dividing the dataset information into a training set and a validation set, use it as training data. Through multiple iterations and cross-validation, and evaluate the model accuracy using evaluation metrics to form a hydrological sequence prediction model.
[0078] S5. Threshold interval calculation. Based on the prediction results, evaluate the accuracy according to the model output effect, and calculate the confidence level as the effective threshold interval for the station water regime data.
[0079] S6. Outlier judgment. Compare the real-time flood reporting data with the data in the predicted threshold interval to check if it exceeds the effective threshold interval.
[0080] If it exceeds the effective threshold interval, it is determined as an outlier and handed over to manual intervention.
[0081] If it does not exceed the effective threshold interval, it is determined as normal and can be directly stored in the database.
[0082] In the above S1, data preprocessing is a process of cleaning and processing the dataset to ensure the integrity and effectiveness of the data quality for model training, so as to achieve better simulation calculations. The main steps include:
[0083] S11. Analyze the physical structures upstream and downstream of the flood reporting stations in the basin, quantitatively characterize the topological composition characteristics of the flood reporting stations in the basin, and determine the equation as:
[0084] R = {r1, r2, r3,..., r i ,..., r N};
[0085] Among them, R is the composition characteristic of the flood reporting stations in the basin; r i is the i-th basic composition unit of the system, that is, the flood reporting station in the topological structure, which is a multi-dimensional vector representing the location and the number of upstream and downstream units. The dimension is determined by the specific number of characterization factors; i is the serial number of the basic composition unit of the system, i = 1, 2, 1,......, N; N is the number of basic composition units, that is, the dataset should contain the characteristics and structures of the flood reporting data in the time scale and space scale.
[0086] S12. Null value detection. Since the number of segments of different flood reporting stations is different, 24 segments, that is, the data step size is 1 hour, and there are partial data gaps in accordance with the unified situation. Mark the detected null values.
[0087] S13. Linear interpolation of time series data. To ensure the continuity of the data, perform data filling processing on the attribute columns containing a small amount of missing data. For continuous data changes (such as water level above the reservoir, discharge from the reservoir, etc.), select adjacent data for linear filling, as shown in formula (1);
[0088]
[0089] where x filled is the data of a certain property that is currently vacant, and x i+1 is the data immediately adjacent to the right of the vacant data, and x i-1 is the data immediately adjacent to the left of the vacant data;
[0090] S14. Filling with the mode of discrete data. For discrete data with a mode, select to use the mode for filling;
[0091] x filled = mode(X) (2)
[0092] where mode(X) represents the mode in the X attribute;
[0093] S15. Removing the entire column of the attribute with a large number of vacant data;
[0094] In the above S2, feature extraction is to sort the flood reporting indicators collected at the basin stations, including historical flood reporting data such as water level, flow rate, and water regime, in a unified time sequence, extract the time series features of the flood reporting stations in the time sequence, and at the same time extract their upstream and downstream relationships and water flow propagation features based on the upstream and downstream relationships of the spatial distribution, effectively increasing the training features, and at the same time performing difference and normalization processing on the extracted features to form an effective feature set. The specific steps are as follows:
[0095] S21. Time series feature extraction. Sort the flood reporting indicators collected at the basin stations, including historical flood reporting data such as water level, flow rate, and water regime, in a unified time sequence, construct multiple branches to process features of different time scales, use a convolutional layer (Conv1D) to extract short-term features, use a fully connected layer to extract long-term features, and use an embedding template to capture time interactions; extract the time series features of the flood reporting stations in the time sequence;
[0096] S22. Utilize the relationship between upstream and downstream stations, add the features of the upstream stations to the attributes of the downstream stations to increase their prediction features;
[0097] S23. Convert the time series features of the flood reporting elements of the upstream stations into lag features and consider their impact on the downstream stations. The main method is to use the water level and flow rate of the upstream stations the day before as the input features of the downstream reservoir according to the flood propagation time.
[0098] In the above S3, feature fusion mainly utilizes the time series characteristics, introduces the upstream and downstream relationships extracted by the spatial feature extraction through the fusion attention mechanism, considers the input delay and prediction step length, and uses a fully connected network for dimension alignment to form a feature set, providing a standardized feature interval for the next model training.
[0099] S31. Using 24 time steps as input features can better reflect the dynamic characteristics of water level changes, thereby enhancing the learning ability of the model. That is, the lag feature X logged ;
[0100] X logged ={X(t - n), X(t - n + 1), …, X(t - 1)} (3)
[0101] where n is the historical time step. X(t - 1) is the target variable at the past time point;
[0102] S32. Feature data difference processing, performing difference processing on the original time series to eliminate the trend component and enhance the stationarity of the data;
[0103] S33. Data normalization processing, scaling the data to the [-1, 1] interval to improve the stability and convergence speed of model training.
[0104]
[0105] where X norm is the normalized data;
[0106] In S4, the prediction model training is a process of training a station prediction model based on the water information flood control characteristic values. The specific steps are as follows:
[0107] S41. Input layer design, designing the input layer to receive the preprocessed multi - variable time series data;
[0108] S42. Dataset division. Divide the data according to 7:3, that is, 70% as the training set and 30% as the test set;
[0109] S43. Output fusion, using the methods of concatenation and weighted average to integrate the information in the time and channel dimensions to form a comprehensive feature representation
[0110] S44. Using the fully - connected layer to output the prediction value and perform denormalization processing on the prediction value;
[0111] S45. Using cross - validation to verify the model accuracy;
[0112] S46. Calculating the MSE and MAE metrics to evaluate and optimize the model;
[0113] Furthermore, in S5, the threshold interval calculation is based on the prediction results, and the accuracy is evaluated according to the model output effect. By calculating the confidence level as the effective threshold interval of the station water information data. The specific steps are as follows:
[0114] S51. According to the input actual data, output the prediction results for the next 24 hours and calculate the standard deviation of the prediction error;
[0115] S52. Calculate the upper and lower limits of the confidence interval using the standard deviation and the quantiles of the normal distribution to determine the interval distribution of the prediction results.
[0116] Furthermore, in the above-mentioned S6, the anomaly detection is to analyze the change patterns and rules of the historical sequence data, and then judge whether the data is an outlier based on the difference between the prediction of the current observed value and the actual observed value. If it is determined to be an anomaly, an outlier reminder will be given in a visual way for manual intervention and modification. The specific steps are as follows:
[0117] S61. Compare the real-time flood reporting data with the data in the predicted threshold interval to check whether it exceeds the effective threshold interval.
[0118] S62. If it exceeds the effective threshold interval, it is determined as an outlier and handed over to manual intervention.
[0119] S63. If it does not exceed the effective threshold interval, it is determined as a normal value and can be directly stored in the database.
[0120] Figure 1 The following is the flow chart of the hydrological flood reporting data anomaly detection method based on intelligent algorithms, in which an MTS-Mixers is used to construct a water level and flow prediction model, and the model structure is as Figure 2 shown. A data set is constructed using the data of the Liantuo Station on the main stream of the Yangtze River as the main station in the past 3 years. After data cleaning, there are 32,562 data in this detection station, including 26,049 in the training set, 4,885 in the validation set, and 1,628 in the test set. After training, the model is obtained. After inputting the upstream and downstream water levels and flow data, as well as relevant water regime information into the model, the predicted data of the water level and flow can be obtained. Here, the last 200 data in the test set are shown, and the result graph is as Figure 3 shown.
[0121] For the 1,628 test sets, the prediction error results of the model are as follows:
[0122] Index Test result MSE 0.0106 RMSE 0.1032
[0123] Among them, the mean squared error (MSE) reflects the average of the squares of the errors between the predicted values and the true values, and the root mean squared error (RMSE) is used to evaluate the average level of the prediction errors.
[0124] According to the above analysis, it can be seen that the method of the present invention has strong practicability and can effectively solve the problem of hydrological flood reporting data anomaly detection by constructing an intelligent algorithm in a time series prediction manner.
[0125] As Figure 4 shown, the embodiment of the present application provides a hydrological flood reporting data intelligent detection system, including,
[0126] The data preprocessing module 1 uses the flood reporting data of the main control section hydrological stations of the main stream of the basin as the model training data set;
[0127] The feature extraction module 2 utilizes the time series characteristics of the hydrological flood reporting data, considers the periodic patterns and trend changes in the data, and effectively extracts the time series characteristics of the time series; considering the spatial distribution, extracts spatial features, marks them according to the upstream and downstream relationships, and at the same time extracts the water volume propagation time according to the peak change situation, and extracts the corresponding upstream and downstream relationship influence features;
[0128] The feature fusion module 3 makes full use of the time series characteristics and introduces the attention mechanism to introduce the upstream and downstream relationships extracted from the spatial features. Considering the input delay and prediction step length, a fully connected network is used for dimension alignment. After normalization processing, a feature set is formed;
[0129] The prediction model training module 4 designs the model input layer, training layer, and output layer. After dividing the data set information into the training set and the validation set as the training data, through multiple iterations and cross-validation, the model accuracy is evaluated by the evaluation index to form a hydrological sequence prediction model;
[0130] The threshold interval calculation module 5 is based on the prediction result, evaluates the accuracy according to the model output effect, and calculates the confidence level as the effective threshold interval of the station water regime data;
[0131] The outlier judgment module 6 compares the real-time flood reporting data with the data in the prediction threshold interval to determine whether it exceeds the effective threshold interval.
[0132] As Figure 5 shown, the embodiment of the present application provides an intelligent detection device for hydrological flood reporting data. The intelligent detection device for hydrological flood reporting data may vary greatly due to configuration or performance, and may include one or more processors (central processing units, CPUs) (for example, one or more processors) and a memory, and one or more storage media for storing application programs or data (for example, one or more mass storage devices). Among them, the memory and the storage medium can be short-term storage or persistent storage. The program stored in the storage medium may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the intelligent detection device for hydrological flood reporting data. Further, the processor can be set to communicate with the storage medium and execute a series of instruction operations in the storage medium on the intelligent detection device for hydrological flood reporting data to implement the steps of the intelligent detection method for hydrological flood reporting data provided by the above method embodiments.
[0133] The intelligent detection device for hydrological flood reporting data may further include one or more power supplies, one or more wired or wireless network interfaces, one or more input / output interfaces, and / or one or more operating systems, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art can understand that Figure 5 The structure of the intelligent detection device for hydrological flood reporting data shown does not constitute a limitation on the intelligent detection device based on hydrological flood reporting data, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0134] The embodiment of the present application provides a computer-readable storage medium, and the computer-readable storage medium stores program codes. When the program codes are executed by a processor, the steps of the above-mentioned method for intelligent detection of hydrological flood reporting data are implemented.
[0135] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0136] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0137] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0138] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, causing a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one process Figure 1 one process or more processes and / or blocks Figure 1 steps for implementing the functions specified in one block or more blocks.
[0139] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0140] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0141] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0142] The above are only examples of the embodiments of the present application and are not intended to limit the scope of protection of the present application. For those skilled in the art, the present application may have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the scope of protection of the present application.
Claims
1. An intelligent detection method for hydrological flood reporting data, characterized in that, It includes the following steps: S1. Data preprocessing: To ensure that the data effectively covers the main hydrological characteristics of the basin, the flood reporting data of the main control section hydrological stations of the main stream of the basin are used as the model training data set. S2. Feature extraction: Hydrological time series feature extraction. Utilize the time series characteristics of the hydrological flood reporting data, consider the periodic patterns and trend changes in the data, and effectively extract the time series features of the time series. Consider the spatial distribution, extract spatial features, mark them according to the upstream and downstream relationships, and at the same time extract the water volume propagation time according to the peak change situation, and extract the corresponding upstream and downstream relationship influence features. S3. Feature fusion: Make full use of the time series characteristics and introduce the upstream and downstream relationships extracted by the attention mechanism for spatial features. Consider the input delay and prediction step length, use a fully connected network for dimension alignment, and after normalization processing, form a feature set. S4. Prediction model training: Design the input layer, training layer, and output layer of the model. Divide the data set information into training data and validation data after partitioning, and through multiple iterations and cross-validation, evaluate the model accuracy through evaluation indicators to form a hydrological sequence prediction model. S5. Threshold interval calculation: Based on the prediction results, evaluate the accuracy according to the model output effect, and calculate the confidence level as the effective threshold interval for the station water regime data. S6. Outlier judgment: Compare the real-time flood reporting data with the data in the predicted threshold interval to check whether it exceeds the effective threshold interval.
2. The intelligent detection method for hydrological flood reporting data according to claim 1, wherein In S1, data preprocessing is a process of cleaning and processing the data set to ensure the integrity and effectiveness of the data quality for model training, so as to achieve better simulation calculations. The main steps include: S11. Analyze the physical structure of the upstream and downstream of the basin flood reporting stations, quantitatively describe the topological composition characteristics of the basin flood reporting stations, and determine the equation as: R={r1,r2,r3,...,r i ,...,r N}; wherein, R is the composition characteristic of basin flood reporting stations; r i is the i-th basic composition unit of the system, that is, the flood reporting station in the topological structure, which is a multi-dimensional vector representing the location and the number of upstream and downstream units, and the dimension is determined by the specific number of characterization factors; i is the serial number of the basic composition unit of the system, i = 1, 2,......, N; N is the number of basic composition units, that is, the dataset should contain the characteristics and structure of flood reporting data in the time scale and space scale; S12. Null value detection: Since the section times of different flood reporting stations are different, there is a situation where some data are missing in a unified manner. Mark the detected null values. S13. Linear interpolation of time series data: To ensure the continuity of the data, perform data filling processing on the attribute columns containing a small amount of missing data. For continuous data changes, select adjacent data for linear filling, as shown in formula (1). where x is the data of a certain property that is currently missing i+1 is the data immediately adjacent to the right of the missing data, x i-1 is the data immediately adjacent to the left of the missing data; S14. Mode filling of discrete data: For discrete data with a mode, choose to use the mode for filling. x filled = mode(X) (2) In the formula, mode(X) represents the mode in the X attribute. S15. Remove the entire column of attributes with a large amount of missing data.
3. The intelligent detection method for hydrological flood reporting data according to claim 1, characterized in that, In S2, feature extraction is to sort the historical flood reporting data collected by the basin stations, including water level, flow rate, water potential, etc., in a unified time sequence, extract the time series features of the flood reporting stations in the time sequence, and at the same time extract their upstream and downstream relationships and water flow propagation features based on the upstream and downstream relationships of the spatial distribution, effectively increasing the training features. At the same time, perform difference and normalization processing on the extracted features to form an effective feature set. The specific steps are as follows: S21. Temporal feature extraction: Using the flood reporting indicators collected at basin stations, including water level, flow rate, and water regime, sort the historical flood reporting data in the unified time sequence, construct multiple branches to process features of different time scales, use convolutional layers to extract short-term features, use fully connected layers to extract long-term features, and use embedding templates to capture temporal interactions; Extract the temporal features of flood reporting stations in chronological order; S22. Utilize the relationship between upstream and downstream stations, add the features of upstream stations to the attributes of downstream stations to increase their predictive features; S23. Convert the temporal features of flood reporting elements of upstream stations into lag features, considering their impact on downstream stations. The main method is to use the water level and flow rate of the upstream station the previous day as the input features of the downstream reservoir according to the flood propagation time.
4. The intelligent detection method for hydrological flood reporting data according to claim 1, characterized in that, In the above S3, feature fusion is to form a feature set through fusion and processing based on the extracted feature set, providing a standardized feature interval for the next model training. The specific steps are as follows: S31. Use a fully connected network for dimension alignment; S32. Design two factorization modules to capture dependencies. Among them, the temporal factorization module: decompose the time series, identify and extract temporal patterns; the channel factorization module: handle the correlation between different variables and reduce redundant information; S33. Convert the temporal features of flood reporting elements of upstream stations into lag features, considering their impact on downstream stations. The main method is to use the water level and flow rate of the upstream station the previous day as the input features of the downstream reservoir according to the flood propagation time; S34. Use 24 time steps as input features to better reflect the dynamic characteristics of water level changes, thereby enhancing the learning ability of the model, that is, the lag feature X logged ; X logged = {X(t - n), X(t - n + 1), …, X(t - 1)} (3) In the formula, n is the historical time step, and X(t - 1) is the target variable at the past time point; S35. Output fusion, incorporate the attention mechanism, and use the methods of concatenation and weighted average to integrate the information in the temporal and channel dimensions to form a comprehensive feature representation; S36. Feature data difference processing: Perform difference processing on the original time series to eliminate the trend component and enhance the stationarity of the data; S37. Data normalization processing: Scale the data to the interval [-1, 1] to improve the stability and convergence speed of model training. where X norm is the normalized data.
5. The intelligent detection method for hydrological flood reporting data according to claim 1, characterized in that, In the above S4, the prediction model training is the process of training a station prediction model based on the flood reporting feature values of water conditions. The specific steps are as follows: S41. Input layer design: Design the input layer to receive the preprocessed multivariate time series data; S42. Dataset division: Divide the data according to 7:3, that is, 70% as the training set and 30% as the test set; S43. Output fusion: Use the methods of concatenation and weighted average to integrate the information in the temporal and channel dimensions to form a comprehensive feature representation; S44. Use a fully connected layer to output the predicted value and perform inverse normalization processing on the predicted value; S45. Use cross-validation to verify the model accuracy; S46. Calculate the mean square error MSE and mean absolute error MAE indicators to evaluate and optimize the model.
6. The intelligent detection method for hydrological flood reporting data according to claim 1, wherein, In the above S5, the threshold interval calculation is based on the prediction results, and the accuracy is evaluated according to the output effect of the model. Calculate the confidence level as the effective threshold interval of the station water condition data. The specific steps are as follows: S51. Output the prediction results for the next 24 hours according to the input actual data, and calculate the standard deviation of the prediction error; S52. Calculate the upper and lower limits of the confidence interval using the standard deviation and the quantiles of the normal distribution to determine the interval distribution of the prediction results.
7. The intelligent detection method for hydrological flood reporting data according to claim 1, wherein, In S5, the outlier judgment is specifically as follows: If it exceeds the effective threshold interval, it is determined as an outlier and handed over to manual intervention. If it does not exceed the effective threshold interval, it is determined as a normal value and directly stored in the database.
8. An intelligent detection system for hydrological flood reporting data, characterized in that, Including A data preprocessing module that uses the flood reporting data of the main control section water regime stations of the main stream of the basin as the model training data set. A feature extraction module that utilizes the time series characteristics of the hydrological flood reporting data, considers the periodic patterns and trend changes in the data, and effectively extracts the time series temporal features. Considering the spatial distribution, extract spatial features, mark them according to the upstream and downstream relationship, and at the same time extract the water volume propagation time according to the peak change situation, and extract the corresponding upstream and downstream relationship influence features. A feature fusion module that makes full use of the time series characteristics and introduces the upstream and downstream relationship extracted by the spatial features through the attention mechanism. Considering the input delay and prediction step size, a fully connected network is used for dimension alignment. After normalization processing, a feature set is formed. A prediction model training module that designs the model input layer, training layer, and output layer. After dividing the data set information into training sets and validation sets as training data, through multiple iterations and cross-validations, and evaluating the model accuracy through evaluation indicators, a hydrological sequence prediction model is formed. A threshold interval calculation module that, based on the prediction results, evaluates the accuracy according to the model output effect, and calculates the confidence level as the effective threshold interval of the station water regime data. An outlier judgment module that compares the real-time flood reporting data with the predicted threshold interval data to determine whether it exceeds the effective threshold interval.
9. An intelligent detection device for hydrological flood reporting data, characterized in that, The hydrological flood reporting data intelligent detection device includes: a memory and at least one processor, and instructions are stored in the memory; at least one of the processors calls the instructions in the memory so that the hydrological flood reporting data intelligent detection device executes each step of the hydrological flood reporting data intelligent detection method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores program codes, and when the program codes are executed by the processor, the steps of the hydrological flood reporting data intelligent detection method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Data analysis management system and method based on false tide prediction model
CN120823695A
Flood session division method and system based on flood characteristics
CN120910483A
Irrigation area data anomaly analysis method and server
CN120995017A
An irrigation district data anomaly analysis method and server
CN120995017B