Supervised Temporal Water Level Data Generation Method, System, and Storage Medium
By building a conditional generation adversarial network (CGAN) and performing conditional feature regulation driven by hydrological variables, the problems of flood data imbalance and hydrological extreme event patterns are solved, and the performance and data quality of the flood warning system are improved.
Patent Information
- Application Number
- CN202510167665.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-17
AI Technical Summary
The existing technology is difficult to effectively identify high-risk flood events in flood warning systems, mainly due to the unbalanced flood data and the difficult pattern of hydrological extreme event.
A supervised time-sequential water level data generation method is adopted to obtain multi-source hydrological monitoring data, outlier value identification and data preprocessing, a conditional generation adversarial network (CGAN) is constructed, and the parameters of the generation network are optimized through hydrological variable-driven conditional feature regulation, and the conditional input feature data with enhanced hydrological adaptability are generated.
The performance of the model in identifying high-risk flood events is improved, the ability to generalize hydrological diversity scenarios is enhanced, the generated water level timing data has higher authenticity and correlation, and the quality and credibility of the data have also been significantly improved.
Smart Images

Figure CN119623531B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning, and in particular, to a supervised time-series water level data generation method, system and storage medium. Background Art
[0002] "Supervised" refers to a learning method in which a model learns from labeled training data. In supervised learning, each training sample has an associated output label, and the goal of the model is to learn the relationship between the input data (features) and the output label (target). Specifically, in the "supervised time-series water level data generation method", "supervised" means that in this case, the time-series water level data is labeled, that is, each water level data at a time point has a corresponding label, which may be the category, level of the water level or some other form of mark. The model learns to predict the pattern of water level changes by analyzing these labeled water level data. For example, it may learn how the water level will change under certain specific environmental conditions. Once the model is trained, it can be used to predict future water level changes or classify new water level data.
[0003] However, in flood warning systems, the current supervised time-series water level data generation methods often have the following problems: The flood data itself has highly imbalanced characteristics, and extreme flood events are relatively rare, resulting in far fewer positive samples (flood events) than negative samples (normal water levels) in the training data. This imbalance seriously affects the performance of the model in identifying high-risk flood events. The occurrence patterns of hydrological extreme events are difficult to predict. Traditional models lack the ability to adapt to this long-term hydrological change trend. Summary of the Invention
[0004] Based on this, it is necessary for the present invention to provide a supervised time-series water level data generation method and system to solve at least one of the above technical problems.
[0005] To achieve the above object, a supervised time-series water level data generation method includes the following steps:
[0006] Step S1: Obtain multi-source hydrological monitoring data; identify outliers and perform data preprocessing according to the hydrological monitoring data to generate standardized hydrological basic data; construct a data generation model of a conditional generative adversarial network based on the standardized hydrological basic data;
[0007] Step S2: Perform conditional feature regulation driven by hydrological variables on the data generation model, optimize the parameters of the generation network, and generate conditional input feature data with enhanced hydrological adaptability; inject the conditional input feature data into the data generation model to generate preliminary water level time-series data;
[0008] Step S3: Perform stratified importance sampling on the preliminary water level time series data, and set weights differentially according to the preset rare extreme hydrological event data to obtain optimized hydrological event sampling data; oversample the rare extreme hydrological event data by the bootstrap method to obtain high-quality extreme event data; perform distribution consistency evaluation based on statistical characteristics according to the extreme event data and the optimized hydrological event sampling data to obtain water level time series data;
[0009] Step S4: Perform multi-scale verification on the water level time series data to obtain water level time series verification data, where the multi-scale verification includes spectral characteristic test and mutual information measurement; eliminate the artifacts in the water level time series verification data, and perform correction processing on the extreme value outliers based on physical constraints, so as to obtain supervised time series water level data.
[0010] The present invention realizes the comprehensive acquisition of multi-channel hydrological information by obtaining multi-source hydrological monitoring data, which helps to comprehensively cover the data characteristics of different sources; on this basis, data noise and outliers are removed by using outlier identification and data preprocessing steps to generate standardized basic hydrological data, ensuring the accuracy and consistency of the data. This provides a high-quality input source for subsequent data modeling and avoids the error propagation of the original data. Further, a data generation model is constructed through a conditional generative adversarial network, enabling the model to have strong adaptability and be able to simulate complex hydrological characteristics and data distribution characteristics, improving the flexibility and controllability of data generation. By introducing conditional feature regulation driven by hydrological variables, the parameters of the data generation model are optimized, which can better capture the influence of hydrological conditions on hydrological changes, thereby generating conditional input feature data with hydrological adaptability. This process improves the generalization ability of the model in hydrological diversity scenarios; by injecting the conditional input feature data into the generation model, the generated preliminary water level time series data has high authenticity and relevance, laying a foundation for subsequent optimization processing. Through stratified importance sampling, the ability to identify key features of the preliminary water level time series data is improved, and differential weight setting is combined with rare extreme hydrological event data to optimize the sampling strategy and enhance the attention to extreme events; the bootstrap method is used to oversample rare event data, which not only expands the sample size but also improves data balance. The distribution consistency evaluation based on statistical characteristics further ensures the coincidence degree of the generated hydrological data with the real scenario, thereby obtaining more reliable water level time series data. Through multi-scale verification, including spectral characteristic test and mutual information measurement, the frequency domain characteristics and the dependence relationship between variables of the water level time series data are strictly tested, enhancing the scientificity and comprehensiveness of data verification; subsequently, artifact elimination and correction of extreme value outliers based on physical constraints are performed to ensure the physical rationality and consistency of the data. These measures effectively reduce the artifact interference in the generated data and improve the quality and credibility of the finally generated supervised time series water level data.
[0011] The present invention also provides a supervised time-series water level data generation system for performing the above-mentioned supervised time-series water level data generation method. The supervised time-series water level data generation system includes:
[0012] A model construction module, configured to obtain multi-source hydrological monitoring data; identify outliers and perform data preprocessing on the hydrological monitoring data to generate standardized hydrological basic data; construct a data generation model of a conditional generative adversarial network based on the standardized hydrological basic data;
[0013] A hydrological condition enhancement module, configured to perform conditional feature regulation driven by hydrological variables on the data generation model, optimize the parameters of the generation network, and generate condition input feature data with enhanced hydrological adaptability; inject the condition input feature data into the data generation model to generate preliminary water level time-series data;
[0014] A sampling optimization module, configured to perform stratified importance sampling on the preliminary water level time-series data, set differential generation weights according to preset rare extreme hydrological event data to obtain optimized sampling data for hydrological events; oversample the rare extreme hydrological event data by the bootstrap method to obtain high-quality extreme event data; perform distribution consistency evaluation based on statistical features according to the extreme event data and the optimized sampling data for hydrological events to obtain water level time-series data;
[0015] A data verification and correction module, configured to perform multi-scale verification on the water level time-series data to obtain water level time-series verification data, where the multi-scale verification includes spectral characteristic test and mutual information measurement; eliminate artifacts in the water level time-series verification data and perform correction processing on extreme abnormal points based on physical constraints to obtain supervised time-series water level data.
[0016] The model construction module in the present invention provides a solid data foundation for the subsequent generation of water level time series data by obtaining multi-source hydrological monitoring data. First, through outlier identification and data preprocessing, this module cleans the original data, removes the noise and errors that do not conform to the normal hydrological change law, and generates standardized hydrological basic data. This standardization process not only improves the consistency and usability of the data, but also ensures that subsequent modeling will not be interfered by abnormal data. Based on these standardized hydrological basic data, a data generation model of conditional generative adversarial network (GAN) is constructed, which can simulate the complex law of water level change and generate time series data that conforms to the actual hydrological characteristics, laying the foundation for the accuracy and reliability of the entire model. The hydrological condition enhancement module plays a key role in generating water level time series data. By regulating the conditional features of the data generation model driven by hydrological variables, this module can optimize the parameters of the generation network according to the actual situation of hydrological changes, so that the generated data is enhanced in hydrological adaptability. This hydrological adaptability-enhanced feature data ensures that the generated water level time series data can more realistically reflect the hydrological changes under different conditions and provides more accurate prediction results. The sampling optimization module is optimized during the data generation process, making the generated water level time series data more in line with the actual application requirements. By performing stratified importance sampling on the preliminary water level time series data, the representativeness of the generated data can be improved. Especially when facing rare extreme hydrological events, through the setting of differential generation weights, these events are preferentially sampled, so as to generate more accurate extreme event data. By oversampling through the bootstrap method, the quality of extreme event data can be further improved, enhancing the model's ability to identify these key events. Finally, these data are processed through distribution consistency evaluation based on statistical features, ensuring the accuracy and consistency of the water level time series data in different scenarios. The data verification and correction module comprehensively checks the generated water level time series data through multi-scale verification methods to ensure the reliability of the data. The spectral characteristic test and mutual information measurement provide two important perspectives for the verification process. The spectral characteristic verification can check the energy distribution characteristics of the data in different frequency ranges, and the mutual information measurement reveals the correlation and information flow between the data, which helps to detect whether the data conforms to the expected hydrological characteristics. Through artifact elimination and extreme value anomaly correction, the quality of the data is further improved, avoiding inaccurate results caused by measurement errors or noise. This process ensures that the generated water level time series data not only conforms to the actual hydrological law, but also has sufficient physical rationality and accuracy, and finally obtains high-quality supervised time series water level data. In summary, these modules cooperate with each other to provide a complete process from data acquisition, generation, optimization to verification and correction, ensuring the accuracy, reliability and applicability of the water level time series data, and being able to provide strong data support for applications such as hydrological prediction and risk assessment.
[0017] The present invention also provides a computer-readable storage medium storing a computer program, and the execution of the computer program implements the supervised time-series water level data generation method as described above. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Other features, objects, and advantages of the present invention will become more apparent by reading the detailed description of the non-limiting embodiments with reference to the following drawings:
[0019] Figure 1 It is a schematic flowchart of the steps of the supervised time-series water level data generation method of the present invention;
[0020] Figure 2 is Figure 1 a detailed schematic flowchart of step S1 in
[0021] Figure 3 is Figure 1 a detailed schematic flowchart of step S2 in DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] The technical method of the present invention will be clearly and completely described below with reference to the drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0023] In addition, the drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings represent the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.
[0024] It should be understood that although terms such as "first", "second", etc. may be used here to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit can be called the second unit, and similarly the second unit can be called the first unit. The term "and / or" used here includes any and all combinations of one or more of the listed related items.
[0025] To achieve the above object, please refer to Figures 1 to 3, the present invention provides a supervised method for generating time-series water level data, and the method includes the following steps:
[0026] Step S1: Obtain multi-source hydrological monitoring data; identify outliers and perform data preprocessing based on the hydrological monitoring data to generate standardized hydrological basic data; construct a data generation model of a conditional generative adversarial network based on the standardized hydrological basic data;
[0027] Step S2: Perform conditional feature regulation driven by hydrological variables on the data generation model, optimize the parameters of the generation network, and generate conditional input feature data with enhanced hydrological adaptability; inject the conditional input feature data into the data generation model to generate preliminary water level time-series data;
[0028] Step S3: Perform stratified importance sampling on the preliminary water level time-series data, and set differential generation weights according to the preset rare extreme hydrological event data to obtain hydrological event optimized sampling data; oversample the rare extreme hydrological event data by the bootstrap method to obtain high-quality extreme event data; perform distribution consistency evaluation based on statistical characteristics according to the extreme event data and the hydrological event optimized sampling data to obtain water level time-series data;
[0029] Step S4: Perform multi-scale verification on the water level time-series data to obtain water level time-series verification data, where the multi-scale verification includes spectral characteristic test and mutual information measurement; eliminate artifacts in the water level time-series verification data and perform correction processing on extreme abnormal points based on physical constraints to obtain supervised time-series water level data.
[0030] In the embodiment of the present invention, refer to Figure 1 As shown, it is a schematic diagram of the step flow of a supervised method for generating time-series water level data of the present invention. In this example, the supervised method for generating time-series water level data includes the following steps:
[0031] Step S1: Obtain multi-source hydrological monitoring data; identify outliers and perform data preprocessing based on the hydrological monitoring data to generate standardized hydrological basic data; construct a data generation model of a conditional generative adversarial network based on the standardized hydrological basic data;
[0032] In an embodiment of the present invention, hydrological monitoring data of a certain basin is obtained, including water level, rainfall, evaporation, and flow information recorded at 5 monitoring stations. The monitoring period is from January 2020 to December 2023, and the acquisition frequency is once per hour. These original hydrological monitoring data are subjected to format conversion and timestamp calibration to ensure that all data times are consistent. Using outlier identification technology based on statistical methods, outliers exceeding 3 times the standard deviation in the water level data are removed, and the Lagrange interpolation method is used to fill in the missing values. Subsequently, based on the time series regression model of adjacent data points, the data is further interpolated and repaired to obtain continuous cleaned hydrological monitoring data. For the cleaned data, the mutual information method is used to screen rainfall and evaporation characteristics with a high correlation with the temporal variation of the water level height, and derived variables such as daily rainfall intensity and cumulative evaporation are generated. Finally, the extended feature data is dimensionless processed and segmented by week to obtain standardized hydrological basic data. Using the standardized data as input, a conditional generative adversarial network is constructed, where the generator adopts a hybrid network structure of a double-layer LSTM and a one-dimensional convolution, and the discriminator is designed with a consistency detection module based on the time and frequency domains, and distribution similarity loss, temporal consistency loss, and physical constraint loss functions are defined to complete the construction of the data generation model.
[0033] Step S2: Perform conditional feature regulation driven by hydrological variables on the data generation model, optimize the parameters of the generation network, and generate conditional input feature data with enhanced hydrological adaptability; inject the conditional input feature data into the data generation model to generate preliminary water level time series data;
[0034] In an embodiment of the present invention, historical hydrological variable data provided by meteorological stations, including temperature (range: -5°C to 40°C), humidity (30% to 90%), precipitation (0 to 200 mm), and atmospheric pressure (950 to 1050 hPa), are subjected to principal component analysis (PCA) to extract key hydrological features with a variance contribution rate of more than 95%. Using a hydrological variable encoding network based on the attention mechanism, the extracted hydrological features are embedded into 32-dimensional vectors. Subsequently, the parameters of the generator network are adaptively optimized through variational inference methods to enhance the model's adaptability to hydrological features. Finally, the generated conditional input feature data is injected into the generator network for adversarial training to generate preliminary water level time series data with a length of 365 days, where the time resolution of the data is hourly and the data distribution simulates the real historical hydrological trend.
[0035] Step S3: Perform stratified importance sampling on the preliminary water level time series data, and set weights differentially according to the preset rare extreme hydrological event data to obtain optimized sampling data for hydrological events; oversample the rare extreme hydrological event data by the bootstrap method to obtain high-quality extreme event data; perform distribution consistency evaluation based on statistical characteristics according to the extreme event data and the optimized sampling data for hydrological events to obtain water level time series data.
[0036] In the embodiment of the present invention, for the preliminarily generated water level time series data, based on the stratified sampling weight strategy, it is divided into a normal interval and a rare interval according to the mean and variance of the data distribution, and a higher sampling weight is given to the rare interval to obtain stratified importance sampling data; taking the historical extreme hydrological event data (such as the flood event caused by a once-in-a-century rainstorm in a certain year in 2021) as a reference, adopting a differential weight generation strategy to increase the sampling probability of the extreme event data; oversampling the rare extreme hydrological event samples by 10 times by the bootstrap method to generate high-quality extreme event data; combining the stratified importance sampling data and the oversampled extreme event data, performing statistical feature consistency evaluation based on the Kolmogorov-Smirnov test, and adjusting the distribution characteristics of the inconsistent data using interpolation and smoothing methods to obtain water level time series data with consistent statistical distributions.
[0037] Step S4: Perform multi-scale verification on the water level time series data to obtain water level time series verification data, where the multi-scale verification includes spectral characteristic test and mutual information measurement; eliminate the artifacts in the water level time series verification data and perform correction processing on the extreme abnormal points based on physical constraints to obtain supervised time series water level data.
[0038] In the embodiment of the present invention, for the generated water level time series data, use wavelet transform to verify the spectral characteristics of the time series characteristics, focus on analyzing the response of the high-frequency components to the rainfall intensity and the reflection of the low-frequency components to the long-term water level changes, and calculate the spectral energy distribution difference between the real water level data and the generated data; use the mutual information measurement method to analyze the non-linear correlation between variables such as water level and rainfall, evaporation, etc., to ensure that the correlation characteristics of the generated data and the real data are consistent; after combining the verification results into the water level time series verification data, use an anomaly recognition algorithm based on a deep residual network to eliminate data artifacts, and correct the abnormal extreme values through a physical constraint method (such as a hydrological continuity equation); finally, obtain high-quality supervised time series water level data that meets the hydrological laws and statistical characteristics.
[0039] The present invention realizes the comprehensive collection of multi-channel hydrological information by acquiring multi-source hydrological monitoring data, which helps to comprehensively cover the data characteristics of different sources. On this basis, data noise and outliers are removed through outlier identification and data preprocessing steps to generate standardized basic hydrological data, ensuring the accuracy and consistency of the data. This provides a high-quality input source for subsequent data modeling and avoids the error propagation of the original data. Further, a data generation model is constructed through a conditional generative adversarial network, enabling the model to have strong adaptability and be able to simulate complex hydrological characteristics and data distribution characteristics, improving the flexibility and controllability of data generation. Through the introduction of conditional feature regulation driven by hydrological variables, the parameters of the data generation model are optimized, which can better capture the impact of hydrological conditions on hydrological changes, thereby generating conditional input feature data with hydrological adaptability. This process improves the generalization ability of the model in hydrological diversity scenarios. By injecting the conditional input feature data into the generation model, the generated preliminary water level time series data has high authenticity and relevance, laying a foundation for subsequent optimization processing. Through hierarchical importance sampling, the ability to identify key features of the preliminary water level time series data is improved. By combining rare extreme hydrological event data to set differential generation weights, the sampling strategy is optimized, enhancing the attention to extreme events. The oversampling of rare event data using the bootstrap method not only expands the sample size but also improves data balance. The distribution consistency evaluation based on statistical characteristics further ensures the coincidence degree between the generated hydrological data and the real scenario, thus obtaining more reliable water level time series data. Through multi-scale verification, including spectral property testing and mutual information measurement, the frequency domain characteristics and the dependence relationship between variables of the water level time series data are strictly tested, enhancing the scientificity and comprehensiveness of data verification. Subsequently, artifact elimination and correction of extreme outliers based on physical constraints are carried out to ensure the physical rationality and consistency of the data. These measures effectively reduce the artifact interference in the generated data and improve the quality and credibility of the finally generated supervised time series water level data.
[0040] Preferably, step S1 includes the following steps:
[0041] Step S11: Collect water level, rainfall, evaporation and flow information to obtain original hydrological monitoring data;
[0042] Step S12: Perform format conversion and timestamp calibration processing on the original hydrological monitoring data to obtain time-aligned multi-source hydrological monitoring data;
[0043] Step S13: Identify outliers in the multi-source hydrological monitoring data based on statistical methods, and perform interpolation repair processing and outlier elimination processing to obtain hydrological monitoring filtered data;
[0044] Step S14: Perform data filling processing on the hydrological monitoring filtered data using a time series regression model based on adjacent data points to obtain the hydrological monitoring cleaned data;
[0045] Step S15: Screen the feature variables highly correlated with the water level time series change for the hydrological monitoring cleaned data through the mutual information method, and generate derived variables to obtain the hydrological feature extended data;
[0046] Step S16: Perform dimensionless processing on the hydrological feature extended data, and segment the data based on different time windows to obtain the standardized hydrological basic data;
[0047] Step S17: Build a data generation model of a conditional generative adversarial network based on the standardized hydrological basic data.
[0048] As an embodiment of the present invention, refer to Figure 2 shown in Figure 1 which is the detailed step flow diagram of Step S1 in
[0049] Step S11: Collect water level, rainfall, evaporation and flow information to obtain the original hydrological monitoring data;
[0050] In the embodiment of the present invention, a certain typical river channel is taken as the research object, and 5 hydrological monitoring stations are arranged to record the water level (range: 0.5 - 5m), rainfall (0 - 300mm), evaporation (0 - 100mm) and flow (0 - 500m³ / s) respectively. The data collection period is from January 2020 to December 2023, once every hour; the data is collected through automated monitoring equipment and transmitted to the cloud data center through the Internet of Things gateway to form the original hydrological monitoring data; the storage format of the monitoring data is a CSV file, including the date, time, station number and the corresponding monitoring values. Some files have problems with inconsistent formats or missing fields.
[0051] Step S12: Perform format conversion and timestamp calibration processing on the original hydrological monitoring data to obtain the time-aligned multi-source hydrological monitoring data;
[0052] In the embodiment of the present invention, for the collected original hydrological monitoring data, first use the pandas library of Python to read the CSV file, standardize the time field to the format of "YYYY-MM-DD HH:MM:SS", and unify it to a 24-hour timestamp; for the records with missing time, automatically generate the complementary timestamp using the time difference between the previous and next two data; sort the multi-source data according to the station number, and convert the data of each station into a unified multi-dimensional table format, with each column corresponding to the water level, rainfall, evaporation and flow respectively, so as to generate the time-aligned multi-source hydrological monitoring data.
[0053] Step S13: Identify outliers in the multi-source hydrological monitoring data based on statistical methods, and perform interpolation repair processing and outlier removal processing to obtain hydrological monitoring filtered data;
[0054] In the embodiment of the present invention, for the multi-source hydrological monitoring data with time alignment, first, a statistical outlier identification method based on the box plot method is used to remove the outliers in the water level data that exceed 1.5 times the range of the upper and lower quartiles; for the interpolation repair processing of outliers, the Lagrange interpolation method is used to estimate a single outlier, and at the same time, the local average method is used to smooth the interpolation results of continuous outliers; for the detected completely incorrect data rows (such as all variables being outliers at the same time), they are directly removed; after the repair is completed, a histogram statistical analysis is performed on the distribution of the repaired data to verify the rationality of the data range, so as to obtain hydrological monitoring filtered data.
[0055] Step S14: Perform data filling processing on the hydrological monitoring filtered data based on the time series regression model of adjacent data points to obtain hydrological monitoring cleaned data;
[0056] In the embodiment of the present invention, for the hydrological monitoring filtered data, a time series regression model based on historical time point data is used to fill in the missing data; the specific method is to use water level, rainfall, evaporation, and flow as multi-variable inputs, and use the time interval of each hour as the prediction step of the regression model to train a multi-layer linear regression model to predict the missing values; for data intervals with continuous missing exceeding 12 hours, the average value based on historical similar days is used to complete the filling; after the filling is completed, the data is re-sorted and error analysis is performed to ensure that the interpolation error does not exceed 5%, so as to generate complete hydrological monitoring cleaned data.
[0057] Step S15: Screen the characteristic variables with high correlation based on the time series change of the water level in the hydrological monitoring cleaned data by the mutual information method, and generate derived variables to obtain hydrological characteristic extended data;
[0058] In the embodiment of the present invention, for the hydrological monitoring cleaned data, the mutual information method is used to analyze the correlation between the time series change of the water level and other variables (rainfall, evaporation, flow); the variables with mutual information values higher than 0.8 are selected as relevant characteristic variables; derived variables are generated according to the selected characteristic variables, such as rainfall intensity (rainfall / hours), cumulative evaporation (accumulated evaporation value), and flow change rate (differential value of flow); the derived variables are added to the original data to form an extended data set, and at the same time, the correlation of the extended variables is tested to ensure that they have a reasonable physical interpretability with the time series change of the water level, and finally, hydrological characteristic extended data is obtained.
[0059] Step S16: Perform dimensionless processing on the hydrological feature extended data and segment the data based on different time windows to obtain the standardized hydrological basic data;
[0060] In the embodiment of the present invention, for all variables in the hydrological feature extended data, dimensionless processing is performed according to the normalization formula to ensure that the value range of all variables is between [0, 1]; for time series data, it is segmented on a weekly basis, and the data is divided into multiple 7-day time windows, each window containing 168 hours of observation data; for data segments spanning different months, boundary data of the corresponding months are supplemented by means of a sliding window, so as to generate the standardized hydrological basic data covering the annual cycle.
[0061] Step S17: Construct a data generation model of a conditional generative adversarial network based on the standardized hydrological basic data.
[0062] In the embodiment of the present invention, the standardized hydrological basic data generated in step S16 is used as the input to construct a conditional generative adversarial network (CGAN); in specific implementation, the generator adopts a network composed of two layers of LSTM units and a one-dimensional convolutional layer, the input is the standardized hydrological data and conditional feature variables (such as rainfall and evaporation), and the output is the predicted water level time series data; the discriminator network extracts time series features through a convolutional layer and discriminates the distribution difference between the generated data and the real data in combination with the conditional features; the loss function of the generation network is defined as the time series distribution difference loss, the conditional constraint loss, and the physical consistency loss. After 1000 rounds of adversarial training, a conditional generative adversarial network model with stable performance is finally generated.
[0063] The present invention acquires the original hydrological monitoring data by collecting key hydrological information such as water level, rainfall, evaporation, and flow rate. This step ensures the comprehensiveness and diversity of the data, providing a reliable basis for subsequent data analysis and modeling. Covering different hydrological variables helps to comprehensively evaluate the driving factors of hydrological changes. Format conversion and timestamp calibration are performed on the original hydrological monitoring data to ensure the time alignment of multi-source data. This step guarantees the timeliness and comparability of the data, avoiding errors caused by inconsistent time, and thus providing a unified time frame for subsequent analysis, enabling effective combination of various hydrological data. Through outlier identification based on statistical methods, combined with interpolation repair and outlier removal processing, unreasonable points in the data are effectively removed. This processing greatly improves the quality of the data, eliminates outliers that may interfere with the analysis results, and ensures the accuracy and reliability of subsequent data. Based on the filtered hydrological monitoring data, the missing data is filled through a time series regression model, further optimizing the continuity and integrity of the data. The time series regression model can effectively recover the missing values by referring to the values of adjacent data points, thus ensuring the stability and usability of the data. By using the mutual information method to perform feature screening and derivative variable generation on the cleaned hydrological monitoring data, key feature variables closely related to the temporal variation of the water level can be accurately extracted. This step improves the relevance and effectiveness of the data, providing valuable input variables for subsequent model construction. Dimensionless processing is performed on the extended hydrological feature data, making the comparison between different features more scientific and reasonable. Through data segmentation based on different time windows, it is ensured that the data can exhibit different hydrological characteristics according to different time scales. This process helps with data standardization, eliminates the influence of dimensions on the analysis results, and improves the effect and accuracy of model training. Based on the standardized hydrological basic data, a data generation model of a conditional generative adversarial network is constructed, further enhancing the model's ability to simulate complex hydrological characteristics. This step can effectively generate data that conforms to actual hydrological laws by using an adversarial network, providing strong data support for subsequent hydrological prediction and analysis.
[0064] Preferably, step S17 includes the following steps:
[0065] Step S171: Construct a generator network structure of a conditional generative adversarial network based on the standardized hydrological basic data, and design a hybrid architecture including multiple layers of long short-term memory and convolutional neural networks to obtain the generator network structure data;
[0066] Step S172: Design a discriminator network structure of a conditional generative adversarial network according to the standardized hydrological basic data, and perform multi-domain consistency test processing using a multi-scale time series discrimination strategy to obtain the discriminator network structure data, where the multi-scale time series discrimination strategy includes time domain and frequency domain feature extraction modules;
[0067] Step S173: Define the loss function of the generative adversarial network according to the discriminator network structure data and the generator network structure data, so as to obtain the loss function data, where the loss function includes the loss functions of distribution similarity loss, temporal consistency loss, and hydro-physical constraint loss;
[0068] Step S174: Integrate the discriminator network structure data, the generator network structure data, and the loss function data into a data generation model.
[0069] In the embodiments of the present invention, based on standardized hydrological basic data, a hybrid architecture of a generator network is designed. The generator network consists of two parts: a long short-term memory (LSTM) network module and a convolutional neural network (CNN) module. First, a two-layer stacked LSTM network is used to extract the time series features of hydrological data. The number of hidden units of the LSTM is set to 128, and the input data is the hourly data segment (with a length of 168 hours) in the hydrological feature extended data. Then, the high-dimensional features output by the LSTM are processed through a one-dimensional convolutional layer with a kernel size of 3, a stride of 1, and an output channel number of 64. Finally, through two fully connected layers, prediction data with the same length as the target water level time series data is generated. To improve the stability of the generator network, batch normalization and the ReLU activation function are adopted, and the sigmoid activation function is used in the last layer to map the output to [0,1]. This structure can capture the long-term dependence and local features of hydrological data in the generator network, thereby obtaining the generator network structure data. When designing the discriminator network structure, a multi-scale time series discrimination strategy is introduced to enhance the detection ability of data characteristics. The discriminator network contains two modules: a time domain feature extraction module and a frequency domain feature extraction module. The time domain module consists of three layers of one-dimensional convolutional neural networks with kernel sizes of 3, 5, and 7 respectively, and strides of 2 for all, to extract time series features of different scales. The frequency domain module converts the time series into frequency spectrum features through the fast Fourier transform (FFT), and then extracts the frequency domain features through a layer of convolutional network. The outputs of the two modules are fused through a fully connected layer, and finally, a discrimination result is output through the sigmoid activation function, indicating whether the input data is real or generated. The input of the discriminator is a combination of the water level time series data output by the generator and the real hydrological data, and the consistency test between the time domain and the frequency domain is realized through multi-scale feature extraction, thereby obtaining the discriminator network structure data. Based on the network structures of the generator and the discriminator, the loss function of the generative adversarial network is defined. The loss function consists of three parts: distribution similarity loss, temporal consistency loss, and hydrological physical constraint loss. The distribution similarity loss calculates the difference between the generated data and the real data distribution through cross-entropy. The temporal consistency loss uses the dynamic time warping (DTW) method to measure the alignment degree of the generated data and the real data on the time axis. The hydrological physical constraint loss constrains the physical meaning of the generated data (such as the upper and lower limits of the water level, the fluctuation trend, and the correlation with rainfall), and calculates the physical consistency of the generated data using a penalty term. The three parts of the loss are weighted and summed, with weights of 0.5, 0.3, and 0.2 respectively, and the generative adversarial network is optimized by combining the Adam optimization algorithm, thereby obtaining the loss function data. The discriminator network structure data, the generator network structure data, and the loss function data are integrated into the conditional generative adversarial network model.In the specific operation, first, the parameters of the generator and discriminator networks are initialized. The weights of the generator are initialized using Xavier initialization, and the weights of the discriminator are randomly initialized using a normal distribution. Then, adversarial training is performed on the generator and discriminator through the definition of the loss function. In each round of training, the discriminator receives real data and generated data to distinguish between true and false, and the generator adjusts its parameters to deceive the discriminator to the greatest extent. During the training process, a learning rate adjustment strategy is adopted. The initial learning rate is set to 0.001 and decreased by 10% every 100 rounds. After 2000 rounds of training, the distribution error of the generative adversarial network model on the validation data is less than 5%, and the temporal consistency error is less than 3%, thus obtaining a hydrological data generation model with stable performance.
[0070] The present invention designs a hybrid architecture incorporating multiple layers of long short-term memory (LSTM) and convolutional neural network (CNN) by constructing the generator network structure of a conditional generative adversarial network (GAN) based on standardized hydrological basic data. This architecture combines the time series modeling ability of LSTM and the feature extraction advantages of CNN, enabling it to effectively capture the temporal and spatial features of hydrological data, thereby generating more accurate and realistic hydrological time series data. The design of this generator network structure provides a strong foundation for the efficient data generation of the conditional generative adversarial network, ensuring the effectiveness and generalization ability of the network in complex hydrological change environments. By designing the discriminator network structure of the conditional generative adversarial network and introducing a multi-scale time series discrimination strategy for multi-domain consistency verification, the comprehensive analysis of the data's temporal and frequency domain features by the model is enhanced. This discrimination strategy includes temporal and frequency domain feature extraction modules, enabling the discriminator to not only identify temporal consistency but also evaluate the authenticity of the generated data from the frequency level. The addition of this strategy allows the network to comprehensively consider the performance of the data at different levels, thereby improving the accuracy and robustness of discrimination and further enhancing the data quality of the generator network. According to the generator and discriminator network structures, the loss function of the generative adversarial network is defined, including distribution similarity loss, temporal consistency loss, and hydrological physical constraint loss. Through the design of these loss functions, the network can optimize the generated data during the training process, not only ensuring the distribution similarity between the generated data and the real data but also ensuring the consistency of the data in the time series, and strengthening the practical applicability of the generated data by combining hydrological physical constraints. This multi-dimensional loss function design ensures that the generated data meets the statistical characteristics while conforming to the physical laws and temporal characteristics of the hydrological system, enhancing the comprehensive effect of the model. Integrating the discriminator network structure data, the generator network structure data, and the loss function data into a complete data generation model completes the construction of the conditional generative adversarial network. The completion of this step means that the model already has the ability to generate hydrological data from input conditions. By optimizing the collaboration between the generator and the discriminator, it can effectively simulate and generate time series data that conform to actual hydrological changes. This integrated generation model provides strong data support for subsequent hydrological prediction and decision-making and has high application value and practical operability.
[0071] Preferably, step S2 includes the following steps:
[0072] Step S21: Obtain historical hydrological variable data, including temperature, humidity, precipitation, and atmospheric pressure information; perform feature dimensionality reduction processing on the historical hydrological variable data and extract key hydrological features to obtain key hydrological feature data;
[0073] Step S22: Design a hydrological variable encoding network based on the attention mechanism using the key hydrological feature data, and perform a non-linear transformation on the key hydrological feature data using the hydrological variable encoding network to obtain conditional input feature data;
[0074] Step S23: Perform hydrological adaptability parameter tuning on the generator network of the data generation model according to the conditional input feature data to obtain hydrological adaptability optimized model parameters, where the hydrological adaptability parameter tuning is specifically to design a variational inference-based generator network parameter self-adaptive optimization strategy;
[0075] Step S24: Optimize the generator network of the data generation model according to the hydrological adaptability optimized model parameters to obtain a data generation optimized model;
[0076] Step S25: Inject the conditional input feature data into the generator network of the data generation optimized model, and generate preliminary water level time series data through the adversarial training of the generator and discriminator of the generative adversarial network.
[0077] As an embodiment of the present invention, refer to Figure 3 shown, for Figure 1 the detailed step flow diagram of step S2 in
[0078] Step S21: Obtain historical hydrological variable data, including temperature, humidity, precipitation, and atmospheric pressure information; perform feature dimensionality reduction processing on the historical hydrological variable data, and perform key hydrological feature extraction to obtain key hydrological feature data;
[0079] In the embodiment of the present invention, hydrological variable data for the most recent 30 years is extracted from the hydrological observation database, including daily records of temperature, humidity, precipitation, and atmospheric pressure information, with a data resolution of daily average values. First, principal component analysis (PCA) is used to perform dimensionality reduction processing on the hydrological variable data, reducing the original 4-dimensional features to 2 dimensions, where the cumulative variance contribution rate of the principal components reaches 95%, reducing redundant information. Then, the mutual information method is used to evaluate the correlation between the dimensionality-reduced features and the target water level change, and key hydrological features are extracted, including the annual average temperature change range and the seasonal variation value of precipitation. These key features are particularly important in drought and flood risk analysis, thus obtaining key hydrological feature data.
[0080] Step S22: Design a hydrological variable encoding network based on the attention mechanism using the key hydrological feature data, and perform a non-linear transformation on the key hydrological feature data using the hydrological variable encoding network to obtain conditional input feature data;
[0081] Based on the key hydrological feature data, an embodiment of the present invention designs a hydrological variable encoding network based on the attention mechanism. The encoding network adopts a two-layer Transformer structure, inputs the key hydrological feature data, and captures the feature dependence relationships at different time steps through the self-attention mechanism. In a specific implementation, the number of heads of the multi-head attention unit in the first Transformer encoding layer is set to 8, and the hidden layer dimension is 64; the number of heads of the attention unit in the second Transformer layer is set to 4, and the hidden layer dimension is 32. Residual connections and layer normalization operations are added to enhance the training stability. Finally, the network embeds the non-linearly transformed hydrological features into a 32-dimensional conditional input feature vector space through a fully connected layer and a ReLU activation function, thereby obtaining the conditional input feature data.
[0082] Step S23: Perform hydrological adaptability parameter tuning on the generator network of the data generation model according to the conditional input feature data, so as to obtain the hydrological adaptability optimized model parameters, where the hydrological adaptability parameter tuning is specifically to design an adaptive optimization strategy for the generator network parameters based on variational inference;
[0083] An embodiment of the present invention uses the conditional input feature data generated in step S22 to perform hydrological adaptability parameter tuning on the generator network. Design an adaptive optimization strategy for the generator network based on variational inference. First, define the conditional input feature data as the conditional input of the generator network, convert the conditional vector distribution into an approximate normal distribution through a variational autoencoder (VAE), and define the KL divergence loss between the distributions. Secondly, during the optimization process, dynamically adjust the initialization of the hidden layer weights of the generator, use the Adam optimizer with a learning rate of 0.001, set the training batch to 64, and perform 5000 iterations. Through this method, the generator network can more effectively learn the influence of hydrological variables on hydrological changes, thereby obtaining the hydrological adaptability optimized model parameters.
[0084] Step S24: Optimize the generator network of the data generation model according to the hydrological adaptability optimized model parameters, so as to obtain the data generation optimized model;
[0085] An embodiment of the present invention further optimizes the generator network according to the hydrological adaptability optimized model parameters. In the optimization, first freeze the discriminator parameters and train the generator alone to make the generated water level time series data more conform to the real hydrological change characteristics. Adopt a cyclic learning rate strategy, set the initial learning rate to 0.0005, adjust it to 0.0001 after 2000 rounds of training, and add L2 regularization to prevent overfitting. During the optimization process, monitor the distribution difference and hydrological correlation of the generated data through the validation set until the distribution error between the generated water level data and the historical real data is less than 3%, and finally obtain a performance-optimized data generation model.
[0086] Step S25: Inject the conditional input feature data into the generator network of the data generation optimization model, and generate preliminary water level time series data through the adversarial training of the generator and discriminator of the generative adversarial network.
[0087] In the embodiment of the present invention, the conditional input feature data in step S22 is injected into the optimized generator network for adversarial training of the generative adversarial network. In specific training, the generator generates preliminary water level time series data according to the conditional input features, and inputs it into the discriminator together with the real water level time series data for true / false discrimination. The optimization goal of the generator is to minimize the discrimination accuracy of the discriminator, while the optimization goal of the discriminator is to improve the discrimination accuracy. In training, the mini-batch stochastic gradient descent method (batch size = 32) is adopted. Each time the generator is updated, it is trained for 1 round, and the discriminator is trained for 2 rounds. After continuous training for 3000 rounds, the preliminary water level time series data generated by the generator has high consistency with the real data in terms of time distribution, and finally the preliminary water level time series data is obtained.
[0088] The present invention effectively extracts key hydrological feature data such as temperature, humidity, precipitation, and atmospheric pressure by obtaining historical hydrological variable data and performing feature dimensionality reduction processing. This process can reduce the interference of redundant information while retaining the hydrological information crucial for generating water level time series data. Through dimensionality reduction and feature extraction, the data becomes more concise and efficient, thus improving the speed and accuracy of subsequent model training. The key hydrological feature data provides a necessary information basis for hydrological adaptability optimization, ensuring that the subsequent steps can perform more effective modeling and optimization based on the refined hydrological features. Using the design of a hydrological variable encoding network based on the attention mechanism, a non-linear transformation is performed on the key hydrological feature data to obtain conditional input feature data. The attention mechanism helps the model focus on the hydrological factors that have a greater impact on water level prediction by weighting the influences of different hydrological variables, thereby improving the sensitivity and accuracy of the model to hydrological changes. The non-linear transformation further enhances the network's ability to model complex hydrological patterns, ensuring that the generated conditional input feature data can comprehensively reflect the impact of hydrological changes on the water level time series and improving the hydrological adaptability of the data generation model. By using the conditional input feature data to optimize the hydrological adaptability parameters of the generator network of the data generation model, an adaptive optimization strategy for generator network parameters based on variational inference is designed. This optimization process enables the generator to better adapt to the data generation requirements under different hydrological conditions. By optimizing the parameters using the variational inference method, the generalization ability of the network is improved. Hydrological adaptability optimization ensures that the generator can generate water level time series data that conforms to the actual conditions in various hydrological change scenarios, thereby improving the hydrological adaptability and reliability of the generation model. After obtaining the hydrological adaptability optimized model parameters, the generator network of the data generation model is optimized. This optimization further improves the performance of the generator network, enabling it to generate more accurate water level time series data that highly matches the actual hydrological conditions. Model optimization ensures that the generator can stably and efficiently generate water level data that meets expectations when facing various hydrological scenarios, thus providing reliable data support for water level prediction and management. By injecting the conditional input feature data into the optimized generator network and through the adversarial training of the generator and discriminator of the generative adversarial network, preliminary water level time series data is generated. This adversarial training process ensures that the generated data not only conforms to the changing trend of hydrological features but also can be verified for its authenticity and consistency by the discriminator. Through this interaction between the generator and the discriminator, the water level time series data is further optimized and verified, providing a high-quality data basis for subsequent water level prediction and hydrological change analysis.
[0089] Preferably, step S25 includes the following steps:
[0090] Step S251: Input the conditional input feature data as a conditional vector into the initial layer of the generator network in the data generation optimization model, thereby constructing the conditional-guided generator network input data;
[0091] Step S252: Perform multi-layer non-linear transformation on the input data of the generator network through the generator network in the data generation optimization model, so as to generate candidate water level time series data;
[0092] Step S253: Input the candidate water level time series data into the discriminator network in the data generation optimization model, and perform authenticity discrimination and adversarial training according to the preset hydrological basic data, so as to obtain adversarial training feedback data;
[0093] Step S254: Perform parameter adaptive adjustment on the generator network in the data generation optimization model according to the adversarial training feedback data of the discriminator network, and optimize the quality of the candidate water level time series data, so as to obtain generator iterative optimization data;
[0094] Step S255: Loop through steps S252 to S254 until the generated generator iterative optimization data meets the preset adversarial training convergence criterion, so as to obtain preliminary water level time series data.
[0095] In the embodiment of the present invention, the conditional input feature data generated in step S22 is used as a conditional vector to be input into the initial layer of the generator network, thereby constructing the input data of the conditional-guided generator network. In a specific implementation, the conditional vector is a 32-dimensional feature embedding vector, and a linear mapping layer is used to expand it to the input dimension of the generator network, which is 256. After the linear mapping layer, there is a Batch Normalization layer and a ReLU activation function, which are used to ensure the stable distribution of the input data. Then, it is concatenated with a random noise vector (a normal distribution with a mean of 0 and a standard deviation of 1, and a dimension of 128) to form the input data tensor of the generator network. The input data tensor generated in step S251 is input into the generator network for multi-layer non-linear transformation to generate candidate water level time series data. The generator network adopts a hybrid architecture that combines a long short-term memory network (LSTM) and a fully connected layer. First, the long-term dependence features in the time series are extracted through 3 layers of LSTM units, and the number of hidden units in each layer of LSTM is 128, 64, and 32 respectively. Then, through two fully connected layers, non-linear mapping and compression of the data dimension are completed, and the water level time series data is output, with the time series length set to 365 days (daily average water level). During the network training process, the Adam optimizer is used, and the learning rate is set to 0.0001. The initial quality of the candidate water level time series data is dynamically adjusted based on the generator parameters. The candidate water level time series data generated in step S252 is input into the discriminator network to compare its authenticity with the preset historical hydrological basic data. The discriminator network adopts a convolutional neural network (CNN) architecture, which includes 3 one-dimensional convolutional layers and 2 fully connected layers. The sizes of the convolutional kernels are 5, 3, and 3 respectively, and the number of channels is 64, 128, and 256 in sequence. Finally, the network outputs the true / false discrimination probability through the Softmax activation function, and at the same time, combines the frequency domain transformation module to additionally examine the spectral distribution characteristics of the candidate data. The loss function for authenticity discrimination is binary cross-entropy, and the adversarial training feedback data is generated by combining the frequency domain consistency discrimination result, which is used to guide the improvement of the generator network. According to the adversarial training feedback data obtained in step S253, the parameters of the generator network are adaptively adjusted. The specific adjustment method is to perform gradient backpropagation based on the adversarial loss (including the generator loss and the discriminator loss) and the time distribution consistency loss to optimize the weights of the generator network. At the same time, post-processing is performed on the output of the candidate water level time series data. High-frequency artifacts are eliminated through a low-pass filter, and outliers are corrected by linear interpolation. After each adjustment, the distribution difference between the data output by the generator and the real data gradually decreases, and the optimized data is called the generator iterative optimization data. Steps S252 to S254 are repeatedly executed to perform multiple rounds of adversarial training and candidate data optimization on the generator network until the preset adversarial training convergence criterion is met.The specific convergence criteria are set as follows: 1) The distribution difference between the output data of the generator and the real hydrological data is less than 2%; 2) The change range of the adversarial loss of the generator is less than 0.01 in 100 consecutive iterations; 3) The statistical characteristics of the generated water level time series data in the time domain and frequency domain are consistent with the historical data. Finally, when the above criteria are met, the iteratively optimized data output by the generator is the high-quality preliminary water level time series data.
[0096] The present invention inputs the conditional input feature data as a conditional vector into the initial layer of the generator network of the data generation optimization model, thereby constructing the input data of the conditional-guided generator network. This process ensures that the generator network can guide the generation of water level time series data based on hydrological features and other relevant conditions. The conditional input feature data provides important information about the external environmental conditions for the generator, enabling the generated water level time series data to accurately reflect the changing trends of hydrological variables, and increasing the accuracy and pertinence of data generation. The input data is subjected to multi-layer non-linear transformation through the generator network to generate candidate water level time series data. The multi-layer non-linear transformation can capture the complex relationships and patterns in the input data, thereby generating more accurate and rich candidate water level time series data. This transformation process makes the generated data not only conform to hydrological conditions but also have a high degree of detail and diversity, which helps to better simulate the dynamic changes of water level time series and improve the quality and usability of the generated data. The candidate water level time series data is input into the discriminator network, and authenticity discrimination and adversarial training are carried out based on the preset hydrological basic data to obtain adversarial training feedback data. Through the evaluation of the discriminator and adversarial training, the generated data has passed the authenticity verification and is ensured to conform to the actual hydrological conditions. The adversarial process between the discriminator and the generator prompts the generator to continuously improve, and the generated data gradually improves in quality, reducing the output of untrue or inappropriate data, thereby enhancing the robustness and credibility of the model. The parameters of the generator network are adaptively adjusted according to the adversarial training feedback data of the discriminator network, and the quality of the candidate water level time series data is optimized to obtain the generator iterative optimization data. This feedback mechanism allows the generator network to continuously adjust its parameters according to the opinions of the discriminator and optimize the data generation process. Through repeated iterative optimization, the generator can generate water level time series data that is more in line with the actual situation and has higher quality. This process helps to improve the accuracy and consistency of the generated data and finally obtain a more satisfactory output. By repeatedly executing step S252 to step S254 until the generated generator iterative optimization data meets the preset adversarial training convergence criterion, preliminary water level time series data is obtained. Through repeated iteration and optimization, the generator continuously improves its output, and the finally generated water level time series data reaches the set convergence standard and has a high degree of accuracy and reliability. This process ensures that the generated water level time series data can provide effective support in practical applications, accurately reflect the changing trends of hydrological conditions, and finally meet the requirements and expectations of the model.
[0097] Preferably, step S3 includes the following steps:
[0098] Step S31: Design a stratified sampling weight strategy according to the distribution characteristics and statistical attributes of the preliminary water level time series data to obtain sampling weight strategy data;
[0099] Step S32: Perform stratified importance sampling on the preliminary water level time series data according to the sampling weight strategy data to obtain importance-stratified sampling data;
[0100] Step S33: Based on the preset characteristics of rare extreme hydrological events, set different weights for the sampling probabilities of prominent extreme hydrological events to obtain differentially generated weight data;
[0101] Step S34: Perform weighted sampling of different types of hydrological events on the extreme hydrological event data according to the differentially generated weight data to obtain optimized sampling data for hydrological events;
[0102] Step S35: Oversample the rare extreme hydrological event data by the bootstrap method to obtain high-quality extreme event data;
[0103] Step S36: Conduct a distribution consistency assessment based on statistical characteristics according to the importance-stratified sampling data, extreme event data, and optimized sampling data for hydrological events to obtain the water level time series data.
[0104] In the embodiments of the present invention, the distribution characteristics and statistical attributes of the preliminary water level time series data are analyzed. First, by calculating statistical indicators such as mean, standard deviation, skewness, and kurtosis, the overall distribution pattern of the data is described. Then, combined with the time series attributes of the data, the change rate and frequency characteristics in different time windows are calculated, with particular attention paid to the regions of sudden increase and decrease in water level. According to these characteristics, the preliminary water level time series data is divided into three layers: low water level interval, medium water level interval, and high water level interval. The sample size weights of each layer are initially allocated according to their proportions in the overall data, and at the same time, the sampling weight of the high water level interval is increased to capture potential extreme events. The designed stratified sampling weight strategy is output in the form of a weighted coefficient matrix, where each row corresponds to the weight of an interval. According to the sampling weight strategy data in step S31, stratified importance sampling is implemented on the preliminary water level time series data. The specific method is to resample the sample distribution of each layer interval according to the weight matrix. During the resampling process, a sampling with replacement strategy is adopted, and different sampling probabilities are set according to the importance weights of each layer. For example, the weight of the low water level interval is set to 0.3, the medium water level interval is 0.5, and the high water level interval is 0.7. This sampling process is implemented through Python, and the random.choice method of the Numpy library is called to generate stratified samples. Finally, the importance stratified sampling data is output. Its sample size is the same as the original data, but the distribution pays more attention to the characteristics of the high water level interval. Combining the preset characteristics of rare extreme hydrological events, such as flood peak water level exceeding the 90th percentile or high water level events with a duration exceeding 7 days, the sampling probability is adjusted differentially. By setting the threshold range of the extreme event characteristics, such as water level ≥ 5 meters and duration ≥ 3 days, the generation weight of the event samples that meet the conditions is increased to 1.5 times the original weight, and at the same time, the adjacent time windows are extended to increase the sampling probability of their associated samples. The differential generation weight is achieved through a weighting function, and the form of the weight function is: , where is the adjustment factor, is the distance attenuation coefficient, Let \(t\) be the time distance between the sample and the extreme event, and the generated differential weights are output in vector form. Apply the differential weight data generated in step S33 to the sampling process of extreme hydrological event data. By setting the weighting coefficients for different types of hydrological events (such as short-term floods, long-term waterlogging, etc.), the sampling probabilities of different types of events are optimized separately. For example, the weight of the short-term flood event is set to 1.2, and the weight of the long-term waterlogging event is set to 1.5. And through the random weighted sampling method, weighted sampling is performed on extreme events to ensure that the distribution ratio of each type of event in the optimized sampling data meets the set target. The sampling process is completed based on the resample function of Sklearn, and finally, optimized sampling data of hydrological events covering multiple types of events and focusing on extreme events are obtained. Oversample the rare extreme hydrological event data by the bootstrap method. First, repeatedly sample the extreme event sample set. The specific operation of the bootstrap method is: randomly draw samples from the original data set to form a new data set, and the samples are put back after each sampling to ensure that each sample may be selected multiple times. Set the oversampling ratio to 2 times, that is, the final number of samples is twice that of the original rare event data set. Use the random.choices method in Python for sampling, and remove duplicates and sort the data in time for the oversampled data to ensure the temporal coherence of the expanded data set, and finally obtain high-quality extreme event data. After merging the importance stratified sampling data in step S32, the optimized sampling data of hydrological events in step S34, and the extreme event data in step S35, perform a distribution consistency assessment based on statistical characteristics. First, perform a distribution fitting analysis on the merged data, and calculate the overlap rate between its frequency distribution histogram and the original water level time series data; then evaluate the distribution consistency of the two through the Kolmogorov-Smirnov (K-S) test, and independently count the distribution proportion of rare events to ensure that the proportion of rare events in the optimized data increases by at least 15%. During the evaluation process, use Matplotlib to draw a distribution comparison graph, and output the final water level time series data that meets the consistency requirements for subsequent processing.
[0105] The present invention designs a stratified sampling weight strategy based on the distribution characteristics and statistical attributes of preliminary water level time series data, thereby obtaining sampling weight strategy data. This strategy design ensures that the sampling weights of different water level time series data can be optimized according to the distribution characteristics of the data, especially for possible biases or imbalances in the data. This strategy provides a reasonable framework for subsequent sampling processing, enabling a better representation of the overall picture of water level time series data in the subsequent process and assigning appropriate weights to different data points, thereby improving the representativeness and reliability of the data. Stratified importance sampling is performed on the preliminary water level time series data according to the sampling weight strategy data, thereby obtaining importance stratified sampling data. Through stratified importance sampling, it can be ensured that the relative importance of different data points is accurately reflected during the processing, especially in water level time series data, where data in certain time periods or conditions is more critical. This step reduces the risk of over-representing or under-representing data in certain time periods by optimizing the sampling process, thereby enhancing the balance and accuracy of the data. Based on the preset characteristics of rare extreme hydrological events, a differential generation weight setting for the sampling probability of highlighting extreme hydrological events is performed, thereby obtaining differential generation weight data. This differential weight setting focuses on extreme hydrological events, ensuring that these rare but important events can be appropriately emphasized during the generation process. In this way, the prediction ability of the model when dealing with extreme events can be effectively enhanced, ensuring that the generation and reflection of extreme events have sufficient accuracy and reliability. Weighted sampling of different types of hydrological events is performed on the extreme hydrological event data according to the differential generation weight data, thereby obtaining optimized sampling data for hydrological events. This process further optimizes the sampling of hydrological events, reasonably enhancing the occurrence frequency and importance of extreme hydrological events during the generation process. Through weighted sampling, it can be ensured that extreme events are appropriately represented in the generated data, making the generated data more in line with the actual occurrence law of hydrological events, thereby enhancing the generalization ability of the model. Oversampling of rare extreme hydrological event data is performed by the bootstrap method, thereby obtaining high-quality extreme event data. The bootstrap method is a commonly used oversampling method that can effectively balance the sample size of rare events in the data, ensuring that these events do not affect the training effect of the model due to insufficient data volume. Through this method, more reliable extreme hydrological event data can be obtained, thereby enhancing the accuracy and stability of the generation model when dealing with extreme situations. Based on the importance stratified sampling data, extreme event data, and optimized sampling data for hydrological events, a distribution consistency assessment based on statistical characteristics is performed to obtain water level time series data. Through the distribution consistency assessment based on statistical characteristics, it can be ensured that the finally generated water level time series data not only conforms to the distribution law of actual data in terms of statistical characteristics but also accurately reflects the processing effects of importance sampling and extreme events.This step helps to verify the quality of the generated data and ensure its high consistency and usability, thus providing a solid foundation for subsequent model applications.
[0106] Preferably, step S36 includes the following steps:
[0107] Step S361: Extract statistical features from importance-stratified sampling data, extreme event data, and optimized sampling data for hydrological events, and construct a multi-dimensional statistical feature vector;
[0108] Step S362: Design a distribution distance calculation method based on statistical features according to the multi-dimensional statistical feature vector, and conduct a distribution similarity test based on Kolmogorov-Smirnov to obtain distribution consistency measurement data;
[0109] Step S363: Perform distribution feature reconciliation processing on different data sets of importance-stratified sampling data, extreme event data, and optimized sampling data for hydrological events based on feature interpolation and smoothing techniques according to the distribution consistency measurement data, so as to obtain aligned water level time series data.
[0110] In the embodiments of the present invention, statistical feature extraction is first performed on importance stratified sampling data, extreme event data, and hydrological event optimized sampling data respectively. The extracted features include mean, variance, kurtosis, skewness, time series autocorrelation, and extreme value ratio, etc. The specific operation is to calculate the features for each data set one by one. For example, the extreme value ratio is obtained by statistically counting the sample ratio of water levels exceeding the 95th percentile; the time series autocorrelation is obtained by calculating the lag correlation coefficient of the autocorrelation function. The extracted statistical features in each dimension are stored in the form of feature vectors, and the data of each sample is represented as a multi-dimensional vector, which contains the above statistical feature values. To ensure the consistency of features of different data sets, a standardization method is used to perform dimensionless processing on the feature vectors, and finally a multi-dimensional statistical feature vector set composed of three groups of feature vectors is generated. Based on the multi-dimensional statistical feature vectors generated in step S361, a distribution distance calculation method is designed. The method is specifically to measure the similarity of the distributions of different data sets through the Kolmogorov-Smirnov (K-S) test. The implementation steps of the K-S test are as follows: First, construct a cumulative distribution function (CDF) for the feature distribution of each group of data. For example, construct a CDF curve of the extreme value ratio feature within the range of [0,1]; then calculate the maximum vertical distance between the corresponding CDF curves of each pair of data sets as the K-S statistic. According to a preset similarity threshold, such as 0.05, determine whether the distributions are significantly different, and at the same time record the distribution distances of each group of data, and store the distribution consistency measurement data in the form of a matrix, where the rows and columns of the matrix represent different data sets respectively, and the element value is the K-S statistic. According to the distribution consistency measurement data obtained in step S362, distribution feature reconciliation processing is performed on importance stratified sampling data, extreme event data, and hydrological event optimized sampling data. The specific method is: for data sets with significant distribution differences, a feature correction technique based on linear interpolation is used to make the data features of the data set with fewer samples approach the data features of the data set with more samples. For example, for the kurtosis feature of extreme event data, use the interpolation function , where and are the feature values of two data sets respectively, is the interpolation weight (the value range is 0.2 to 0.8). At the same time, to eliminate local fluctuations, cubic spline interpolation is used to smooth the time series features. By adjusting the feature distributions of each data set through the above method, making its K-S statistic reach within the preset similarity threshold (such as 0.03), finally generate water level time series data with consistent and aligned features.
[0111] The present invention extracts statistical features from importance stratified sampling data, extreme event data, and optimized sampling data for hydrological events, and constructs a multi-dimensional statistical feature vector. The main purpose of this process is to deeply analyze various types of data, extract their key statistical characteristics, and transform these characteristics into multi-dimensional feature vectors. These feature vectors can comprehensively reflect the internal laws of the data, including the central tendency, dispersion degree, and distribution characteristics of the data, providing necessary statistical information support for subsequent analysis and processing, and ensuring that the multi-dimensional characteristics of the data are fully captured. A distribution distance calculation method based on statistical features is designed according to the multi-dimensional statistical feature vector, and a distribution similarity test based on Kolmogorov-Smirnov is performed to obtain distribution consistency measurement data. This step can effectively measure the distribution consistency between different data sets by constructing an appropriate distribution distance calculation method and combining the Kolmogorov-Smirnov test for distribution similarity evaluation. The Kolmogorov-Smirnov test is a classical statistical method that can objectively evaluate the distribution differences of data sets and provide a quantitative basis for further data adjustment. This step helps to ensure the consistency of all sampling data in statistical features and lays a solid foundation for subsequent data fusion and modeling. According to the distribution consistency measurement data, the distribution characteristics of different data sets of importance stratified sampling data, extreme event data, and optimized sampling data for hydrological events are harmonized based on feature interpolation and smoothing techniques to obtain aligned water level time series data. This step solves the differences in distribution characteristics among different data sets through feature interpolation and smoothing processing techniques, enabling these data to be fused and aligned under the same framework. By harmonizing the distribution characteristics of different data sets, potential biases and noises can be eliminated, ensuring that the statistical attributes and distribution characteristics of the water level time series data are consistent. This process greatly improves the quality and stability of the data, enabling the generated water level time series data to better reflect the actual hydrological conditions, thereby improving the prediction ability and accuracy of the model.
[0112] Preferably, step S4 includes the following steps:
[0113] Step S41: Verify the spectral characteristics of the water level time series data to generate spectral characteristic verification data, where the spectral characteristic verification specifically uses wavelet transform to decompose multi-scale time series features, analyze the energy distribution in different time frequency bands, and evaluate the spectral characteristic consistency between the water level time series data and the pre-acquired real water level data;
[0114] Step S42: Analyze the correlation and information transfer characteristics between different variables of the water level time series data based on the mutual information measurement method to obtain mutual information enhanced verification data;
[0115] Step S43: Combine the spectrum characteristic verification data and the mutual information enhancement verification data into water level time series verification data;
[0116] Step S44: Eliminate the artifacts in the water level time series verification data based on the anomaly recognition algorithm, so as to obtain the optimized data with artifacts removed;
[0117] Step S45: Perform correction processing on the optimized data with artifacts removed for the extreme anomaly points based on physical constraints, so as to obtain the supervised time series water level data.
[0118] The embodiments of the present invention verify the spectral characteristics of water level time series data. The specific method is to perform multi-scale decomposition on the water level time series data using wavelet transform. By selecting an appropriate wavelet basis (such as the Morlet wavelet), the water level time series data is decomposed into multiple frequency sub-bands. Then, calculate the energy distribution in each frequency sub-band and analyze the spectral characteristics of different time frequency bands. For each sub-band, calculate the ratio of its energy to the total energy, and analyze the consistency between the energy distribution of the water level time series data in each frequency band and the pre-acquired real water level data through comparison. The spectral characteristic consistency evaluation uses the mean square error (MSE) or the Pearson Correlation Coefficient (PCC) to quantify the spectral similarity of the two data sets, and then generates spectral characteristic verification data. Based on the mutual information measurement method, analyze the correlation and information transfer characteristics between different variables of the water level time series data. First, evaluate the information transfer relationship between different variables by calculating the mutual information between each variable in the water level time series data. For example, use the Shannon mutual information formula to calculate the mutual information between each pair of variables. By analyzing the mutual information between multiple variables such as water level data, rainfall, and flow, the dependence relationship between different hydrological characteristics can be revealed. Sort the mutual information of each pair of variables, and select the variable pairs with high mutual information as the core variables of the verification data. Finally, generate mutual information enhanced verification data, and further analyze the potential correlation patterns and characteristics in the water level time series data through this data. Combine the spectral characteristic verification data and the mutual information enhanced verification data into water level time series verification data. First, perform standardization processing on these two sets of data to ensure that their dimensions are consistent, and use the weighted average method for fusion. For example, a weight of 0.6 can be assigned to the spectral characteristic verification data and a weight of 0.4 can be assigned to the mutual information enhanced verification data, and then the two are weighted and summed to obtain a comprehensive verification data set. The combined water level time series verification data contains information on spectral consistency and inter-variable correlation, and can provide a comprehensive verification basis for subsequent artifact removal and outlier correction. Use an anomaly recognition algorithm to eliminate artifacts from the water level time series verification data. First, adopt algorithms such as the Isolation Forest or Local Outlier Factor (LOF) to identify outliers in the water level time series verification data. These outliers may be artifacts caused by measurement errors, equipment failures, or other abnormal factors. The algorithm automatically identifies outliers that deviate from the normal trend by detecting the local density deviation of the data. For these outliers, repair them through interpolation or smoothing methods. For example, use a sample-based interpolation method (such as spline interpolation) to fill in the outliers, or use the weighted average method to smooth the fluctuations in the data, so as to obtain artifact removal optimized data. Perform extreme outlier correction processing based on physical constraints on the artifact removal optimized data.First, by defining physical constraints, such as the maximum and minimum water levels, reasonable ranges of flow and precipitation, etc., the data is screened in combination with the actual hydrological conditions. For the extreme points that exceed these physical constraints, they are adjusted by a processing method based on correction rules. For example, using the maximum-minimum constraint method or other correction algorithms, the extreme value data is adjusted to a reasonable range. In addition, for some mutant extreme abnormal points, data smoothing techniques (such as weighted moving average method) can be used for further correction to ensure the smoothness and physical rationality of the water level time series data, and finally the supervised time series water level data is obtained.
[0119] The present invention verifies the spectral characteristics of water level time series data, decomposes the multi-scale time series features by using wavelet transform, analyzes the energy distribution in different time-frequency segments, and conducts an evaluation on the consistency of the spectral characteristics between the water level time series data and the pre-acquired real water level data. Through spectral analysis, the variation characteristics of the water level time series data can be deeply understood from different frequency scales, revealing potential periodic patterns or abnormal fluctuations. Wavelet transform can effectively capture the local features of the water level time series data, enabling spectral analysis to not only identify the overall trend but also handle local anomalies, thereby improving the accuracy and reliability of the water level data. The comparison of the spectral characteristics consistency evaluation with the real data can ensure that the generated data conforms to the physical laws of the actual water level changes. Analyze the correlation and information transfer characteristics between different variables of the water level time series data based on the mutual information measurement method to obtain mutually information enhanced verification data. Through the mutual information measurement analysis, the non-linear relationships and mutual influences between variables can be revealed, thereby providing a more comprehensive correlation analysis for the water level time series data. The advantage of this analysis method is that it can capture the complex non-linear and dependence relationships between variables, being more flexible and accurate than traditional linear methods. Through the enhanced verification data, the accuracy of the model can be effectively improved, especially when dealing with complex multi-variable relationships, which helps to generate time series data that is more in line with the actual situation. Combine the spectral characteristics verification data and the mutually information enhanced verification data into water level time series verification data. By integrating these two types of verification data, the analysis results at the frequency domain and information transfer levels can be combined, thereby comprehensively improving the verification validity of the data. The spectral characteristics verification and the mutually information enhancement verification provide verification information from different perspectives, and the combination of the two makes the final verification of the water level time series data more comprehensive and reliable, effectively eliminating the blind spots of a single verification method. Eliminate the artifacts in the data of the water level time series verification data based on the anomaly recognition algorithm to obtain artifact-removed optimized data. Artifacts are data anomalies introduced by factors such as noise and measurement errors, which may have an adverse impact on the analysis of the water level time series data and model training. Through the anomaly recognition algorithm, these artifacts can be automatically identified and removed to ensure that the water level data is more real and accurate, avoiding the misleading influence of the artifacts on the results. Conduct a correction process for the extreme anomaly points of the artifact-removed optimized data based on physical constraints to obtain supervised time series water level data. Extreme anomaly points are usually deviations caused by measurement errors or abnormal events. Correcting these extreme values can significantly improve the reliability of the water level time series data. By introducing physical constraints, such as the physical range and variation law of the water level, it can be ensured that the corrected data conforms to the actual hydrological conditions, avoiding the adverse impact of extreme values on the model, and thus obtaining more accurate and reliable water level time series data. This correction process improves the quality of the data, ensuring that the generated water level time series data can effectively reflect the real hydrological changes.
[0120] The present invention also provides a supervised time-series water level data generation system for implementing the above-mentioned supervised time-series water level data generation method. The supervised time-series water level data generation system includes:
[0121] A model construction module, configured to obtain multi-source hydrological monitoring data; identify outliers and perform data preprocessing on the hydrological monitoring data to generate standardized hydrological basic data; construct a data generation model of a conditional generative adversarial network based on the standardized hydrological basic data;
[0122] A hydrological condition enhancement module, configured to perform conditional feature regulation driven by hydrological variables on the data generation model, optimize the parameters of the generation network, and generate condition input feature data with enhanced hydrological adaptability; inject the condition input feature data into the data generation model to generate preliminary water level time-series data;
[0123] A sampling optimization module, configured to perform stratified importance sampling on the preliminary water level time-series data, set differential generation weights according to preset rare extreme hydrological event data to obtain optimized sampling data for hydrological events; oversample the rare extreme hydrological event data by the bootstrap method to obtain high-quality extreme event data; perform distribution consistency evaluation based on statistical features according to the extreme event data and the optimized sampling data for hydrological events to obtain water level time-series data;
[0124] A data verification and correction module, configured to perform multi-scale verification on the water level time-series data to obtain water level time-series verification data, where the multi-scale verification includes spectral characteristic test and mutual information measurement; eliminate artifacts in the water level time-series verification data and perform extreme value outlier correction processing based on physical constraints to obtain supervised time-series water level data.
[0125] The model construction module in the present invention provides a solid data foundation for the subsequent generation of water level time series data by obtaining multi-source hydrological monitoring data. This module first cleans the original data through outlier identification and data preprocessing, removing noise and errors that do not conform to the normal hydrological change law, and generating standardized hydrological basic data. This standardization process not only improves the consistency and usability of the data, but also ensures that subsequent modeling will not be interfered by abnormal data. Based on these standardized hydrological basic data, a data generation model of a conditional generative adversarial network (GAN) is constructed, which can simulate the complex law of water level change and generate time series data that conforms to the actual hydrological characteristics, laying the foundation for the accuracy and reliability of the entire model. The hydrological condition enhancement module plays a key role in generating water level time series data. By regulating the conditional features of the data generation model driven by hydrological variables, this module can optimize the parameters of the generation network according to the actual situation of hydrological changes, so that the generated data is enhanced in hydrological adaptability. This hydrological adaptability-enhanced feature data ensures that the generated water level time series data can more realistically reflect the hydrological changes under different hydrological conditions and provides more accurate prediction results. The sampling optimization module is optimized during the data generation process, making the generated water level time series data more in line with the actual application requirements. By performing stratified importance sampling on the preliminary water level time series data, the representativeness of the generated data can be improved. Especially when facing rare extreme hydrological events, through the setting of differential generation weights, these events are preferentially sampled, thereby generating more accurate extreme event data. By oversampling through the bootstrap method, the quality of extreme event data can be further improved, enhancing the model's ability to identify these key events. Finally, these data are processed through distribution consistency evaluation based on statistical features, ensuring the accuracy and consistency of the water level time series data in different scenarios. The data verification and correction module comprehensively checks the generated water level time series data through multi-scale verification methods to ensure the reliability of the data. Spectrum characteristic test and mutual information measurement provide two important perspectives for the verification process. Spectrum characteristic verification can check the energy distribution characteristics of the data in different frequency ranges, and mutual information measurement reveals the correlation and information flow between the data, which helps to detect whether the data conforms to the expected hydrological characteristics. Through artifact elimination and extreme value outlier correction, the quality of the data is further improved, avoiding inaccurate results caused by measurement errors or noise. This process ensures that the generated water level time series data not only conforms to the actual hydrological law, but also has sufficient physical rationality and accuracy, and finally obtains high-quality supervised time series water level data. In summary, these modules cooperate with each other to provide a complete process from data acquisition, generation, optimization to verification and correction, ensuring the accuracy, reliability and applicability of the water level time series data, and can provide strong data support for applications such as hydrological prediction and risk assessment.
[0126] The present invention also provides a computer-readable storage medium storing a computer program, and the computer program is executed to implement the supervised time-series water level data generation method as described above.
[0127] Therefore, from any perspective, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Thus, all changes falling within the meaning and scope of the equivalent elements of the application documents are intended to be encompassed within the present invention.
[0128] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features invented herein.
Claims
1. A supervised method for generating time series water level data, characterized in that: The following steps are involved: Step S1: Acquire multi-source hydrological monitoring data; perform outlier identification and data preprocessing based on the hydrological monitoring data to generate standardized hydrological basic data; construct a data generation model of a conditional generative adversarial network based on the standardized hydrological basic data; Step S2: The data generation model is regulated by the conditional characteristics driven by hydrological variables, the parameters of the generation network are optimized, and the conditional input characteristic data with enhanced hydrological adaptability are generated; the conditional input characteristic data is injected into the data generation model to generate preliminary water level time series data. Step S2 is specifically as follows: Step S21: acquiring historical hydrological variable data, including temperature, humidity, precipitation, and atmospheric pressure information; performing feature dimensionality reduction processing on the historical hydrological variable data, and extracting key hydrological features, thereby obtaining key hydrological feature data; Step S22: designing a hydrological variable encoding network based on an attention mechanism based on the key hydrological characteristic data, and using the hydrological variable encoding network to perform nonlinear transformation on the key hydrological characteristic data, thereby obtaining conditional input characteristic data; Step S23: According to the conditional input feature data, the generator network of the data generation model is tuned for hydrological adaptability parameters, thereby obtaining hydrological adaptability optimization model parameters, wherein the hydrological adaptability parameter tuning is specifically to design a generation network parameter adaptive optimization strategy based on variational inference; Step S24: optimizing the generator network of the data generation model according to the hydrological adaptability optimization model parameters, thereby obtaining a data generation optimization model; Step S25: injecting the conditional input feature data into the generator network of the data generation optimization model, and generating preliminary water level time series data through adversarial training of the generator and the discriminator of the adversarial network; Step S3: Perform stratified importance sampling based on the preliminary water level time series data, and perform differentiated generation weight setting based on preset rare extreme hydrological event data to obtain optimized sampling data for hydrological events; The rare extreme hydrological event data are oversampled by the bootstrap method to obtain high-quality extreme event data; the distribution consistency evaluation based on statistical characteristics is performed on the extreme event data and the optimized sampling data of hydrological events to obtain water level time series data; Step S4: Perform multi-scale verification on the water level time series data to obtain water level time series verification data, wherein the multi-scale verification includes spectral characteristic verification and mutual information measurement; perform artifact elimination on the water level time series verification data, and perform extreme value anomaly correction processing based on physical constraints to obtain supervised time series water level data.
2. The supervised time series water level data generation method according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: collecting water level, rainfall, evaporation and flow information to obtain original hydrological monitoring data; Step S12: performing format conversion and time stamp calibration processing on the original hydrological monitoring data, thereby obtaining time-aligned multi-source hydrological monitoring data; Step S13: performing outlier identification based on statistical methods on the multi-source hydrological monitoring data, and performing interpolation repair processing and outlier removal processing to obtain hydrological monitoring filtered data; Step S14: performing data filling processing on the hydrological monitoring filtered data based on the time series regression model of adjacent data points, thereby obtaining hydrological monitoring cleansing data; Step S15: screening the hydrological monitoring cleaning data for characteristic variables highly correlated with the water level time series changes through the mutual information method, and generating derived variables, thereby obtaining hydrological characteristic extended data; Step S16: performing dimensionless processing on the hydrological characteristic extended data and performing data segmentation based on different time windows, thereby obtaining standardized hydrological basic data; Step S17: Construct a data generation model of a conditional generative adversarial network based on standardized hydrological basic data.
3. The supervised time series water level data generation method according to claim 2, characterized in that: Step S17 includes the following steps: Step S171: constructing a generator network structure of a conditional generative adversarial network based on standardized hydrological basic data, designing a hybrid architecture including multi-layer long short-term memory and convolutional neural network, thereby obtaining generator network structure data; Step S172: generating a discriminator network structure of the adversarial network according to the standardized hydrological basic data design conditions, and performing multi-domain consistency verification processing using a multi-scale time series discrimination strategy to obtain discriminator network structure data, wherein the multi-scale time series discrimination strategy includes time domain and frequency domain feature extraction modules; Step S173: defining the loss function of the generative adversarial network according to the discriminator network structure data and the generator network structure data, thereby obtaining loss function data, wherein the loss function includes loss functions of distribution similarity loss, temporal consistency loss, and hydrological physical constraint loss; Step S174: Integrate the discriminator network structure data, the generator network structure data and the loss function data into a data generation model.
4. The supervised time series water level data generation method according to claim 3, characterized in that: Step S25 includes the following steps: Step S251: using the conditional input feature data as conditional vector input data to generate the initial layer of the generator network in the optimization model, thereby constructing conditionally guided generator network input data; Step S252: performing multi-layer nonlinear transformation on the generator network input data through the generator network in the data generation optimization model, thereby generating candidate water level time series data; Step S253: inputting the candidate water level time series data into the discriminator network in the data generation optimization model, and performing authenticity discrimination and adversarial training according to the preset hydrological basic data, thereby obtaining adversarial training feedback data; Step S254: adaptively adjusting the parameters of the generator network in the data generation optimization model according to the adversarial training feedback data of the discriminator network, and optimizing the quality of the candidate water level time series data, thereby obtaining generator iterative optimization data; Step S255: Loop through steps S252 to S254 until the generated generator iterative optimization data meets the preset adversarial training convergence criteria, thereby obtaining preliminary water level time series data.
5. The supervised time series water level data generation method according to claim 4, characterized in that: Step S3 includes the following steps: Step S31: Designing a stratified sampling weight strategy according to the distribution characteristics and statistical attributes of the preliminary water level time series data, thereby obtaining sampling weight strategy data; Step S32: performing stratified importance sampling processing on the preliminary water level time series data according to the sampling weight strategy data, thereby obtaining importance stratified sampling data; Step S33: setting the differentiated generation weights of the sampling probabilities of the prominent extreme hydrological events based on the preset rare extreme hydrological event characteristics, thereby obtaining differentiated generation weight data; Step S34: performing weighted sampling processing of different types of hydrological events on the extreme hydrological event data according to the differentiated generated weight data, thereby obtaining optimized sampling data of hydrological events; Step S35: oversampling rare extreme hydrological event data by a bootstrap method, thereby obtaining high-quality extreme event data; Step S36: Perform distribution consistency assessment based on statistical characteristics according to importance stratified sampling data, extreme event data and hydrological event optimized sampling data to obtain water level time series data.
6. The supervised time series water level data generation method according to claim 5, characterized in that: Step S36 includes the following steps: Step S361: extracting statistical features from importance stratified sampling data, extreme event data, and hydrological event optimized sampling data, and constructing a multidimensional statistical feature vector; Step S362: designing a distribution distance calculation method based on statistical characteristics according to the multidimensional statistical feature vector, and performing a distribution similarity test based on Kolmogorov-Smirnov, thereby obtaining distribution consistency measurement data; Step S363: According to the distribution consistency measurement data, the importance stratified sampling data, the extreme event data and the hydrological event optimized sampling data are subjected to distribution feature reconciliation processing based on feature interpolation and smoothing technology for different data sets, thereby obtaining aligned water level time series data.
7. The supervised time series water level data generation method according to claim 6, characterized in that: Step S4 includes the following steps: Step S41: verifying the spectrum characteristics of the water level time series data to generate spectrum characteristics verification data, wherein the spectrum characteristics verification specifically uses wavelet transform to decompose multi-scale time series characteristics, analyzes the energy distribution of different time frequency segments, and evaluates the consistency of the spectrum characteristics of the water level time series data with the pre-acquired real water level data; Step S42: analyzing the correlation and information transfer characteristics between different variables of the water level time series data based on the mutual information measurement method, thereby obtaining mutual information enhanced verification data; Step S43: merging the spectrum characteristic verification data and the mutual information enhancement verification data into water level time series verification data; Step S44: Eliminate artifacts in the water level time series verification data based on an abnormality recognition algorithm, thereby obtaining artifact elimination optimized data; Step S45: Perform extreme value abnormal point correction processing based on physical constraints on the artifact removal optimization data to obtain supervised time series water level data.
8. A supervised time series water level data generation system, characterized in that: Used to execute the supervised time series water level data generation method according to claim 1, the supervised time series water level data generation system comprises: Model building module, used to obtain multi-source hydrological monitoring data; perform outlier identification and data preprocessing based on hydrological monitoring data to generate standardized hydrological basic data; build a data generation model of conditional generative adversarial network based on standardized hydrological basic data; The hydrological condition enhancement module is used to regulate the conditional characteristics driven by hydrological variables in the data generation model, optimize the parameters of the generation network, and generate conditional input characteristic data with enhanced hydrological adaptability; the conditional input characteristic data is injected into the data generation model to generate preliminary water level time series data; The sampling optimization module is used to perform stratified importance sampling based on preliminary water level time series data, and to perform differentiated generation weight settings based on preset rare extreme hydrological event data to obtain optimized sampling data for hydrological events; to oversample rare extreme hydrological event data through the self-help method to obtain high-quality extreme event data; to perform distribution consistency evaluation based on statistical characteristics based on extreme event data and hydrological event optimized sampling data to obtain water level time series data; The data verification and correction module is used to perform multi-scale verification on the water level time series data to obtain water level time series verification data, where the multi-scale verification includes spectral characteristic verification and mutual information measurement; the water level time series verification data is used to eliminate artifacts in the data and perform extreme value anomaly correction processing based on physical constraints to obtain supervised time series water level data.
9. A computer-readable storage medium storing a computer program, characterized in that: The computer program is executed to implement the supervised time-series water level data generation method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Automatic modeling system for data imbalance
CN113177642A
Drainage basin multi-point water level prediction and early warning method based on generative adversarial network
CN115688579A