Hydrological abnormal data detection method and system based on transfer learning and medium

Through transfer learning and dynamic distribution alignment technology, the problem of consistency of hydrological monitoring data across river basins has been solved, the accuracy and stability of hydrological anomaly data detection have been improved, the annotation cost has been reduced, and the accuracy and timeliness of hydrological flood reporting data have been improved.

CN120671029APending Publication Date: 2025-09-19CHINA THREE GORGES CORPORATION +2
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510623824.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing hydrological monitoring technologies face the difficulty of ensuring data quality consistency across river basins. Traditional supervised learning methods have a sharp drop in performance when applied in different river basins, and abnormal data misleads water conservancy scheduling, leading to increased risks of floods and droughts, affecting water resource management and ecological security.

Method used

A transfer learning-based method is used to obtain source and target hydrological datasets, calculate multiple features and perform dynamic distribution alignment, combine a small amount of manual annotation to improve model performance, and use random forest, support vector machine or deep learning models for anomaly detection.

Benefits of technology

It improves the accuracy of hydrological anomaly data detection, reduces the dependence on target data set annotation, reduces the cost of manual annotation, is suitable for the stable detection of large-scale time series data, and improves the accuracy and timeliness of hydrological flood reporting data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671029A_ABST
    Figure CN120671029A_ABST
Patent Text Reader

Abstract

The invention relates to a hydrological abnormal data detection method and system based on transfer learning and a medium. The method comprises the following steps: acquiring a source hydrological data set and a target hydrological data set; calculating various characteristics of the source data and the target data; aligning the feature distribution of the source data set with the feature distribution of the target data set through a dynamic distribution alignment algorithm, and migrating the source basin model to a target basin; key feature samples of the target drainage basin are extracted for manual labeling; and identifying the abnormal data of the target data by using the updated anomaly detection model. The method is especially suitable for anomaly detection of cross-basin water level and flow data, can efficiently and accurately detect time sequence anomaly data in hydrological data, helps to improve the reliability of the hydrological data, and effectively solves the problem of model failure caused by basin hydrological feature differences.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of hydrological monitoring technology, and in particular to a method, system and medium for detecting hydrological anomaly data based on transfer learning. Background Art

[0002] Hydrological monitoring data, especially water level data, is directly related to flood prevention and disaster reduction, water resource allocation, and ecological security. It is a crucial foundation for protecting people's lives and property and supporting major water conservancy decisions. Abnormal fluctuations in water level data are often precursors to floods, droughts, or engineering hazards. The accuracy and timeliness of water level monitoring directly impact the reliability of disaster warnings.

[0003] Abnormal data can mislead water and rainfall forecasts and reservoir operations, significantly increasing the risk of dam failure and urban flooding, and posing a potential threat to watershed security. Deviations in key parameters such as water levels and flows can trigger erroneous flood evolution model calculations, leading to inappropriate storage and discharge strategies for reservoirs. For example, underestimating peak flood flows can delay flood discharge orders, creating the risk of exceeding reservoir capacity limits; while overestimating incoming water can lead to unnecessary water discharge, exacerbating the risk of overflowing downstream rivers. Furthermore, the failure of rainfall-flood coupling analysis under the interference of abnormal data can cause an imbalance between urban drainage system operations and surface runoff control, triggering a chain reaction of disasters in extreme weather conditions.

[0004] In water resource management, abnormal data can distort water supply and demand balance analysis and lead to inaccurate scheduling of inter-basin water transfer projects. For example, deviations in flow monitoring can cause a misalignment between reservoir storage plans and regional water demand, resulting in water shortages for agricultural irrigation or insufficient water replenishment for ecological conservation. Furthermore, the long-term accumulation of abnormal data can disrupt the continuity of hydrological sequences, undermining the reliability of water resource carrying capacity assessments and threatening water supply security and the stability of aquatic ecosystems. Such risks not only threaten the structural safety of water conservancy facilities but also potentially undermine the harmonious relationship between people and water within the basin, endangering the lives and property of coastal residents and the sustainable development of the ecological environment.

[0005] In the field of hydrological monitoring, river monitoring data (water level, flow, etc.) has significant temporal and spatial heterogeneity:

[0006] 1. Basin specificity: Different basins have different parameters such as topographic slope, riverbed roughness, and tributary inflow ratio, which lead to nonlinear differences in the water level-discharge relationship under the same rainfall conditions;

[0007] 2. Equipment heterogeneity: The water level meters (including bubble and radar types) and flow meters (including acoustic Doppler and electromagnetic types) used at various hydrological stations have fundamental differences, resulting in widely varying data noise distributions.

[0008] 3. Sudden changes in hydrological conditions: Affected by human activities, some river channels have experienced "abnormalization".

[0009] However, current hydrological monitoring faces significant challenges. First, with the upgrading of monitoring equipment (for example, traditional float-type water level gauges are gradually being replaced by high-precision sensors such as radar and pressure sensors), the frequency and volume of data collection have increased significantly, with single stations generating millions of data points annually. However, anomaly annotation relies on comprehensive analysis by professionals based on multi-dimensional information such as rainfall patterns and upstream and downstream water balances, resulting in high costs and low efficiency. Second, differences in topography and riverbed characteristics across different river basins lead to nonlinear variations in the water level-discharge relationship. Furthermore, due to differing monitoring equipment principles (for example, bubble-type and acoustic Doppler instruments produce significantly different data noise distributions), ensuring data consistency across river basins is difficult. When existing supervised learning methods are directly applied to river basin B, performance degrades significantly after training a model in river basin A due to offsetting basin characteristic distributions. These factors include differences in the slopes of the water level-discharge relationship curves and mismatched flood peak lag time distributions.

[0010] Traditional methods often misinterpret normal fluctuations, such as water conservancy scheduling, as anomalies due to their neglect of inter-basin feature migration and hydrophysical constraints, severely restricting the in-depth application of hydrological big data. Therefore, there is an urgent need to develop intelligent detection methods that integrate device heterogeneity adaptation, cross-basin knowledge transfer, and physical rules to overcome data quality bottlenecks and strengthen hydrological security. Summary of the Invention

[0011] The purpose of the embodiments of the present application is to overcome the shortcomings of the existing technology and provide an intelligent detection method, system and medium for hydrological flood reporting data, which is fully based on the time series characteristics of water flood reporting data of river sections, and takes into account the evolution process of water flow in upstream and downstream relationships, to enhance its prediction characteristics, so as to replace manual experience judgment and improve detection accuracy.

[0012] To achieve the above objectives, this application provides the following technical solutions:

[0013] In a first aspect, an embodiment of the present application provides a method for detecting hydrological anomaly data based on transfer learning, comprising the following steps:

[0014] S1. Obtain a source hydrological dataset and a target hydrological dataset, wherein the source hydrological dataset includes labeled time series data, and the target hydrological dataset includes unlabeled or partially labeled time series data;

[0015] S2. Calculate various inherent characteristics of the source data and target data, including statistical characteristics, prediction error characteristics, time characteristics, and hydrological characteristics, and calculate the difference in characteristic distribution between the source data and the target data through distance measurement;

[0016] S3, align the feature distribution of the source dataset with the feature distribution of the target dataset through the dynamic distribution alignment algorithm, and migrate the source watershed model to the target watershed;

[0017] S4. Extract key feature samples from the target watershed and manually annotate them. The amount of annotations should not exceed 3% of the total data volume. Add the training data to the target data to train the anomaly detection model to improve model performance.

[0018] S5. Apply the updated anomaly detection model to identify abnormal data in the target data.

[0019] The time series data in step S1 includes a combination of at least one element of water level, flow, water storage capacity, rainfall, evaporation, water temperature and sediment content.

[0020] The statistical features in step S2 include mean, variance, number of intersections, first-order autocorrelation, residual autocorrelation, trend strength, linearity, curvature, entropy, ARCH test p-value, and GARCH test p-value features; the prediction error features use SARIMA and Holt-Winters models to perform time series prediction, and calculate the average error, root mean square error, mean absolute error, mean percentage error, and mean absolute percentage error through weighted integrated prediction results; the prediction error features include calculated differential value, maximum variance offset, maximum horizontal offset, maximum KL divergence offset, the degree of variance change of the residual part, and the maximum continuous length in the discretization bucket; the hydrological features include water level fluctuation, flow acceleration, and reporting interval.

[0021] The difference in characteristic distribution between the source data and the target data in step S2 is calculated using Wasserstein distance and Bhattacharyya distance.

[0022] The dynamic distributed alignment algorithm in step S2 is enabled in at least one of the following scenarios:

[0023] (1) There are significant seasonal differences between the source and target domains;

[0024] (2) The data acquisition equipment in the target domain is different from that in the source domain;

[0025] (3) Detection of sudden hydrological events such as floods or droughts in the target area.

[0026] In step S3, the dynamic distribution alignment algorithm adopts the Deep CORAL algorithm to make the source dataset and the target dataset closer in the feature space.

[0027] In step S4, the key feature samples adopt a strategy combining uncertainty and context diversity to select the most representative data points for annotation. The sampling strategy includes:

[0028] (1) Uncertainty sampling: select the first k samples with the highest model prediction entropy;

[0029] (2) Diversity sampling: Select representative samples of different feature clusters based on the K-Means clustering algorithm.

[0030] The anomaly detection model in step S4 adopts one of random forest, support vector machine or deep learning model or an integrated model of any combination.

[0031] In a second aspect, an embodiment of the present application provides a hydrological anomaly data detection system based on transfer learning, comprising:

[0032] A hydrological data set acquisition module is configured to acquire a source hydrological data set and a target hydrological data set, wherein the source hydrological data set includes annotated time series data and the target hydrological data set includes unannotated or partially annotated time series data;

[0033] The self-feature calculation module calculates various self-features of the source data and the target data, including statistical features, prediction error features, time features, and hydrological features, and calculates the feature distribution differences between the source data and the target data through distance measurement;

[0034] The distribution alignment module aligns the feature distribution of the source dataset with that of the target dataset through a dynamic distribution alignment algorithm, and migrates the source watershed model to the target watershed;

[0035] The anomaly detection model training module extracts key feature samples from the target watershed and manually annotates them. The annotated amount does not exceed 3% of the total data volume. The training data is added to the target data to train the anomaly detection model to improve model performance.

[0036] The abnormal data identification module applies the updated anomaly detection model to identify abnormal data in the target data.

[0037] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, which stores program code. When the program code is executed by a processor, the steps of the hydrological anomaly data detection method based on transfer learning as described above are implemented.

[0038] Compared with the prior art, the present invention has the following beneficial effects:

[0039] 1. Improve detection accuracy through transfer learning and reduce dependence on labeled data of the target dataset.

[0040] 2. The proposed watershed feature mapping network can separate general hydrological features from watershed-specific features, effectively reducing the cost of manual annotation.

[0041] 3. Suitable for anomaly detection of large-scale time series data and improving the stability of reporting data. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.

[0043] Figure 1 The figure is a flowchart of the overall process of the method of the present invention.

[0044] Figure 2 Schematic diagram of feature alignment during transfer learning.

[0045] Figure 3 Schematic diagram of the sample selection mechanism in data annotation. DETAILED DESCRIPTION

[0046] The technical solutions in the embodiments of the present application will be described below in conjunction with the accompanying drawings. It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.

[0047] The terms "comprises," "comprising," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not preclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0048] The terms "first," "second," etc. are only used to distinguish one entity or operation from another entity or operation, and are not to be understood as indicating or implying relative importance, nor are they to be understood as requiring or implying any actual relationship or order between these entities or operations.

[0049] See also Figures 1 to 3 ,A hydrological anomaly data detection method based on transfer learning, comprising the following steps:

[0050] S1. Obtain a source hydrological dataset and a target hydrological dataset, wherein the source hydrological dataset includes labeled time series data, and the target hydrological dataset includes unlabeled or partially labeled time series data;

[0051] S2. Calculate various inherent characteristics of the source data and target data, including statistical characteristics, prediction error characteristics, time characteristics, and hydrological characteristics, and calculate the difference in characteristic distribution between the source data and the target data through distance measurement;

[0052] S3, align the feature distribution of the source dataset with the feature distribution of the target dataset through the dynamic distribution alignment algorithm, and migrate the source watershed model to the target watershed;

[0053] S4. Extract key feature samples from the target watershed and manually annotate them. The amount of annotations should not exceed 3% of the total data volume. Add the training data to the target data to train the anomaly detection model to improve model performance.

[0054] S5. Apply the updated anomaly detection model to identify abnormal data in the target data.

[0055] In step S2, multidimensional time series features, including statistical features, prediction error features, and temporal features, are extracted from the source dataset, and the data is sub-domained using the K-means clustering method. Here, five years of historical water level, rainfall, and flow data from a hydrological station in the middle reaches of the Yangtze River (designated Site A) are used as an example. Professional hydrologists at this site have annotated abnormal data, including sudden water level increases due to heavy rain, abnormal flow fluctuations due to equipment failures, and missing data due to data collection system failures. The target hydrological dataset, on the other hand, comes from a site in the middle reaches of the Yangtze River (designated Site B) collected over a period of one year. Due to a lack of professional staff at this site to annotate abnormal data, the majority of the data remains unlabeled.

[0056] Typical examples of two site datasets are shown in the following table:

[0057]

[0058] For statistical characteristics, the mean, variance, number of crossing points, first-order autocorrelation, residual autocorrelation, trend strength, linearity, curvature, entropy, ARCH test p-value, and GARCH test p-value are calculated based on a sliding window whose window size is determined by the main period estimated by Fourier transform. These characteristics describe the statistical characteristics of the data, as well as hydrological characteristics such as water level fluctuation, flow acceleration, and reporting interval. The mean is the average of all data, and the formula is as follows:

[0059] The variance is the variance of all data, and the formula is as follows:

[0060]

[0061] Crossover counts the number of times the series crosses the moving average line.

[0062]

[0063] Linearity is the linear strength of the trend decomposition based on STL (Seasonal-Trend Decomposition using Loess). Curvature is the curvature strength of the trend decomposition based on STL. Given a time series , the trend component is obtained by STL decomposition:

[0064]

[0065] in, is the trend component, is the seasonal term, is the residual. For the trend component Fit a linear regression model:

[0066]

[0067] in is the time index, is the intercept, is the slope, is the error term.

[0068] Linearity is also called the coefficient of determination, which is the coefficient of determination of the regression model. , for STL decomposition, is defined as:

[0069]

[0070] in,

[0071] is the model prediction value, is the mean of the trend component.

[0072] Curvature is used to quantify the trend component of a time series The nonlinear fluctuation intensity of is approximated by calculating the second-order derivative of the trend component. The curvature is defined as the root mean square of the second-order difference of the trend component:

[0073]

[0074] The second-order difference for:

[0075]

[0076] The ARCH test p-value is the Lagrange Multiplier (LM) test p-value of the ARCH model. The full name of the ARCH model (Autoregressive Conditional Heteroskedasticity Model) is the "Autoregressive Conditional Heteroskedasticity Model". At a certain moment, the occurrence of a noise obeys the normal distribution. The mean of the normal distribution is zero, and the variance is a quantity that changes with time. And this variance that changes with time is a linear combination of the squares of the finite noise values ​​in the past. This constitutes the autoregressive conditional heteroskedasticity model. If the square of the error term obeys process, namely:

[0077]

[0078] The above model is called autoregressive conditional heteroskedasticity model, or ARCH model for short. , construct LM statistics:

[0079]

[0080] in, is the sample size, is the coefficient of determination of the auxiliary regression model, The degrees of freedom are The chi-square distribution of Down, The value corresponds to the highest order The value of , so that the LM statistic still satisfies:

[0081]

[0082] GARCH test Lagrange Multiplier Test for GARCH Model with Value Characteristics value.

[0083] Multidimensional time series features were extracted from the hydrological data at Sites A and B. Taking flow data as an example, based on Fourier transform analysis, the main period of flow data at Site A is 24 hours. Based on this, a sliding window size of 24 hours was set to calculate statistical features. For the flow data from January 1 to January 7, 2022, the following statistical features were calculated:

[0084]

[0085] For the prediction error characteristics, SARIMA, Holt, Holt-Winters and STL models are used for time series forecasting. The five types of time error characteristics, namely average error, root mean square error, mean absolute error, average percentage error and mean absolute percentage error, are calculated by weighted integrated forecast results and then analyzed in three rolling time windows (lengths are , , , Calculated separately on the main period, a total of 15 features are generated, reflecting the deviation between the actual value and the predicted value. The average error formula is as follows:

[0086]

[0087] The root mean square error formula is as follows:

[0088]

[0089] The mean absolute error formula is as follows:

[0090]

[0091] The formula for mean percentage error is as follows:

[0092]

[0093] The mean absolute percentage error formula is as follows:

[0094]

[0095] Since different models have different prediction accuracy, it is necessary to assign weights to the models. The formula for calculating the weighted prediction error for a model is as follows:

[0096]

[0097] in, It's a model In time predictions, It's a model In time The prediction error, It's time ensemble predictions.

[0098] For water level data, multiple time series forecasting models are applied:

[0099] SARIMA model: Considering the seasonality of water level data, the SARIMA(2,1,1)(1,1,1)24 model was used, where the seasonal period was set to 24 h to capture the intraday periodic variations.

[0100] Holt-Winters Model: Setup , the seasonal period is 24, which is suitable for processing water level data with trends and seasonality.

[0101] Based on the above model, the prediction error characteristics are calculated and illustrated using traffic data as an example:

[0102]

[0103] It can be seen that the prediction error on January 7 was significantly higher than that on other dates. Taking all factors into consideration, the anomaly may have come from a false alarm on that day, causing a traffic surge beyond the model's prediction range.

[0104] For time features, since a sharp change in an indicator is likely to be an anomaly, in order to analyze the changes in time series data over time, features are identified by comparing the data in two consecutive windows. At the same time, the difference between the values ​​in the two windows is compared. For example, setting the window , and corresponding differences are obtained, including six indicators: difference value, maximum variance offset, maximum horizontal offset, maximum KL divergence offset, variance change degree of the residual part, and maximum continuous length in the discretization bucket, to capture the dynamic changes of time series.

[0105] For the water level data at site A, calculate the change characteristics of adjacent time windows:

[0106]

[0107] All time characteristics on January 7 were significantly higher than usual, especially the differential value reached 0.78 meters, indicating that there were abnormal situations such as false alarms.

[0108] Transfer learning requires narrowing the gap between the source domain and the target domain. Therefore, by selecting samples that are similar between the source domain and the target domain, the gap can be reduced. The difference in feature distribution between the source data and the target data is calculated, including Wasserstein distance, Bhattacharyya distance, etc. Wasserstein distance measures the difference between the source distribution and the target data. Convert to target distribution Minimum effort required, considering feature space geometry:

[0109]

[0110] in is a joint distribution set. Introducing feature weight vector Later improved to:

[0111]

[0112] Here is the one-dimensional Wasserstein distance.

[0113] Bhattacharyya distance quantifies the probability density function of two distributions The degree of overlap:

[0114]

[0115] For discrete distributions it simplifies to:

[0116]

[0117] After combining the feature importance weights:

[0118]

[0119] Use Wasserstein distance and Bhattacharyya distance to calculate the difference in feature distribution between site A and site B data:

[0120]

[0121] From the calculation results, it can be seen that the two stations have the greatest difference in the distribution of hydrological characteristics, which may be due to the fact that the stations are located in different geographical locations in the river, and have different topography, river section, upstream catchment area and other factors.

[0122] In step S3, the sub-source domain is then constructed and aligned. The feature alignment method Deep CORAL algorithm is used to narrow the feature distribution differences between the source dataset and the target dataset, and the benchmark anomaly detection model is trained. After extracting features from the source domain labeled dataset, the K-means algorithm is used to divide it into K sub-source domains, each of which represents a similar data distribution. After extracting the same features from the unlabeled data in the target domain, the Euclidean distance between it and the center of each sub-source domain is calculated and assigned to the nearest sub-source domain. The Deep CORAL algorithm is used to align the feature covariance matrix of the sub-source domain and the target domain, minimize the distribution difference, and generate the aligned sub-source domain feature set. For the source domain data

[0123] , target domain data , the covariance distance of the DeepCORAL algorithm is defined as follows:

[0124]

[0125] Among them, the feature covariance matrix and Defined as:

[0126]

[0127] Gradient with respect to input features:

[0128]

[0129] The model is optimized using a gradient descent algorithm to align the source and target domain distributions. The uncertainty of each sample in the target dataset is calculated, and representative samples are selected for manual annotation based on contextual information.

[0130] In this example, the hydrological data of site A is constructed into sub-source domains, and the K-means algorithm is used to divide the data of site A into four sub-source domains:

[0131] Sub-source domain 1: normal hydrological period (water level 30% below the warning line)

[0132] Sub-source domain 2: Early warning hydrological period (water level 0-30% below the warning line)

[0133] Sub-source domain 3: Super-alert hydrological period (water level exceeds the warning line by 0-30%)

[0134] Sub-source domain 4: Flood season (water level exceeds the warning line by more than 30%)

[0135] The Deep CORAL algorithm is used to align the feature distributions of the source and target domains. Taking the water level data in the range of 12.0-25.0m as an example, the covariance matrix before and after alignment is as follows:

[0136] Before alignment:

[0137] Source domain covariance matrix :

[0138]

[0139] Target domain covariance matrix :

[0140]

[0141] After alignment:

[0142] Source domain covariance matrix :

[0143]

[0144] Target domain covariance matrix :

[0145]

[0146] After alignment, the Frobenius norm distance of the covariance matrix decreases from 0.82 to 0.11, indicating that the difference in feature distribution is significantly reduced.

[0147] Use a small amount of labeled data to incrementally train the model to improve detection accuracy. First, calculate the uncertainty of the target domain sample. The calculation formula is as follows:

[0148]

[0149] Sort by descending order of uncertainty, giving priority to samples with low model confidence.

[0150] Defining a time context window , select the sample. Traverse the sorted samples, if the current sample and the context window of the selected sample If there is no overlap, it is added to the candidate set to avoid redundant annotations. For each round of annotation, 60 samples are manually annotated and added to the training set. The anomaly monitoring model parameters are updated, and after three rounds of iteration, the final detection model is obtained. The annotation volume only requires 1%-5% of the target data. The anomaly detection model uses a random forest, support vector machine, or deep learning model, or an ensemble of any combination.

[0151] Apply the updated model to perform anomaly detection on the target dataset and adjust the detection threshold based on business needs to optimize detection performance. The online detection process includes inputting the time series to be detected, extracting multidimensional features according to the same process, inputting the features into the final model, calculating the anomaly probability, and setting the threshold (0.6-0.8) to identify anomalies. The detection threshold includes dynamically adjusting the number of clusters K and the context window. .

[0152] The embodiment of the present application provides a hydrological anomaly data detection system based on transfer learning, comprising:

[0153] A hydrological data set acquisition module is configured to acquire a source hydrological data set and a target hydrological data set, wherein the source hydrological data set includes annotated time series data and the target hydrological data set includes unannotated or partially annotated time series data;

[0154] The self-feature calculation module calculates various self-features of the source data and the target data, including statistical features, prediction error features, time features, and hydrological features, and calculates the feature distribution differences between the source data and the target data through distance measurement;

[0155] The distribution alignment module aligns the feature distribution of the source dataset with that of the target dataset through a dynamic distribution alignment algorithm, and migrates the source watershed model to the target watershed;

[0156] The anomaly detection model training module extracts key feature samples from the target watershed and manually annotates them. The annotated amount does not exceed 3% of the total data volume. The training data is added to the target data to train the anomaly detection model to improve model performance.

[0157] The abnormal data identification module applies the updated anomaly detection model to identify abnormal data in the target data.

[0158] An embodiment of the present application provides a computer-readable storage medium, which stores program code. When the program code is executed by a processor, the steps of the hydrological anomaly data detection method based on transfer learning as described above are implemented.

[0159] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0160] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0161] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0162] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0163] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0164] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0165] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can be implemented using any method or technology to store information. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change RAM (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0166] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. For those skilled in the art, various modifications and variations of the present application are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A method for detecting hydrological anomaly data based on transfer learning, characterized in that: The following steps are involved: S1. Obtain a source hydrological dataset and a target hydrological dataset, wherein the source hydrological dataset includes labeled time series data, and the target hydrological dataset includes unlabeled or partially labeled time series data; S2. Calculate various inherent characteristics of the source data and target data, including statistical characteristics, prediction error characteristics, time characteristics, and hydrological characteristics, and calculate the difference in characteristic distribution between the source data and the target data through distance measurement; S3, align the feature distribution of the source dataset with the feature distribution of the target dataset through the dynamic distribution alignment algorithm, and migrate the source watershed model to the target watershed; S4. Extract key feature samples from the target watershed and manually annotate them. The amount of annotations should not exceed 3% of the total data volume. Add the training data to the target data to train the anomaly detection model to improve model performance. S5. Apply the updated anomaly detection model to identify abnormal data in the target data.

2. The method for detecting hydrological anomaly data based on transfer learning according to claim 1, characterized in that: The time series data in step S1 includes a combination of at least one element of water level, flow, water storage capacity, rainfall, evaporation, water temperature and sediment content.

3. The method for detecting hydrological anomaly data based on transfer learning according to claim 1, characterized in that: The statistical features in step S2 include mean, variance, number of intersections, first-order autocorrelation, residual autocorrelation, trend strength, linearity, curvature, entropy, ARCH test p-value, and GARCH test p-value features; the prediction error features use SARIMA and Holt-Winters models to perform time series prediction, and calculate the average error, root mean square error, mean absolute error, mean percentage error, and mean absolute percentage error through weighted integrated prediction results; the prediction error features include calculated differential value, maximum variance offset, maximum horizontal offset, maximum KL divergence offset, the degree of variance change of the residual part, and the maximum continuous length in the discretization bucket; the hydrological features include water level fluctuation, flow acceleration, and reporting interval.

4. The method for detecting hydrological anomaly data based on transfer learning according to claim 1, characterized in that: The difference in characteristic distribution between the source data and the target data in step S2 is calculated using Wasserstein distance and Bhattacharyya distance.

5. The method for detecting hydrological anomaly data based on transfer learning according to claim 1, characterized in that: The dynamic distributed alignment algorithm in step S2 is enabled in at least one of the following scenarios: (1) There are significant seasonal differences between the source and target domains; (2) The data acquisition equipment in the target domain is different from that in the source domain; (3) Detection of sudden hydrological events such as floods or droughts in the target area.

6. The method for detecting hydrological anomaly data based on transfer learning according to claim 1, characterized in that: In step S3, the dynamic distribution alignment algorithm adopts the Deep CORAL algorithm to make the source dataset and the target dataset closer in the feature space.

7. The method for detecting hydrological anomaly data based on transfer learning according to claim 1, characterized in that: In step S4, the key feature samples adopt a strategy combining uncertainty and context diversity to select the most representative data points for annotation. The sampling strategy includes: (1) Uncertainty sampling: select the first k samples with the highest model prediction entropy; (2) Diversity sampling: Select representative samples of different feature clusters based on the K-Means clustering algorithm.

8. The method for detecting hydrological anomaly data based on transfer learning according to claim 1, characterized in that: The anomaly detection model in step S4 adopts one of random forest, support vector machine or deep learning model or an integrated model of any combination.

9. A hydrological anomaly data detection system based on transfer learning, characterized in that: include, A hydrological data set acquisition module is configured to acquire a source hydrological data set and a target hydrological data set, wherein the source hydrological data set includes annotated time series data and the target hydrological data set includes unannotated or partially annotated time series data; The self-feature calculation module calculates various self-features of the source data and the target data, including statistical features, prediction error features, time features, and hydrological features, and calculates the feature distribution differences between the source data and the target data through distance measurement; The distribution alignment module aligns the feature distribution of the source dataset with that of the target dataset through a dynamic distribution alignment algorithm, and migrates the source watershed model to the target watershed; The anomaly detection model training module extracts key feature samples from the target watershed and manually annotates them. The annotated amount does not exceed 3% of the total data volume. The training data is added to the target data to train the anomaly detection model to improve model performance. The abnormal data identification module applies the updated anomaly detection model to identify abnormal data in the target data.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores program code, and when the program code is executed by a processor, the steps of the hydrological anomaly data detection method based on transfer learning according to any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Abnormity detection method and system based on context residual learning

    CN121479701A

  • An anomaly detection method and system based on contextual residual learning

    CN121479701B