Hydrological data automatic classification and dynamic label generation method and system

CN122286454BActive Publication Date: 2026-09-11JIANGXI CHANGDA QINGKE INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610758390.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-29
Publication Date
2026-09-11
Estimated Expiration
2046-05-29

AI Technical Summary

Technical Problem

这种极端的类别不平衡问题,导致传统的分类算法(如决策树、支持向量机等)在训练时严重偏向多数类,难以有效学习和识别少数类样本的复杂特征模式,造成预警事件的漏报率较高

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122286454B_ABST
    Figure CN122286454B_ABST
Patent Text Reader

Abstract

The application provides a hydrological data automatic classification and dynamic label generation method and system, which comprises the following steps: performing noise identification and data processing on an unbalanced data set constructed by hydrological monitoring data, and using a local overlap index of samples to traverse all minority class samples to obtain optimized samples; using a diffusion probability model to perform distribution reconstruction and oversampling on the optimized samples to form a balanced data set; using a distance weighted roulette wheel selection algorithm to perform probability undersampling on majority class samples to form a sampling balanced data set; taking the sampling balanced data set as a source domain and taking an original hydrological data set as a target domain, and training a base classifier through a transfer learning algorithm; and performing iterative training on the base classifier according to an improved adaptive boosting algorithm to obtain a label classification model, and inputting real-time hydrological data to be classified into the label classification model to generate a hydrological state label. The application can significantly improve the identification accuracy of early warning events.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method and system for automated classification and dynamic label generation of hydrological data. Background Technology

[0002] Hydrological monitoring is a crucial foundation for flood control and disaster reduction, water resource management, and water ecological protection. With the widespread application of technologies such as the Internet of Things and remote sensing, the hydrological monitoring network has been continuously improved, accumulating massive amounts of multi-source, high-frequency hydrological monitoring data (such as water level, flow rate, and rainfall). This data provides valuable resources for understanding hydrological patterns and issuing early warnings of floods. However, how to automatically and accurately identify key anomalies or early warning states (such as exceeding warning levels or sudden floods) from this massive amount of data is a core challenge currently facing the field of hydrological data processing.

[0003] In practical applications, hydrological status data often exhibits significant class imbalance. Most of the time, hydrological conditions remain within the normal or safe range (i.e., the majority class), while warnings or hazardous conditions requiring special attention (i.e., the minority class) occur extremely infrequently. This extreme class imbalance causes traditional classification algorithms (such as decision trees and support vector machines) to be severely biased towards the majority class during training, making it difficult for them to effectively learn and identify the complex feature patterns of minority class samples, resulting in a high false negative rate for warning events. Furthermore, hydrological data commonly contains noise and class overlap, such as outliers caused by sensor malfunctions or samples with blurred boundaries between normal and critical states. These factors further interfere with classifier training, reducing the model's robustness and generalization ability. Summary of the Invention

[0004] Based on this, the purpose of the present invention is to provide a method and system for automated classification and dynamic label generation of hydrological data, so as to at least solve the shortcomings of the above-mentioned technologies.

[0005] This invention proposes a method for automated classification and dynamic label generation of hydrological data, comprising: Collect hydrological monitoring data uploaded from various platforms, and construct an imbalanced dataset containing majority class and minority class samples based on the hydrological monitoring data; The imbalanced dataset is subjected to noise identification and data processing. For each sample in the imbalanced dataset, its Euclidean distance with other samples is calculated, and a local overlap index of the samples is constructed. The local overlap index is used to traverse all minority class samples to obtain optimized samples. A diffusion probability model is constructed, and the distribution of the optimized samples is reconstructed and oversampled using the diffusion probability model to generate minority class samples that are in balance with the number of majority class samples, so as to form a balanced dataset. The majority class samples are probabilistically downsampled using a distance-weighted roulette wheel selection algorithm to form a sampled balanced dataset. The sampled balanced dataset is used as the source domain, and the original labeled hydrological dataset is used as the target domain. A base classifier is trained using a transfer learning algorithm. The improved adaptive boosting algorithm and the base classifier are iteratively trained to obtain a label classification model. The real-time hydrological data to be classified is then input into the label classification model to generate the corresponding hydrological status label.

[0006] Furthermore, the steps of collecting hydrological monitoring data uploaded from various platforms and constructing an imbalanced dataset containing majority and minority class samples based on the hydrological monitoring data include: Read the original observation sequence of each monitoring section within a time period from the hydrological database or real-time transmission system; The original observation sequence is cleaned and standardized to obtain a standard observation sequence; Based on hydrological operational rules, minority and majority class labels are defined, and the standard observation sequences are processed according to the minority and majority class labels to obtain an imbalanced dataset.

[0007] Furthermore, the imbalanced dataset undergoes noise identification and data processing. For each sample in the imbalanced dataset, its Euclidean distance to other samples is calculated, and a local overlap index is constructed. The optimized samples are then obtained by iterating through all minority class samples using this local overlap index. For each sample in the imbalanced dataset, calculate its Euclidean distance to other samples and construct a local overlap index for the samples. Use the local overlap index to traverse all minority class samples and divide the minority class samples into dangerous samples, safe samples, and noisy samples. Delete the noisy samples and retain dangerous samples and safe samples to obtain the cleaned minority class samples. By traversing all majority class samples using the local overlap index, majority class samples in which the number of minority class samples in the neighborhood exceeds the number of majority class samples are deleted to obtain purified majority class samples. The purified minority class samples and the purified majority class samples are merged to obtain optimized samples.

[0008] Furthermore, the steps for constructing a diffusion probability model include: Define the diffusion step number and the noise variance sequence, and perform forward noise addition on the initial minority class data according to the noise variance sequence and the diffusion step number to obtain noisy data; The noise-added data and the corresponding time step are input into a preset neural network model to obtain the corresponding predicted noise, and a corresponding noise loss function is constructed based on the predicted noise. A classification loss function is constructed for the neural network model. The classification loss function is then weighted and fused with the noise loss function to obtain a total loss function. The total loss function is then used to optimize the neural network model to obtain a diffusion probability model.

[0009] Furthermore, the step of using a distance-weighted roulette wheel selection algorithm to probabilistically downsample the majority class samples to form a balanced sampled dataset includes: Define the weighted Euclidean distance sum from the majority class sample to all minority class samples in the balanced dataset as the fitness function, and calculate the corresponding fitness value according to the fitness function; Based on the fitness value, the majority class samples to be retained are selected by the distance-weighted roulette wheel selection algorithm to obtain the majority class sample set. The majority class sample set is then fused with the minority class samples in the balanced dataset to form a sampled balanced dataset.

[0010] This invention also proposes an automated classification and dynamic label generation system for hydrological data, comprising: The data acquisition module is used to collect hydrological monitoring data uploaded from various platforms and construct an imbalanced dataset containing majority class and minority class samples based on the hydrological monitoring data. The sample optimization module is used to identify noise and process data in the imbalanced dataset. For each sample in the imbalanced dataset, it calculates the Euclidean distance between it and other samples, constructs a local overlap index of the samples, and uses the local overlap index to traverse all minority class samples to obtain optimized samples. The model building module is used to build a diffusion probability model and use the diffusion probability model to reconstruct the distribution and oversample the optimized samples to generate minority class samples that are balanced with the number of majority class samples, so as to form a balanced dataset. The classifier training module is used to perform probability downsampling on the majority class samples using a distance-weighted roulette wheel selection algorithm to form a sampled balanced dataset. The sampled balanced dataset is used as the source domain, and the original hydrological dataset with labels is used as the target domain. A base classifier is trained through a transfer learning algorithm. The label classification module is used to iteratively train the improved adaptive boosting algorithm and the base classifier to obtain a label classification model, and input the real-time hydrological data to be classified into the label classification model to generate the corresponding hydrological status label.

[0011] Furthermore, the data acquisition module is specifically used for: Read the original observation sequence of each monitoring section within a time period from the hydrological database or real-time transmission system; The original observation sequence is cleaned and standardized to obtain a standard observation sequence; Based on hydrological operational rules, minority and majority class labels are defined, and the standard observation sequences are processed according to the minority and majority class labels to obtain an imbalanced dataset.

[0012] Furthermore, the sample optimization module is specifically used for: For each sample in the imbalanced dataset, calculate its Euclidean distance to other samples and construct a local overlap index for the samples. Use the local overlap index to traverse all minority class samples and divide the minority class samples into dangerous samples, safe samples, and noisy samples. Delete the noisy samples and retain dangerous samples and safe samples to obtain the cleaned minority class samples. By traversing all majority class samples using the local overlap index, majority class samples in which the number of minority class samples in the neighborhood exceeds the number of majority class samples are deleted to obtain purified majority class samples. The purified minority class samples and the purified majority class samples are merged to obtain optimized samples.

[0013] Furthermore, the model building module is specifically used for: Define the diffusion step number and the noise variance sequence, and perform forward noise addition on the initial minority class data according to the noise variance sequence and the diffusion step number to obtain noisy data; The noise-added data and the corresponding time step are input into a preset neural network model to obtain the corresponding predicted noise, and a corresponding noise loss function is constructed based on the predicted noise. A classification loss function is constructed for the neural network model. The classification loss function is then weighted and fused with the noise loss function to obtain a total loss function. The total loss function is then used to optimize the neural network model to obtain a diffusion probability model.

[0014] Furthermore, the classifier training module is specifically used for: Define the weighted Euclidean distance sum from the majority class sample to all minority class samples in the balanced dataset as the fitness function, and calculate the corresponding fitness value according to the fitness function; Based on the fitness value, the majority class samples to be retained are selected by the distance-weighted roulette wheel selection algorithm to obtain the majority class sample set. The majority class sample set is then fused with the minority class samples in the balanced dataset to form a sampled balanced dataset.

[0015] The present invention also proposes a storage medium on which a computer program is stored, which, when executed by a processor, implements the above-described method for automated classification and dynamic label generation of hydrological data.

[0016] The present invention also proposes a computer, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described method for automated classification and dynamic label generation of hydrological data.

[0017] Compared with existing technologies, the hydrological data automated classification and dynamic label generation method and system of this invention, by constructing a local overlap index, accurately identifies and processes noise in imbalanced datasets, thereby purifying training data and reducing fuzzy overlap between categories; it employs an improved diffusion probability model for oversampling, and through forward noise addition and reverse denoising processes, learns the true data distribution of minority class samples, generating high-quality, diverse synthetic minority class samples, effectively avoiding overfitting problems; it uses the processed balanced dataset as the source domain and combines it with the original labeled dataset as the target domain, using a transfer learning algorithm to train the base classifier, effectively utilizing the rich information of the source domain while adapting to the inherent data distribution of the target domain, overcoming model bias that may be caused by training on a single dataset; and iteratively training the final label classification model through an improved adaptive boosting algorithm, which can significantly improve the accuracy of identifying early warning events. Attached Figure Description

[0018] Figure 1 This is a flowchart of the method for automated classification and dynamic label generation of hydrological data in the first embodiment of the present invention; Figure 2 This is a structural block diagram of the hydrological data automated classification and dynamic label generation system in the second embodiment of the present invention; Figure 3 This is a structural block diagram of the computer in the third embodiment of the present invention.

[0019] The following detailed description, in conjunction with the accompanying drawings, will further illustrate the present invention. Detailed Implementation

[0020] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Several embodiments of the invention are illustrated in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete.

[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0022] Example 1 Please see Figure 1 The figure shows the method for automated classification and dynamic label generation of hydrological data in the first embodiment of the present invention, which specifically includes steps S101 to S105: S101, Collect hydrological monitoring data uploaded from various platforms, and construct an imbalanced dataset containing majority class samples and minority class samples based on the hydrological monitoring data; Furthermore, step S101 specifically includes steps S1011 to S1013: S1011, read the original observation sequence of each monitoring section within a time period from the hydrological database or real-time transmission system; S1012, perform data cleaning and standardization on the original observation sequence to obtain a standard observation sequence; S1013, define minority class labels and majority class labels based on hydrological business rules, and process the standard observation sequence according to the minority class labels and the majority class labels to obtain an imbalanced dataset.

[0023] In practical implementation, raw observation sequences of each monitoring section within a time period are collected from hydrological monitoring stations (hydrological database) or sensors (set in a real-time transmission system) at a fixed sampling frequency (in this embodiment, the sampling frequency is once every 15 minutes or hour). These raw observation sequences are presented in the form of a multi-dimensional time matrix.

[0024] In the formula, Indicates the first The timestamp of each record This is the water level value. For flow rate value, This represents the cumulative rainfall over a period of time. This is the flow rate value. This represents the total number of original observation records.

[0025] Furthermore, to eliminate outliers caused by sensor malfunctions and transmission errors, data cleaning is performed on each dynamic variable, employing a sliding median filter and threshold truncation method for any time series. ,in, Indicates the timestamp The original observed value of a certain hydrological variable (water level) at a certain location Flow value Cumulative rainfall over a period of time Flow rate value (one of them), if Then, linear interpolation is used for replacement, where, The median within the local window. The standard deviation of the sequence. Z-score standardization is performed on each of the cleaned variables using a preset multiplier (in this embodiment, the multiplier can be selected as 3) to give features of different dimensions a uniform scale, thus obtaining the corresponding time series. .

[0026] Based on hydrological industry standards and historical experience, warning thresholds are set for the aforementioned observations. A minority of categories are marked as alert states requiring attention, while the majority are marked as normal or safe states. The labels are generated based on whether the observed value (e.g., water level) exceeds the preset threshold.

[0027] In the formula, This represents the standardized water level value. This represents the standardized water level warning threshold. This is a minority category label, indicating a warning status. This indicates a majority category label, representing a normal or safe state.

[0028] Furthermore, to capture the temporal evolution of hydrological conditions and enhance the classifier's ability to perceive dynamic patterns, a sliding window technique is employed to extract features from historical sequences. The window length is set. and sliding step size (Recommended) That is, sliding point by point. Regarding time... Construct feature vectors Including current and previous Standardized variable values ​​at each time point:

[0029] The dimension of the feature vector is ,in, This refers to the number of dynamic variables used (such as water level, flow rate, and rainfall). This approach ensures that each sample contains a "snapshot" of a past period, providing the model with temporal context.

[0030] At the same time, the corresponding label within the window is taken from the original label of the window at the last moment, that is:

[0031] To ensure that the labels correspond to the latest state and are aligned with future prediction tasks, a sliding window is applied to all valid time points, resulting in a window of size [size missing]. Sample-label set:

[0032] in, This is the start time of the data.

[0033] Furthermore, the number of majority class samples is labeled as The number of minority class samples is labeled as Since the number of majority class samples is much larger than the number of minority class samples in hydrological early warning events, the ratio of the two is defined as the imbalance ratio. The larger this value, the more severe the class imbalance, resulting in an imbalanced dataset. .

[0034] S102, noise identification and data processing are performed on the imbalanced dataset. For each sample in the imbalanced dataset, the Euclidean distance between it and other samples is calculated, and a local overlap index of the samples is constructed. The local overlap index is used to traverse all minority class samples to obtain optimized samples. Furthermore, step S102 specifically includes steps S1021 to S1023: S1021, calculate the Euclidean distance between each sample and other samples in the imbalanced dataset, construct a local overlap index for the samples, traverse all minority class samples through the local overlap index, divide the minority class samples into dangerous samples, safe samples and noisy samples, delete the noisy samples, and retain dangerous samples and safe samples to obtain the cleaned minority class samples. S1022, by traversing all majority class samples through the local overlap index, delete the majority class samples in which the number of minority class samples in the neighborhood is greater than the number of majority class samples, so as to obtain the purified majority class samples. S1023, the purified minority class samples and the purified majority class samples are merged to obtain optimized samples.

[0035] In specific implementation, the imbalanced dataset... Each sample in Calculate its Euclidean distance from other samples to obtain the sample. of The nearest neighbor set .

[0036] Quantify the degree of class mixing in the region where the sample is located, and define the sample. Local overlap for:

[0037] In the formula, Represents the majority class sample set. Represents the minority class sample set. This represents the number of elements in the set, where... This means that all neighbors are in the majority class, and the sample is located in the core region of the majority class. This means that all neighbors are in the minority class, and the sample is located in the core region of the minority class. This means that the majority class is dominant in all neighboring classes; This means that a minority class is dominant among all its neighbors; This means that the two types of data are close in size among all neighbors and are in a highly overlapping area; Iterate through all minority class samples and perform a hierarchical operation based on the value of local overlap: like If all the neighbors of the minority class sample are of the majority class, it is completely surrounded by the majority class in the feature space and lacks connection with samples of the same class. Therefore, it is identified as a noise sample and is directly deleted.

[0038] like If the majority class has more neighbors than the minority class, but there are still minority class neighbors, the sample is in the class overlap region and is marked as a dangerous sample and retained.

[0039] like If the number of minority class samples in all their nearest neighbors exceeds or equals the number of majority class samples, the sample is located in a minority class-dense region and is marked as a safe sample, which is then retained, thus obtaining the purified minority class sample set. .

[0040] Furthermore, iterate through all majority class samples and calculate their local set overlap: like ,Right now This indicates that the minority class has more neighbors than the majority class. A value indicating that the majority class sample has penetrated deep into the minority class distribution or is on a highly overlapping decision boundary, potentially interfering with the classifier; therefore, this sample is directly deleted.

[0041] like If the majority class samples are located in the region where their own class dominates, they are retained, thus obtaining the purified majority class sample set. .

[0042] The purified minority class sample set and the purified majority class sample set The samples are then merged to obtain an optimized sample.

[0043] S103, Construct a diffusion probability model, and use the diffusion probability model to reconstruct the distribution and oversample the optimized samples to generate minority class samples that are balanced with the number of majority class samples, so as to form a balanced dataset; Furthermore, step S103 specifically includes steps S1031 to S1033: S1031, Define the diffusion step number and noise variance sequence, and perform forward noise addition processing on the initial minority class data according to the noise variance sequence and the diffusion step number to obtain noisy data; S1032, The noise-added data and the corresponding time step are input into a preset neural network model to obtain the corresponding predicted noise, and a corresponding noise loss function is constructed based on the predicted noise; S1033, construct the classification loss function of the neural network model, weight and fuse the classification loss function with the noise loss function to obtain the total loss function, and use the total loss function to optimize the neural network model to obtain the diffusion probability model.

[0044] In practical implementation, the number of diffusion steps is defined. and noise variance sequence ,satisfy ,and . Control the first The higher the noise intensity of the step data injection, the faster the single-step signal-to-noise ratio decays. In this embodiment, the scheduling method adopts a linear scheduling form.

[0045] Furthermore, the definition starts from... arrive The conditional probability distribution is used to inject a small amount of Gaussian noise into the data each time.

[0046] In the formula, Indicates a multivariate Gaussian distribution. Indicates the first Step-by-step noisy data, Indicates the first Step-by-step noisy data, This is a decay factor used to shrink the mean towards zero to maintain bounded variance. The covariance matrix of isotropic Gaussian noise. It represents an identity matrix where noise in each dimension is independent and has the same variance.

[0047] Specifically, from the purified minority class sample set Clean minority class samples were sampled from the middle. From the clean minority class samples, the Markov property is used. Get any step The conditional distribution is used to obtain noisy data. In the formula, Represents the signal component coefficients, as... Increase and decrease Represents the noise component coefficient, as... Increase and grow. This represents pure noise sampled from a standard Gaussian distribution, denoted as a vector with zero mean and covariance equal to the identity matrix. Gaussian distribution.

[0048] Building a neural network model (In this embodiment, the neural network model can be a BP neural network model or a multilayer perceptron), wherein, As model parameters, the noisy data and corresponding time steps obtained above are input into the neural network model for inverse denoising and neural network prediction to obtain the corresponding prediction noise. .

[0049] Furthermore, to ensure the authenticity of the generated samples and improve the ability of subsequent classifiers to identify minority classes, a cost-sensitive term regarding the classification performance of the generated samples is introduced into the standard diffusion loss. A corresponding noise loss function is constructed based on the predicted noise:

[0050] in, Indicates time step Clean minority class samples and noise Calculate the expected value from three random sources. Let represent the squared Euclidean norm of a vector.

[0051] This neural network model is used to generate a set of synthetic minority class samples from pure noise, which are then merged with majority class samples to train a temporary classifier. Thus, the corresponding classification loss function is constructed:

[0052] In the formula, This represents the AUC value of the temporary classifier, and its range is... ; This represents a monotonically increasing penalty function. In this embodiment, the penalty function is selected as a linear penalty function, which maps the insufficient AUC to the actual loss cost.

[0053] The two loss functions are weighted and fused to obtain the total loss function:

[0054] in, This is a balancing coefficient used to control classification performance.

[0055] The obtained total loss function is used to train and optimize the neural network model, using a purified minority class sample set. Using real data, iterative processing is performed until convergence to obtain a diffusion probability model. This diffusion probability model is then used to reconstruct the distribution of the optimized samples and perform oversampling to generate minority class samples that are in balance with the majority class samples, thus forming a balanced dataset. .

[0056] S104, the majority class samples are probabilistically downsampled using the distance-weighted roulette wheel selection algorithm to form a sampled balanced dataset. The sampled balanced dataset is used as the source domain, and the original hydrological dataset with labels is used as the target domain. A base classifier is trained using a transfer learning algorithm. Furthermore, step S104 specifically includes steps S1041 to S1042: S1041, Define the weighted Euclidean distance from the majority class sample to all minority class samples in the balanced dataset as the fitness function, and calculate the corresponding fitness value according to the fitness function; S1042, Based on the fitness value, the majority class samples to be retained are selected by the distance-weighted roulette wheel selection algorithm to obtain the majority class sample set. The majority class sample set is then fused with the minority class samples in the balanced dataset to form a sampled balanced dataset.

[0057] In practice, the weighted Euclidean distance from the majority class samples to all minority class samples in the balanced dataset is defined as the fitness function. The fitness function is used to calculate the corresponding fitness value. The fitness sum of each candidate sample is calculated based on the fitness, and the selection probability of each candidate sample is calculated. The cumulative probability interval is calculated to construct the roulette wheel selection probability. A random number uniformly distributed in the interval (0,1) is generated. The index that satisfies the cumulative probability interval is found. The sample corresponding to this index is the majority class sample selected this time. It is added to the balanced dataset and merged with its minority class samples to form a sampled balanced dataset.

[0058] Furthermore, the sampled balanced dataset is used as the source domain, and the labeled original hydrological dataset is used as the target domain. The source domain and the target domain are merged into a unified training set. The weight of each sample is initialized, and the majority and minority classes are assigned initial weights in their respective domains. Then, normalization is performed so that the sum of the weights of all samples is 1. For each iteration, a base classifier is trained using all samples under the current weight distribution. In this embodiment, the base classifier can be a weak learner such as a decision tree or a linear support vector machine. During training, the weight of each sample is considered, so that the classifier pays more attention to high-weight samples.

[0059] Specifically, considering only the target domain, calculate the weighted classification error rate of the base classifiers in the target domain:

[0060] In the formula, Indicates the first The classification error rate of each iteration on the target domain arrive This indicates the index range of the target domain samples in the merged dataset. Indicates the first In the first iteration The weights of each sample, This indicates that the base classifier is effective for the samples. Predicted labels, Indicates sample The true label; Introducing tiny positive numbers The weight update coefficients are defined as follows:

[0061] Different weight update strategies are adopted depending on the sample source domain: For source domain samples (index) ):

[0062] For target domain samples (index) ):

[0063] in, Indicates the first In the first iteration The weights of each sample.

[0064] The updated weights are normalized and processed. After several rounds of iteration, the base classifier corresponding to the round with the lowest error rate in the target domain is selected as the final base classifier.

[0065] S105, iteratively train the improved adaptive boosting algorithm and the base classifier to obtain a label classification model, and input the real-time hydrological data to be classified into the label classification model to generate the corresponding hydrological status label.

[0066] In specific implementation, the adaptive boosting algorithm (in this embodiment, the adaptive boosting algorithm is the AdaBoost algorithm) is optimized. Specifically, the error rate of the algorithm is optimized by introducing an AUC cost sensitivity factor into the error function for multiplication. The coefficient is 2x(1-AUC), where AUC is the AUC value of the current classifier. Furthermore, the original logarithmic function form is retained in the weight formula of the algorithm, and an improved error function and an exponential reward term are used. The independent variable of the reward term is the difference between the weight of the minority class samples correctly identified by the classifier and a preset value (in this embodiment, the value is 0.5), thereby obtaining the improved adaptive boosting algorithm. Furthermore, an improved adaptive boosting algorithm and base classifier are used for iterative training to obtain a label classification model. The real-time hydrological data to be classified is input into the label classification model. Standardized hydrological variable values ​​from several past time steps are extracted from the real-time data stream. A feature vector is generated using a sliding window feature construction method. This feature vector is input into the label classification model to obtain the minority class posterior probability. A conventional label probability threshold is set (in this embodiment, the threshold is set to 0.5). For generating conventional binary labels (minority and majority classes) at the current moment, since the evolution of hydrological conditions has continuity and inertia, the instantaneous high probability at a single point may be caused by sensor noise or sporadic fluctuations. Therefore, a dual-threshold rule based on a continuous time window is introduced, setting the window length and two probability thresholds (both probability thresholds are greater than the conventional label probability threshold). The first probability threshold is less than the second probability threshold. Two probability thresholds are used to calculate the real-time hydrological data to be classified. The final hydrological status label is generated based on the calculation results. When the probability value of each value in the real-time hydrological data to be classified is not lower than the first probability threshold (condition 1), and the average probability of the real-time hydrological data to be classified is not lower than the second probability threshold (condition 2), it means that the current and past consecutive moments have continuously shown a high-confidence warning feature, and the overall intensity is sufficient. In this case, the current moment label is generated or upgraded to "high-level alarm" (or flood alarm, high water level alarm, etc.). If the value of the conventional binary classification label is 1, and at least one of conditions 1 and 2 is not met, the current moment label is marked as "attention" (or "low-level warning", "caution", etc.), indicating that continuous monitoring is required but the highest level response is not triggered for the time being. If the value of the conventional binary classification label is 0, the current moment label is marked as "normal".

[0067] In summary, the automated hydrological data classification and dynamic label generation method in the above embodiments of the present invention, by constructing a local overlap index, accurately identifies and processes noise in imbalanced datasets, thereby purifying training data and reducing fuzzy overlap between categories; by employing an improved diffusion probability model for oversampling, and through forward noise addition and reverse denoising processes, it learns the true data distribution of minority class samples, generating high-quality and diverse synthetic minority class samples, effectively avoiding overfitting problems; by using the processed balanced dataset as the source domain and combining it with the original labeled dataset as the target domain, and using a transfer learning algorithm to train the base classifier, it can effectively utilize the rich information of the source domain while adapting to the inherent data distribution of the target domain, overcoming model bias that may be caused by training on a single dataset; and by iteratively training the improved adaptive boosting algorithm to obtain the final label classification model, it can significantly improve the accuracy of identifying early warning events.

[0068] Example 2 In another aspect, this invention proposes an automated classification and dynamic label generation system for hydrological data; please refer to [link / reference needed]. Figure 2 The figure shows an automated hydrological data classification and dynamic label generation system according to a second embodiment of the present invention. The system includes: The data acquisition module 11 is used to collect hydrological monitoring data uploaded from various platforms and construct an imbalanced dataset containing majority class samples and minority class samples based on the hydrological monitoring data. The sample optimization module 12 is used to perform noise identification and data processing on the imbalanced dataset. For each sample in the imbalanced dataset, it calculates the Euclidean distance between the sample and other samples, constructs a local overlap index of the samples, and uses the local overlap index to traverse all minority class samples to obtain optimized samples. The model building module 13 is used to build a diffusion probability model and use the diffusion probability model to reconstruct the distribution and oversample the optimized samples to generate minority class samples that are balanced with the number of majority class samples, so as to form a balanced dataset. The classifier training module 14 is used to perform probability downsampling on the majority class samples using a distance-weighted roulette wheel selection algorithm to form a sampled balanced dataset. The sampled balanced dataset is used as the source domain, and the original hydrological dataset with labels is used as the target domain. A base classifier is trained through a transfer learning algorithm. The label classification module 15 is used to iteratively train the improved adaptive boosting algorithm and the base classifier to obtain a label classification model, and input the real-time hydrological data to be classified into the label classification model to generate the corresponding hydrological status label.

[0069] Furthermore, the data acquisition module 11 is specifically used for: Read the original observation sequence of each monitoring section within a time period from the hydrological database or real-time transmission system; The original observation sequence is cleaned and standardized to obtain a standard observation sequence; Based on hydrological operational rules, minority and majority class labels are defined, and the standard observation sequences are processed according to the minority and majority class labels to obtain an imbalanced dataset.

[0070] Furthermore, the sample optimization module 12 is specifically used for: For each sample in the imbalanced dataset, calculate its Euclidean distance to other samples and construct a local overlap index for the samples. Use the local overlap index to traverse all minority class samples and divide the minority class samples into dangerous samples, safe samples, and noisy samples. Delete the noisy samples and retain dangerous samples and safe samples to obtain the cleaned minority class samples. By traversing all majority class samples using the local overlap index, majority class samples in which the number of minority class samples in the neighborhood exceeds the number of majority class samples are deleted to obtain purified majority class samples. The purified minority class samples and the purified majority class samples are merged to obtain optimized samples.

[0071] Furthermore, the model building module 13 is specifically used for: Define the diffusion step number and the noise variance sequence, and perform forward noise addition on the initial minority class data according to the noise variance sequence and the diffusion step number to obtain noisy data; The noise-added data and the corresponding time step are input into a preset neural network model to obtain the corresponding predicted noise, and a corresponding noise loss function is constructed based on the predicted noise. A classification loss function is constructed for the neural network model. The classification loss function is then weighted and fused with the noise loss function to obtain a total loss function. The total loss function is then used to optimize the neural network model to obtain a diffusion probability model.

[0072] Furthermore, the classifier training module 14 is specifically used for: Define the weighted Euclidean distance sum from the majority class sample to all minority class samples in the balanced dataset as the fitness function, and calculate the corresponding fitness value according to the fitness function; Based on the fitness value, the majority class samples to be retained are selected by the distance-weighted roulette wheel selection algorithm to obtain the majority class sample set. The majority class sample set is then fused with the minority class samples in the balanced dataset to form a sampled balanced dataset.

[0073] The functions or operation steps implemented by the above modules and units are largely the same as those in the above method embodiments, and will not be repeated here.

[0074] The hydrological data automated classification and dynamic label generation system provided in this embodiment of the invention has the same implementation principle and technical effects as the aforementioned method embodiment. For the sake of brevity, any parts not mentioned in the system embodiment can be referred to the corresponding content in the aforementioned method embodiment.

[0075] Example 3 This invention also proposes a computer, please refer to [link / reference]. Figure 3 The computer shown in the third embodiment of the present invention includes a memory 10, a processor 20, and a computer program 30 stored in the memory 10 and executable on the processor 20. When the processor 20 executes the computer program 30, it implements the above-mentioned method for automated classification and dynamic label generation of hydrological data.

[0076] The memory 10 includes at least one type of storage medium, such as flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 10 can be an internal storage unit of a computer, such as the computer's hard disk. In other embodiments, the memory 10 can be an external storage device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Furthermore, the memory 10 can include both internal and external storage units of the computer. The memory 10 can be used not only to store application software and various types of data installed on the computer, but also to temporarily store data that has been output or will be output.

[0077] In some embodiments, the processor 20 may be an electronic control unit (ECU), a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip, used to run program code stored in the memory 10 or process data, such as executing access restriction programs.

[0078] It should be pointed out that, Figure 3 The structure shown does not constitute a limitation on the computer. In other embodiments, the computer may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0079] This invention also proposes a storage medium storing a computer program that, when executed by a processor, implements the above-described method for automated classification and dynamic label generation of hydrological data.

[0080] Those skilled in the art will understand that the logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can mean any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0081] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.

[0082] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0083] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0084] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A method for automated classification and dynamic label generation of hydrological data, characterized in that, include: Collect hydrological monitoring data uploaded from various platforms, and construct an imbalanced dataset containing majority class and minority class samples based on the hydrological monitoring data; The imbalanced dataset is subjected to noise identification and data processing. For each sample in the imbalanced dataset, its Euclidean distance with other samples is calculated, and a local overlap index of the samples is constructed. The local overlap index is used to traverse all minority class samples to obtain optimized samples. A diffusion probability model is constructed, and the distribution of the optimized samples is reconstructed and oversampled using the diffusion probability model to generate minority class samples that are in balance with the number of majority class samples, so as to form a balanced dataset. The majority class samples are probabilistically downsampled using a distance-weighted roulette wheel selection algorithm to form a sampled balanced dataset. The sampled balanced dataset is used as the source domain, and the original labeled hydrological dataset is used as the target domain. A base classifier is trained using a transfer learning algorithm. The improved adaptive boosting algorithm and the base classifier are iteratively trained to obtain a label classification model. The real-time hydrological data to be classified is then input into the label classification model to generate the corresponding hydrological status label. The steps include: noise identification and data processing of the imbalanced dataset; calculating the Euclidean distance between each sample and other samples in the imbalanced dataset; constructing a local overlap index for the samples; and using the local overlap index to traverse all minority class samples to obtain optimized samples. For each sample in the imbalanced dataset, calculate its Euclidean distance to other samples and construct a local overlap index for the samples. Use the local overlap index to traverse all minority class samples and divide the minority class samples into dangerous samples, safe samples, and noisy samples. Delete the noisy samples and retain dangerous samples and safe samples to obtain the cleaned minority class samples. By traversing all majority class samples using the local overlap index, majority class samples in which the number of minority class samples in the neighborhood exceeds the number of majority class samples are deleted to obtain purified majority class samples. The purified minority class samples and the purified majority class samples are merged to obtain optimized samples; The steps involved in constructing the diffusion probability model include: Define the diffusion step number and the noise variance sequence, and perform forward noise addition on the initial minority class data according to the noise variance sequence and the diffusion step number to obtain noisy data; The noise-added data and the corresponding time step are input into a preset neural network model to obtain the corresponding predicted noise, and a corresponding noise loss function is constructed based on the predicted noise. A classification loss function for the neural network model is constructed, and the classification loss function is weighted and fused with the noise loss function to obtain a total loss function. The total loss function is then used to optimize the neural network model to obtain a diffusion probability model. The step of using a distance-weighted roulette wheel selection algorithm to probabilistically downsample the majority class samples to form a balanced sampled dataset includes: Define the weighted Euclidean distance sum from the majority class sample to all minority class samples in the balanced dataset as the fitness function, and calculate the corresponding fitness value according to the fitness function; Based on the fitness value, the majority class samples to be retained are selected by the distance-weighted roulette wheel selection algorithm to obtain the majority class sample set. The majority class sample set is then fused with the minority class samples in the balanced dataset to form a sampled balanced dataset.

2. The method for automated classification and dynamic label generation of hydrological data according to claim 1, characterized in that, The steps of collecting hydrological monitoring data uploaded from various platforms and constructing an imbalanced dataset containing majority and minority class samples based on the hydrological monitoring data include: Read the original observation sequence of each monitoring section within a time period from the hydrological database or real-time transmission system; The original observation sequence is cleaned and standardized to obtain a standard observation sequence; Based on hydrological operational rules, minority and majority class labels are defined, and the standard observation sequences are processed according to the minority and majority class labels to obtain an imbalanced dataset.

3. A system for automated classification and dynamic label generation of hydrological data, characterized in that, include: The data acquisition module is used to collect hydrological monitoring data uploaded from various platforms and construct an imbalanced dataset containing majority class and minority class samples based on the hydrological monitoring data. The sample optimization module is used to identify noise and process data in the imbalanced dataset. For each sample in the imbalanced dataset, it calculates the Euclidean distance between it and other samples, constructs a local overlap index of the samples, and uses the local overlap index to traverse all minority class samples to obtain optimized samples. The model building module is used to build a diffusion probability model and use the diffusion probability model to reconstruct the distribution and oversample the optimized samples to generate minority class samples that are balanced with the number of majority class samples, so as to form a balanced dataset. The classifier training module is used to perform probability downsampling on the majority class samples using a distance-weighted roulette wheel selection algorithm to form a sampled balanced dataset. The sampled balanced dataset is used as the source domain, and the original hydrological dataset with labels is used as the target domain. A base classifier is trained through a transfer learning algorithm. The label classification module is used to iteratively train the improved adaptive boosting algorithm and the base classifier to obtain the label classification model, and input the real-time hydrological data to be classified into the label classification model to generate the corresponding hydrological status label. Specifically, the sample optimization module is used for: For each sample in the imbalanced dataset, calculate its Euclidean distance to other samples and construct a local overlap index for the samples. Use the local overlap index to traverse all minority class samples and divide the minority class samples into dangerous samples, safe samples, and noisy samples. Delete the noisy samples and retain dangerous samples and safe samples to obtain the cleaned minority class samples. By traversing all majority class samples using the local overlap index, majority class samples in which the number of minority class samples in the neighborhood exceeds the number of majority class samples are deleted to obtain purified majority class samples. The purified minority class samples and the purified majority class samples are merged to obtain optimized samples; Specifically, the model building module is used for: Define the diffusion step number and the noise variance sequence, and perform forward noise addition on the initial minority class data according to the noise variance sequence and the diffusion step number to obtain noisy data; The noise-added data and the corresponding time step are input into a preset neural network model to obtain the corresponding predicted noise, and a corresponding noise loss function is constructed based on the predicted noise. A classification loss function for the neural network model is constructed, and the classification loss function is weighted and fused with the noise loss function to obtain a total loss function. The total loss function is then used to optimize the neural network model to obtain a diffusion probability model. Specifically, the classifier training module is used for: Define the weighted Euclidean distance sum from the majority class sample to all minority class samples in the balanced dataset as the fitness function, and calculate the corresponding fitness value according to the fitness function; Based on the fitness value, the majority class samples to be retained are selected by the distance-weighted roulette wheel selection algorithm to obtain the majority class sample set. The majority class sample set is then fused with the minority class samples in the balanced dataset to form a sampled balanced dataset.

4. The hydrological data automated classification and dynamic label generation system according to claim 3, characterized in that, The data acquisition module is specifically used for: Read the original observation sequence of each monitoring section within a time period from the hydrological database or real-time transmission system; The original observation sequence is cleaned and standardized to obtain a standard observation sequence; Based on hydrological operational rules, minority and majority class labels are defined, and the standard observation sequences are processed according to the minority and majority class labels to obtain an imbalanced dataset.

5. A readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements the method for automated classification and dynamic label generation of hydrological data as described in any one of claims 1 to 2.

6. A computer comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method for automated classification and dynamic label generation of hydrological data as described in any one of claims 1 to 2.

Citation Information

Patent Citations

  • Electric energy meter fault classification method and system based on confrontation comparison

    CN119557756A

  • Oversampling method, system and equipment for test data of avionics equipment and medium

    CN120429636A