Training data enhancement method, device and storage medium
By dividing the drainage network dataset into ordered and disordered datasets and performing linear interpolation and Gaussian noise processing, balanced data is generated, which solves the problem of low accuracy in drainage network regulation caused by unbalanced data distribution and achieves more accurate and timely regulation effects.
Patent Information
- Application Number
- CN202511022583.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-07-24
AI Technical Summary
The low accuracy of drainage network regulation is caused by the imbalance of data distribution. Traditional prediction models find it difficult to fully learn the drainage characteristics under abnormal conditions, which affects the response inaccurate and timely.
By obtaining historical monitoring data of drainage network nodes, the dataset is divided into ordered and unordered datasets based on the target balancing parameters, and linear interpolation and Gaussian noise are added to generate balanced enhanced data to improve the balance and representativeness of the dataset.
It improves the accuracy and reliability of drainage network regulation, ensures timely response in extreme events, and improves the effectiveness of data processing.
Smart Images

Figure CN120525085B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a method, device, and storage medium for enhancing training data. Background Art
[0002] The operation of drainage networks is affected by a variety of factors, such as rainfall and fluctuations in sewage discharge. Therefore, timely and accurate regulation is required to ensure drainage quality. Current drainage network regulation methods suffer from an imbalanced data distribution during the data processing phase. Monitoring data in drainage networks is typically high in normal operation, while during extreme events such as heavy rain and pipe blockages, data on abnormal events is relatively scarce. This imbalance in data distribution makes it difficult for traditional prediction models to fully learn the drainage characteristics under abnormal conditions during training. Consequently, in actual prediction and regulation, responses to abnormal events are not accurate or timely, impacting the accuracy of drainage network regulation.
[0003] The above content is only used to assist in understanding the technical solution of this application and does not constitute an admission that the above content is prior art. Summary of the Invention
[0004] The main purpose of this application is to provide a training data enhancement method, device and storage medium, aiming to solve the technical problem of low accuracy in drainage network control caused by unbalanced data distribution.
[0005] To achieve the above objectives, the present application proposes a method for enhancing training data, the method comprising:
[0006] Obtaining historical monitoring data obtained from each node of the drainage network, taking the historical monitoring data corresponding to each node as a data point to obtain an original data set;
[0007] Based on a target first equalization parameter, dividing the original data set into an ordered data set and an unordered data set according to a proximity distance between each data point in the original data set and a neighboring data point corresponding to the target first equalization parameter;
[0008] Linear interpolation is performed on the ordered data set, and Gaussian noise is added to each data point in the unordered data set based on a target second equalization parameter to obtain enhanced data after equalization.
[0009] In one embodiment, before the step of dividing the original data set into an ordered data set and an unordered data set based on the target first equalization parameter and according to the proximity distance between each data point in the original data set and a neighboring data point corresponding to the target first equalization parameter, the step further includes:
[0010] traversing each candidate first equalization parameter, determining a proximity distance between each data point and a neighboring data point in the original data set, and dividing the original data set into a plurality of groups of ordered data points and unordered data points based on the proximity distance between each data point and the neighboring data point in the original data set, wherein the first equalization parameter represents the number of the neighboring data points;
[0011] For each group of ordered data points and disordered data points, linearly interpolate the ordered data points to obtain new ordered data points, and traverse each candidate second equalization parameter to add Gaussian noise to the disordered data points to generate new disordered data points;
[0012] Each new group of the ordered data points and the unordered data points is used as a new original data set. According to the density distribution of each data point in the new original data set, the target first equalization parameter is determined from the candidate first equalization parameters, and the target second equalization parameter is determined from the candidate second equalization parameters.
[0013] In one embodiment, after the steps of performing linear interpolation on the ordered data set and adding Gaussian noise to each data point in the unordered data set based on the target second equalization parameter to obtain equalized enhanced data, the method further includes:
[0014] After training the prediction model according to the enhanced data, determining the prediction quality result based on the current monitoring data obtained from each node of the drainage network by the prediction model;
[0015] A control plan for the drainage network is determined based on the deviation between the predicted quality result and the target quality result.
[0016] In one embodiment, the step of training the prediction model based on the enhanced data includes:
[0017] Training a preset model based on the enhanced data to obtain an initial prediction model;
[0018] After equalizing the disordered data set in the enhanced data according to the target second equalization parameter, the initial prediction model is trained again based on the equalized enhanced data to obtain the prediction model.
[0019] In one embodiment, the step of determining a control scheme for the drainage network based on the deviation between the predicted quality result and the target quality result includes:
[0020] Determining a predicted value corresponding to a target quality indicator from the predicted quality result, and determining a target value corresponding to the target quality indicator from the target quality result;
[0021] Determining the contribution score of each characteristic indicator of each data point in the current monitoring data obtained from each node of the drainage network to the target quality indicator;
[0022] Determining a target control indicator from the characteristic indicators of each data point in the current monitoring data according to the contribution score;
[0023] Determining the control amount of the target control indicator based on the deviation between the predicted value and the target value corresponding to the target quality indicator;
[0024] A control scheme for the drainage network is determined based on the target control index and the control amount of the target control index.
[0025] In one embodiment, the step of determining the control amount of the target control indicator based on the deviation between the predicted value and the target value corresponding to the target quality indicator includes:
[0026] Based on the current monitoring data, the target control indicator is simulated and regulated by the prediction model according to the preset control value corresponding to each target control indicator, so as to obtain the value of the target quality indicator after the simulated control;
[0027] Determine the reference control value of each target control indicator when the value of the target quality indicator after the simulated control reaches an extreme value;
[0028] Determining a control coefficient for each target control indicator based on a deviation between a detection value of each target control indicator in the current monitoring data and the reference control value;
[0029] The control amount of each target control indicator is determined according to the control coefficient of each target control indicator and the deviation between the predicted value and the target value corresponding to the target quality indicator.
[0030] In one embodiment, after the step of determining the control amount of each target control indicator based on the control coefficient of each target control indicator and the deviation between the predicted value and the target value corresponding to the target quality indicator, the step further includes:
[0031] Dividing the drainage pipe network into pipe segments according to each node of the drainage pipe network;
[0032] Determining a time delay parameter of an upstream node corresponding to the pipeline segment according to the length of the pipeline segment and the drainage flow rate;
[0033] According to the delay parameter, a control component of the target control indicator corresponding to each of the nodes is determined.
[0034] In one embodiment, the step of determining the control component of the target control indicator corresponding to each node according to the delay parameter includes:
[0035] Determining the weight of the current node according to the distance between the current node and the end node of the drainage network;
[0036] The control component of the target control indicator corresponding to the current node is determined according to the weight and the delay parameter of the current node and the control amount of the target control indicator.
[0037] In addition, to achieve the above-mentioned purpose, the present application also proposes a training data enhancement device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the computer program is configured to implement the steps of the training data enhancement method as described above.
[0038] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the training data enhancement method described above are implemented.
[0039] The present application provides a method for enhancing training data, which obtains historical monitoring data obtained from each node of the drainage network, takes the historical monitoring data corresponding to each node as a data point, obtains an original data set, and divides the original data set into an ordered data set and an unordered data set based on the target first balance parameter according to the proximity distance between each data point in the original data set and the adjacent data point corresponding to the target first balance parameter. Then, linear interpolation is performed on the ordered data set, and Gaussian noise is added to each data point in the unordered data set based on the target second balance parameter to obtain enhanced data after balance. This method can insert new data points between ordered data points through linear interpolation, thereby maintaining the continuity and trend of the ordered data, and at the same time, adds Gaussian noise to each data point in the unordered data set to increase the diversity and complexity of the data, so as to improve the balance and representativeness of the data set, thereby solving the technical problem of low accuracy in drainage network regulation caused by unbalanced data distribution. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0041] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0042] Figure 1 A flowchart of the first embodiment of the method for enhancing training data of this application is provided;
[0043] Figure 2 A flowchart of the third embodiment of the method for enhancing training data of this application is provided;
[0044] Figure 3 Schematic diagram of the device structure of the hardware operating environment involved in the training data enhancement method in the embodiment of the present application.
[0045] The purpose, features and advantages of this application will be further explained with reference to the accompanying drawings in conjunction with the embodiments. DETAILED DESCRIPTION
[0046] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not intended to limit the present application.
[0047] In order to better understand the technical solution of this application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0048] The operation of drainage networks is affected by a variety of factors, such as rainfall and fluctuations in sewage discharge. Therefore, timely and accurate regulation is required to ensure drainage quality. Current drainage network regulation methods suffer from an imbalanced data distribution during the data processing phase. Monitoring data in drainage networks is typically high in normal operation, while during extreme events such as heavy rain and pipe blockages, data on abnormal events is relatively scarce. This imbalance in data distribution makes it difficult for traditional prediction models to fully learn the drainage characteristics under abnormal conditions during training. Consequently, in actual prediction and regulation, responses to abnormal events are not accurate or timely, impacting the accuracy of drainage network regulation.
[0049] In view of the above problems, the present application proposes a method for enhancing training data, which obtains historical monitoring data obtained from each node of the drainage network, takes the historical monitoring data corresponding to each node as a data point, and obtains the original data set. Based on the target first balance parameter, the original data set is divided into an ordered data set and an unordered data set according to the proximity distance between each data point in the original data set and the adjacent data point corresponding to the target first balance parameter. Then, linear interpolation is performed on the ordered data set, and Gaussian noise is added to each data point in the unordered data set based on the target second balance parameter to obtain enhanced data after balance. This method can insert new data points between ordered data points through linear interpolation, thereby maintaining the continuity and trend of the ordered data. At the same time, Gaussian noise is added to each data point in the unordered data set to increase the diversity and complexity of the data, so as to improve the balance and representativeness of the data set, and solve the technical problem of low accuracy in drainage network regulation caused by unbalanced data distribution.
[0050] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication and program running functions, such as a tablet computer, personal computer, etc., or an electronic device that can realize the above functions, a drainage network control system, etc.
[0051] Based on this, the first embodiment proposed in this application provides a method for enhancing training data, referring to Figure 1 In this embodiment, the training data enhancement method includes steps S10 to S40:
[0052] Step S10: acquiring historical monitoring data acquired from each node of the drainage network, taking the historical monitoring data corresponding to each node as a data point, and obtaining an original data set.
[0053] Historical monitoring data refers to data collected from various nodes in the drainage network over a period of time, including but not limited to parameters such as flow rate, flow velocity, liquid level, hydrogen sulfide concentration, and pH value. The historical monitoring data set corresponding to each node serves as an independent data point for subsequent data analysis and model training. Drainage network nodes can include inspection wells, pumping stations, network intersections, or other key connection points. Each data point contains at least one characteristic indicator that reflects drainage quality, such as flow velocity or pH value.
[0054] Step S20 : Based on the target first equalization parameter, the original data set is divided into an ordered data set and an unordered data set according to the proximity distance between each data point in the original data set and a neighboring data point corresponding to the target first equalization parameter.
[0055] Understandably, when regulating drainage networks, drainage quality is typically predicted based on monitoring data, and regulation plans are then formulated based on the predicted drainage quality. However, this regulation process often faces challenges with small data volumes and uneven data distribution. Especially during extreme events, the ratio of normal operating data to abnormal event data is severely unbalanced, making it difficult for traditional prediction models to effectively process data, impacting the accuracy and reliability of predictions and regulation. Therefore, raw data preprocessing is necessary to improve data quality and model performance.
[0056] Optionally, the density distribution of each data point in the original dataset is calculated and, based on this density distribution, the original dataset is divided into an ordered dataset and an unordered dataset. New sample data points are generated within the ordered dataset to maintain data continuity and trend. New sample data points are also generated within the unordered dataset to expand coverage of low-density areas. The ordered and unordered datasets, after generating the new sample data points, are combined to obtain balanced enhanced data.
[0057] An ordered dataset refers to a dataset with a relatively stable data density trend. Within these regions, the data points are relatively evenly distributed. An unordered dataset refers to a dataset with a sparse data density, large fluctuations, or abnormal distribution.
[0058] Optionally, the number of adjacent data points corresponding to each data point is determined based on the target first equalization parameter, and then the proximity distance between each data point and these adjacent data points is determined. , calculate its The distance between neighboring data points The proximity distance can be calculated using the following formula:
[0059]
[0060] Among them, M is the total number of characteristic indicators, and The data points and The value of the characteristic index m in , the characteristic index m can be oxygen concentration, flow rate, pH value, etc.
[0061] After determining the proximity distance between each data point and its neighboring data points, the original data set is divided into an ordered data set and an unordered data set according to the proximity distance.
[0062] For example, an "orderliness" metric is determined for each data point based on its proximity to neighboring data points. This metric can be based on the average distance, minimum distance, or maximum distance between the data point and its neighbors. A density threshold is then set, and data points with an "orderliness" metric above the threshold are classified as an ordered dataset, while data points with an "orderliness" metric below the threshold are classified as an unordered dataset. Alternatively, clustering methods are used to group data points with close proximity into one category, forming an ordered dataset, while data points that are isolated or far away from other data points are classified as an unordered dataset.
[0063] For example, suppose the average distance between each data point and its neighbors is used as an indicator of "orderliness":
[0064]
[0065] Setting density thresholds ,Will The data points are divided into ordered data sets, The data points are divided into unordered data sets. Density threshold It can be determined by statistical analysis or experience, for example, taking all of the median.
[0066] Step S30 , performing linear interpolation on the ordered data set, and adding Gaussian noise to each data point in the unordered data set based on a target second equalization parameter to obtain enhanced data after equalization.
[0067] After the original data set is divided into an ordered data set and an unordered data set, linear interpolation is performed on the ordered data set to generate new sample data points. Linear interpolation can insert new data points between adjacent data points to maintain the continuity and trend of the data. For example, for the data points in the ordered data set, Perform linear interpolation to insert new sample data points between every two adjacent data points .
[0068] For the disordered dataset in the original dataset, new sample data points can be generated by adding Gaussian noise. The amount of Gaussian noise added is controlled by the target second equalization parameter, which represents the standard deviation of the Gaussian noise. For example, for each data point in the disordered dataset , according to the candidate second equalization parameter , generate a Gaussian distribution N(0, ) and add it to the disordered data set to obtain new sample data points , is the identity matrix.
[0069] Optionally, local density adaptive noise can be introduced into the disordered data set in the original data set. The local standard deviation As the second target equalization parameter .in, , k is the first equilibrium parameter, that is, the data point The number of neighboring data points, Represents a data point The set of k neighboring data points, Represents data points The jth neighboring data point of Represents data points Its neighboring data points The Euclidean distance between them.
[0070] Finally, the ordered dataset after linear interpolation and the unordered dataset after adding Gaussian noise are combined to form the enhanced dataset after equalization. This embodiment effectively addresses the issues of small sample sizes and long-tail distributions while maintaining physical consistency of the data by combining structured preservation with controllable perturbations. This provides a reliable data foundation for drainage network regulation and improves its accuracy.
[0071] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the first embodiment can be referred to the above introduction and will not be repeated hereafter. On this basis, before step S20, the training data enhancement method further includes steps S21 to S23:
[0072] Step S21, traverse each candidate first equalization parameter, determine the proximity distance between each data point in the original data set and the adjacent data points, and divide the original data set into multiple groups of ordered data points and unordered data points based on the proximity distance between each data point in the original data set and the adjacent data points, wherein the first equalization parameter represents the number of the adjacent data points.
[0073] For example, assuming there are 5 candidate first equalization parameters, traverse each candidate first equalization parameter For each data point in the original dataset, use distance metrics such as Euclidean distance, Manhattan distance, etc. to calculate its distance to The distance between adjacent data points. Assuming that the average distance between each data point and its adjacent data points is used as the "orderliness" indicator, the original data set is divided into 5 groups of ordered data points and disordered data points.
[0074] Step S22: for each group of ordered data points and disordered data points, linear interpolation is performed on the ordered data points to obtain new ordered data points, and each candidate second equalization parameter is traversed to add Gaussian noise to the disordered data points to generate new disordered data points.
[0075] For example, assuming that according to a candidate first equalization parameter The original data set is divided into a set of ordered data points [A1, A2] and disordered data points [B1, B2]. Linear interpolation of ordered data points A1 and A2 to obtain A3 can obtain new ordered data points [A1, A2, A3]. Then, for the disordered data points, each candidate second equalization parameter is traversed, assuming that . According to , and Adding Gaussian noise to the disordered data points B1 and B2 can produce new disordered data points: [B1+ , B2+ ]、[B1+ , B2+ ] and [B1+ , B2+ ]. 、 and The standard deviations are Gaussian noise.
[0076] Step S23, taking each new group of the ordered data points and the unordered data points as the new original data set, and determining the target first equalization parameter from the candidate first equalization parameters according to the density distribution of each data point in the new original data set, and determining the target second equalization parameter from the candidate second equalization parameters.
[0077] Taking the above example as an example, after balancing a set of ordered data points [A1, A2] and disordered data points [B1, B2] according to the candidate first balancing parameters and the second balancing parameters, three new original data sets [A1, A2, A3, B1+ , B2+ ]、[A1,A2,A3,B1+ , B2+ ] and [A1, A2, A3, B1+ , B2+ Then, each set of candidate equalization parameters is evaluated based on the density distribution of each data point in the new original data set ( , )、( , )and( , ) effect. The density distribution of data points can be obtained by calculating the local density of the data points, such as using kernel density estimation or a clustering algorithm. Based on the evaluation results, the set of balancing parameters with the best balancing effect is used as the target first and second balancing parameters. Evaluation criteria may include the balance of data distribution, the prediction accuracy of the model, and the generalization ability.
[0078] Optionally, to ensure that the enhanced data after equalization of the equalization parameters can improve the prediction performance in the actual model, a grid search and a cross-validation method are used to jointly determine the target first equalization parameter and the target second equalization parameter.
[0079] For example, for each set of candidate equalization parameters (k, ) perform K-fold cross validation to train the prediction model. K-fold cross validation is a commonly used model evaluation method that divides the dataset into K subsets, then uses K-1 subsets for training and the remaining 1 subset for validation to evaluate the performance of the model. Then, by maximizing the objective function To find the target equilibrium parameters ( , ). The objective function J is related to the performance indicators of the model, such as accuracy, recall rate, F1 value, etc. In this embodiment, Indicates the use of parameter combinations (k, ) in augmented data When the objective function reaches its maximum value, it means that the corresponding equilibrium parameters can make the density distribution of each data point reach the optimal equilibrium state.
[0080] In this embodiment, the data is divided and enhanced multiple times by traversing different first equalization parameters and second equalization parameters, and then the optimal parameter combination is selected through cross-validation to find the best equalization parameters, so that the enhanced data performs best in model training, which can improve the enhancement effect of the training data.
[0081] Based on the above embodiments of the present application, in the third embodiment of the present application, the same or similar contents as those in the above embodiments can be referred to the above introduction and will not be described in detail later. Figure 2 After step S30, the training data enhancement method further includes steps S40 to S50:
[0082] Step S40: After the prediction model is trained according to the enhanced data, a prediction quality result is determined by the prediction model based on the current monitoring data obtained from each node of the drainage network.
[0083] Optionally, a training set and a validation set are divided from the enhanced data. The training set is used for model learning, and the validation set is used to adjust model parameters and evaluate model performance. Select a suitable prediction model based on the characteristics and prediction objectives of the drainage network monitoring data. For example, you can choose an integrated learning algorithm such as XGBoost or LightGBM, which performs well in processing nonlinear relationships and high-dimensional features. Specifically, initialize the model parameters first, including the learning rate, the maximum depth of the tree, the number of leaf nodes, etc. The training set data is input into the model, and the model begins to learn the patterns and regularities in the data. During the training process, the model updates the model parameters through continuous iteration to minimize the difference between the predicted value and the true value, such as the mean square error. Then, use the validation set to verify the model and evaluate the performance of the model. Based on the verification results, adjust and optimize the model's hyperparameters. For example, you can use methods such as grid search to find the hyperparameter combination that optimizes model performance.
[0084] For example, using the balanced training data set Train the model, where is the input feature vector, i.e. the augmented data used for training, is the prediction target, N is the number of data points in the training data set. The training goal is to minimize the mean square error :
[0085]
[0086] in, The predicted value output by the model.
[0087] After the model training is completed, the performance of the model is evaluated by K-fold cross validation. For example, the coefficient of determination is used. And the root mean square error RMSE measures the model performance:
[0088] ,
[0089] choose The model with the highest value and the lowest RMSE is selected as the final prediction model.
[0090] After obtaining a trained prediction model, the current monitoring data collected from each node in the drainage network is preprocessed, including data cleaning and format conversion, to ensure that the data meets the model's input requirements. This preprocessed data is then fed into the trained prediction model. Based on the learned patterns and rules, the model predicts the input data and outputs a prediction quality result. This prediction quality result includes the predicted value of each characteristic indicator at each data point.
[0091] Step S50: determining a control scheme for the drainage network according to the deviation between the predicted quality result and the target quality result.
[0092] It should be noted that the target quality result includes the target value of each characteristic indicator. For example, a nonlinear relationship model is established between the input of multiple characteristic indicators and the target quality indicator, such as hydrogen sulfide concentration, to capture the complex interactions between the parameters. The contribution of each other characteristic indicator to the predicted value of the target quality indicator is calculated, and the characteristic indicators that need to be regulated are identified. Based on the nonlinear relationship model and the deviation quantification results, the corresponding regulation strategy is recommended. For example, if the prediction results show that the hydrogen sulfide concentration in the drainage pipe is too high compared to the target value, the oxygen supply can be increased to increase the oxygen concentration and reduce the hydrogen sulfide concentration.
[0093] In this embodiment, the original data set is constructed by collecting multi-parameter node data, and the balanced parameters are used to divide the ordered and disordered data sets and perform interpolation and other processing to improve the data quality; an integrated learning algorithm is used in combination with cross-validation to train the prediction model to accurately predict the drainage quality; based on the prediction deviation, the contribution of each characteristic indicator and the nonlinear relationship are combined to recommend the control strategy to achieve precise control, effectively solve the problem of data distribution imbalance, improve the accuracy and reliability of prediction and control, and ensure the stable operation of the drainage network.
[0094] Based on the above embodiments of the present application, in the fourth embodiment of the present application, the same or similar contents as those in the above embodiments can be referred to the above introduction and will not be repeated hereafter. On this basis, step S40 includes steps S41~S42:
[0095] Step S41: training a preset model based on the enhanced data to obtain an initial prediction model.
[0096] Step S42 : After equalizing the disordered data set in the enhanced data according to the target second equalization parameter, the initial prediction model is trained again based on the equalized enhanced data to obtain the prediction model.
[0097] For example, a pre-set model that excels at handling nonlinear relationships and high-dimensional features and is suitable for drainage network monitoring data, such as LightGBM and XGBoost, is selected. The balanced augmented data is divided into a training set and a validation set, and the pre-set model is trained and optimized to obtain an initial prediction model.
[0098] The initial prediction model can initially have the ability to predict key parameters of the drainage network such as hydrogen sulfide concentration, flow rate, pH value, etc., providing a basis for subsequent model optimization and application. After the first round of training, the disordered data set in the enhanced data is balanced again according to the target second balance parameter. For each data point in the disordered data set, a Gaussian distribution N(0, ) and adds this noise to the unordered dataset of augmented data to further balance it. This new augmented data contains more sample data points, helping the model better learn the distribution characteristics of the data. The initial prediction model is retrained using this new augmented data. During retraining, the model structure can remain unchanged, with only the model parameters adjusted to further optimize the model's predictive performance. The model structure can also be fine-tuned as needed.
[0099] It's understandable that augmented data includes more diverse samples, especially those in low-density areas. Using augmented data for training allows the model to better learn the data's distribution characteristics, improving its generalization capabilities. The resulting prediction model, further trained, possesses even stronger predictive capabilities, enabling more accurate predictions of various drainage network quality indicators and the development of more precise control plans.
[0100] Based on the above embodiments of the present application, in the fifth embodiment of the present application, the same or similar contents as those in the above embodiments can be referred to the above introduction and will not be repeated hereafter. On this basis, step S50 includes steps S51 to S55:
[0101] Step S51 : determining a predicted value corresponding to a target quality indicator from the predicted quality result, and determining a target value corresponding to the target quality indicator from the target quality result.
[0102] From the predicted quality results output by the prediction model, identify and extract the predicted value corresponding to the target quality indicator. For example, if the target quality indicator is hydrogen sulfide concentration, obtain the predicted value of hydrogen sulfide concentration from the predicted quality results. From the preset target quality results, find the target value corresponding to the target quality indicator. This target value is typically a desired quality standard to be achieved or maintained, such as a safety threshold for hydrogen sulfide concentration.
[0103] Step S52: determining the contribution score of each characteristic indicator of each data point in the current monitoring data acquired by each node of the drainage network to the target quality indicator.
[0104] Step S53: determining a target control index from the characteristic indexes of each data point in the current monitoring data according to the contribution score.
[0105] To more effectively improve drainage quality and enhance regulation and control effects, the impact of characteristic indicators on target quality indicators can be expressed by the contribution score of each data point's characteristic indicator to the predicted value corresponding to the target quality indicator. Characteristic indicators with the greatest impact on the target quality indicator are selected as target regulation indicators. By regulating these key target regulation indicators, the target quality indicator can be more effectively improved, bringing it closer to the target quality result.
[0106] Optionally, the contribution score of each feature indicator of the data point to the predicted value of the target quality indicator is expressed by the SHAP (SHapley Additive exPlanations) value. The SHAP value quantifies the marginal contribution of each feature indicator to the target quality result by calculating the difference between the target quality result with and without the feature indicator. For the target quality indicator n and feature indicator m, the SHAP value Indicates the contribution of feature index m to target quality index n. SHAP value It can be a positive or negative number. A positive number indicates that the characteristic indicator has an increasing effect on the target quality indicator, and a negative number indicates that the characteristic indicator has a decreasing effect on the target quality indicator.
[0107] For example, the SHAP value The calculation process can be summarized as the following steps: first determine the target quality index n, and then take each feature index in the data point as the feature set F. Determine all possible feature index combinations S, where S is a subset of F. For example, F can be {flow rate, liquid level, oxygen concentration, pipe diameter, pH value}, and S can be {oxygen concentration, pipe diameter, pH value}. For each feature index combination S, calculate the prediction result f(S) of the target quality index when the prediction model includes the features in S. For each feature index m, calculate its marginal contribution MC(m, S) in the feature combination S, that is, f(S∪{m})−f(S). Among them, S∪{m} means adding feature index m to the feature combination S. Afterwards, according to the probability of occurrence of feature combination S, that is, the number of combinations of the number of features in S, the marginal contribution MC(m, S) is weighted averaged, and finally the SHAP value of feature index m is obtained. :
[0108]
[0109] Where M represents the total number of feature indicators in the feature set F. SHAP value The absolute value of reflects the importance of feature index m to target quality index n. The larger the absolute value, the greater the influence of feature index m on the prediction result of target quality index n. The preset characteristic indicators with the highest ranking are used as target control indicators.
[0110] Step S54: determining the control amount of the target control indicator according to the deviation between the predicted value and the target value corresponding to the target quality indicator.
[0111] Optionally, step S54 includes steps S541 to S544:
[0112] Step S541 , based on the current monitoring data, simulate the control of the target control indicator according to the preset control value corresponding to each target control indicator through the prediction model to obtain the value of the target quality indicator after the simulated control.
[0113] For example, after determining the target control indicator, a series of preset control values are set for each target control indicator based on the current monitoring data collected from each node in the drainage network. For each target control indicator, its preset control values are iterated over, and the preset control value of the target indicator and the current monitoring data are input into the prediction model. The prediction model then outputs the value of the target quality indicator after adjustment based on the preset control value of the target indicator and the current monitoring data. In each iteration, the value of the current target control indicator can be set to the preset control value, while the values of other control indicators remain unchanged.
[0114] Step S542 , determining a reference control value of each target control indicator corresponding to when the value of the target quality indicator after the simulated control reaches an extreme value.
[0115] Step S543 : determining a control coefficient of each target control indicator according to a deviation between the detection value of each target control indicator in the current monitoring data and the reference control value.
[0116] Step S544 : determining the control amount of each target control indicator according to the control coefficient of each target control indicator and the deviation between the predicted value and the target value corresponding to the target quality indicator.
[0117] For example, the simulation results are analyzed to find the reference control value of each target control indicator when the target quality indicator reaches the extreme value. Assuming that the value of the target quality indicator after simulation control is closest to the target value, that is, when the simulation control effect is the best, the corresponding reference control value is Assuming that the deviation between the target quality index value after simulation control and the target value is the largest, that is, when the simulation control effect is the worst, the corresponding reference control value is .
[0118] For example, if the target quality indicator is hydrogen sulfide concentration, and its target value is 1 mg / m³, if the hydrogen sulfide concentration is simulated and regulated according to the preset control value [7, 7.5, 8, 8.5] of the characteristic indicator pH value, when the simulation control result is displayed as [1.8 mg / m³, 1.6 mg / m³, 1.5 mg / m³, 1.2 mg / m³], then is 8.5, is 7.
[0119] For each target control indicator, calculate the deviation between its current detection value Z and the reference control value corresponding to the extreme value of the target quality indicator after simulated control:
[0120]
[0121] Then, the control coefficient of each target control indicator is determined based on the deviation between the detection value of each target control indicator in the current monitoring data and the reference control value. : . Control coefficient The larger the value is, the greater the difference between the target control indicator and the ideal state, and a higher control priority is required.
[0122] Finally, the control amount of each target control indicator is determined based on the deviation between the detection value of each target control indicator in the current monitoring data and the reference control value, as well as the control coefficient of each target control indicator calculated above. For example, assuming that the target control indicator is flow rate, the deviation between its detection value and the reference control value is , then the flow rate control amount P, where is the flow rate control coefficient.
[0123] Step S55: determining a control scheme for the drainage network according to the target control index and the control amount of the target control index.
[0124] In this embodiment, the control coefficient can be used to quantify the degree of influence of each target control indicator on the target quality indicator. Combined with the deviation between the predicted value and the target value, the amount of adjustment required can be accurately calculated, thereby improving the control efficiency.
[0125] Based on the above embodiments of the present application, in the sixth embodiment of the present application, the same or similar contents as those in the above embodiments can be referred to the above introduction and will not be repeated hereafter. On this basis, after step S544, the training data enhancement method further includes steps S545 to S547:
[0126] Step S545 : dividing the drainage network into pipe segments according to the nodes of the drainage network.
[0127] For example, pipeline network data may include attributes such as node coordinates, pipeline connectivity, pipe diameter, and length. Using graph-theoretic algorithms such as depth-first search, the network topology is traversed to identify all connected regions. Within each connected region, pipeline segments are divided based on node connectivity, with each segment connecting two nodes. Each pipeline segment is assigned a unique identifier, and its attributes, such as length, diameter, upstream node, and downstream node, are recorded.
[0128] Step S546: Determine the time delay parameter between upstream nodes corresponding to the pipeline segment according to the length of the pipeline segment and the drainage flow rate.
[0129] The delay parameter represents the time required for water flow to propagate from the upstream node of the pipeline section to the downstream node, reflecting the transmission lag of water flow in the pipeline network. It is understandable that there is a lag in the transmission of water flow in the pipeline network. Directly regulating a certain node according to the real-time monitoring data of the current node may lead to over-regulation or under-regulation. When the control amount is uniformly allocated to each node, the execution effect of the control instructions at different nodes may be out of sync due to the delay difference, leading to other secondary problems. For example, when adjusting the flow, a node closes the valve too early or the valve is closed too small, causing a surge in upstream pressure. Therefore, for each node, the control amount ratio can be allocated according to the delay parameter to ensure the coordinated execution of the control instructions between different nodes.
[0130] For example, the length L of the pipeline section is obtained, and the drainage flow rate v in the pipeline section is determined by monitoring data at the nodes at both ends of the pipeline section, or by calculating the hydraulic model Where t is the roughness of the pipe section, R is the hydraulic radius, and S is the pipe slope. Then, the time delay parameter is determined based on the length L of the pipe section and the drainage velocity v in the pipe section. .
[0131] Step S547: Determine the control component of the target control indicator corresponding to each node according to the delay parameter.
[0132] Optionally, step S547 includes steps S5471 and S5472:
[0133] Step S5471: Determine the weight of the current node according to the distance between the current node and the end node of the drainage network.
[0134] Step S5472: Determine the control component of the target control indicator corresponding to the current node according to the weight and the delay parameter of the current node and the control amount of the target control indicator.
[0135] Control components are specific control quantities assigned to each node based on node weight, latency parameters, and the total control quantity. This ensures coordinated execution of control instructions across different nodes. It's understandable that the terminal node of a drainage network is the final outlet for drainage, and its status directly impacts the drainage capacity of the entire network. Nodes closer to the terminal node have their control effects reflected more quickly, and therefore should be given higher priority.
[0136] For example, the end node is identified from the pipe network data, and the Dijkstra algorithm is used to calculate the shortest path distance D from the current node to the end node. If there are multiple end nodes, such as a multi-outlet pipe network, the minimum distance is taken as the distance D between the current node and the end node of the drainage pipe network. The weight is assigned using an inverse proportional function, with nodes closer to the end node having a higher weight: Normalize all node weights to the range [0,1] to ensure that the sum of the weights is 1.
[0137] Then, get the weight of the current node , the delay parameters of the current node The total control amount ΔU of the target control index is calculated. According to the weight and delay parameters, the control component of the current node is allocated using the weighted average method. .in, is the sum of the delay parameters of all nodes. By dividing the delay parameter of the current node by the sum of the delay parameters of all nodes, a relative delay ratio can be obtained, thereby allocating the control components more reasonably.
[0138] In this embodiment, by dividing the drainage network into pipe segments and determining the delay parameters, and allocating the control components of the target control indicators according to the node weights and delay parameters, not only the refinement and coordination of the control are improved, but also the lag of water flow transmission in the drainage network is fully considered, avoiding over-regulation or under-regulation problems caused by delay differences, thereby improving the accuracy and reliability of drainage network control.
[0139] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the method for enhancing the training data of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.
[0140] The present application provides a training data enhancement device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the training data enhancement method of the above-mentioned embodiment 1.
[0141] Reference below Figure 3, which shows a schematic diagram of the structure of a training data augmentation device suitable for implementing the embodiments of the present application. The training data augmentation device in the embodiments of the present application may include, but is not limited to, mobile terminals such as laptop computers, tablet computers (PADs, Portable Application Descriptions), portable multimedia players (PMPs, Portable Media Players), and fixed terminals such as desktop computers. Figure 3 The training data enhancement device shown is only an example and should not limit the functions and scope of use of the embodiments of the present application.
[0142] like Figure 3 As shown, the training data augmentation device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 1002 or programs loaded from a storage device 1003 into a random access memory (RAM) 1004. RAM 1004 also stores various programs and data required for the operation of the training data augmentation device. Processing device 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems may be connected to I / O interface 1006: input device 1007 including, for example, a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output device 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage device 1003 including, for example, a magnetic tape, hard disk, etc.; and communication device 1009. Communication device 1009 may allow the training data augmentation device to communicate with other devices wirelessly or by wire to exchange data. While the figure illustrates a training data augmentation device with various systems, it should be understood that implementation or presence of all illustrated systems is not required. More or fewer systems may alternatively be implemented or present.
[0143] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0144] The training data enhancement device provided in this application, utilizing the training data enhancement method described in the aforementioned embodiment, can address the technical issue of low drainage network control accuracy caused by unbalanced data distribution. Compared to the prior art, the training data enhancement device provided in this application offers the same beneficial effects as the training data enhancement method described in the aforementioned embodiment. Other technical features of this training data enhancement device are the same as those disclosed in the aforementioned embodiment and are not further elaborated upon here.
[0145] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0146] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0147] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, a computer program) stored thereon, wherein the computer-readable program instructions are used to execute the training data enhancement method in the above-mentioned embodiment.
[0148] The computer-readable storage medium provided herein may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including, but not limited to, wires, optical cables, radio frequency (RF), etc., or any suitable combination thereof.
[0149] The computer-readable storage medium may be included in the training data enhancement device, or may exist independently without being incorporated into the training data enhancement device.
[0150] The computer-readable storage medium carries one or more programs that, when executed by a training data augmentation device, enable the training data augmentation device to write computer program code for performing the operations of the present application in one or more programming languages, or a combination thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, or as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0151] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of code, and the module, program segment or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or can be implemented using a combination of dedicated hardware and computer instructions.
[0152] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0153] The computer-readable storage medium provided in this application stores computer-readable program instructions (i.e., a computer program) for executing the aforementioned training data augmentation method. This computer-readable storage medium can address the technical issue of low drainage network control accuracy caused by unbalanced data distribution. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are similar to those of the training data augmentation method provided in the aforementioned embodiments and will not be further elaborated here.
[0154] The present application also provides a computer program product, comprising a computer program, which implements the steps of the above-mentioned training data enhancement method when executed by a processor.
[0155] The computer program product provided in this application can address the technical issue of low drainage network control accuracy caused by unbalanced data distribution. Compared to the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the training data enhancement method provided in the above-mentioned embodiment, and will not be further elaborated here.
[0156] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A method for enhancing training data, characterized in that: The training data enhancement method includes: Obtaining historical monitoring data obtained from each node of the drainage network, taking the historical monitoring data corresponding to each node as a data point to obtain an original data set; traversing each candidate first equalization parameter, determining a proximity distance between each data point and a neighboring data point in the original data set, and dividing the original data set into a plurality of groups of ordered data points and unordered data points based on the proximity distance between each data point and the neighboring data point in the original data set, wherein the first equalization parameter represents the number of the neighboring data points; For each group of ordered data points and disordered data points, linearly interpolate the ordered data points to obtain new ordered data points, and traverse each candidate second equalization parameter to add Gaussian noise to the disordered data points to generate new disordered data points; Taking each new group of the ordered data points and the unordered data points as a new original data set, and determining a target first equalization parameter from among the candidate first equalization parameters and a target second equalization parameter from among the candidate second equalization parameters according to the density distribution of each data point in the new original data set; Based on the target first equalization parameter, dividing the original data set into an ordered data set and an unordered data set according to a proximity distance between each data point in the original data set and a neighboring data point corresponding to the target first equalization parameter; performing linear interpolation on the ordered data set, and adding Gaussian noise to each data point in the unordered data set based on the target second equalization parameter to obtain enhanced data after equalization; After training the prediction model according to the enhanced data, determining the prediction quality result based on the current monitoring data obtained from each node of the drainage network by the prediction model; Determining a predicted value corresponding to a target quality indicator from the predicted quality result, and determining a target value corresponding to the target quality indicator from the target quality result; Determining the contribution score of each characteristic indicator of each data point in the current monitoring data obtained from each node of the drainage network to the target quality indicator; Determining a target control indicator from the characteristic indicators of each data point in the current monitoring data according to the contribution score; Determining the control amount of the target control indicator based on the deviation between the predicted value and the target value corresponding to the target quality indicator; A control scheme for the drainage network is determined based on the target control index and the control amount of the target control index.
2. The method for enhancing training data according to claim 1, wherein: The step of training the prediction model according to the enhanced data comprises: Training a preset model based on the enhanced data to obtain an initial prediction model; After equalizing the disordered data set in the enhanced data according to the target second equalization parameter, the initial prediction model is trained again based on the equalized enhanced data to obtain the prediction model.
3. The method for enhancing training data according to claim 1, wherein: The step of determining the control amount of the target control indicator according to the deviation between the predicted value and the target value corresponding to the target quality indicator includes: Based on the current monitoring data, the target control indicator is simulated and regulated by the prediction model according to the preset control value corresponding to each target control indicator, so as to obtain the value of the target quality indicator after the simulated control; Determine the reference control value of each target control indicator when the value of the target quality indicator after the simulated control reaches an extreme value; Determining a control coefficient for each target control indicator based on a deviation between a detection value of each target control indicator in the current monitoring data and the reference control value; The control amount of each target control indicator is determined according to the control coefficient of each target control indicator and the deviation between the predicted value and the target value corresponding to the target quality indicator.
4. The method for enhancing training data according to claim 3, wherein: After the step of determining the control amount of each target control indicator according to the control coefficient of each target control indicator and the deviation between the predicted value and the target value corresponding to the target quality indicator, the method further includes: Dividing the drainage pipe network into pipe segments according to each node of the drainage pipe network; Determining a time delay parameter of an upstream node corresponding to the pipeline segment according to the length of the pipeline segment and the drainage flow rate; According to the delay parameter, a control component of the target control indicator corresponding to each of the nodes is determined.
5. The method for enhancing training data according to claim 4, wherein: The step of determining the control component of the target control indicator corresponding to each node according to the delay parameter includes: Determining the weight of the current node according to the distance between the current node and the end node of the drainage network; The control component of the target control indicator corresponding to the current node is determined according to the weight and the delay parameter of the current node and the control amount of the target control indicator.
6. A training data enhancement device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the training data enhancement method according to any one of claims 1 to 5.
7. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the training data enhancement method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Engineering data source processing method and device, electronic equipment and storage medium
CN117874306A
Real-time vessel navigation tracking
US20200264268A1