A machine learning-based real-time push stream model construction method
By using a machine learning-based real-time flow propagation model to process multi-source hydrological data, identify backwater conditions, and optimize model parameters, the problem of insufficient flow calculation accuracy under backwater conditions in traditional methods is solved, and stable and continuous flow prediction is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA THREE GORGES PROJECTS DEV CO LTD
- Filing Date
- 2026-02-26
- Publication Date
- 2026-06-12
AI Technical Summary
Traditional index velocity methods are insufficient to accurately describe complex nonlinear and dynamic flow relationships in reservoir sections where downstream backwater and upstream inflow combine, leading to decreased accuracy in cross-sectional average velocity calculations and failing to meet the accuracy requirements of hydrological monitoring.
A machine learning-based real-time flow propagation model is adopted. By acquiring and processing multi-source hydrological data, backwater flow condition samples are identified, feature importance scores are calculated, a support vector machine model is established, and a particle swarm optimization algorithm with weighted fitness for different working conditions is introduced to optimize the model parameters. Combined with an adaptive filtering mechanism, flow prediction is performed.
It enables stable, continuous, and physically consistent real-time estimation of station flow under complex hydrological scenarios, improving the reliability and accuracy of flow estimation results and solving the problems of insufficient accuracy and large output fluctuation of traditional models under top-support conditions.
Smart Images

Figure CN122197547A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of hydrological monitoring technology, and in particular to a method for constructing a real-time flow propagation model based on machine learning. Background Technology
[0002] In the field of hydrological measurement and flow monitoring, flow rate is typically used as a crucial basic parameter for water resource allocation, flood warning, and river management. Its acquisition has long relied on manual flow measurement or calculation methods based on water level-flow curves. With the continuous improvement of automation and informatization levels at monitoring stations, online monitoring equipment such as ultrasonic current meters and water level gauges have been widely deployed. How to process multi-source hydrological data in real time and output stable and reliable flow rate results within a computer system has gradually become one of the core issues in hydrological information systems.
[0003] Traditional index velocity methods have significant limitations in test sections of reservoirs where both downstream backwater and upstream inflow act. When the monitoring section is unaffected by backwater, the index velocity and the average velocity of the section typically exhibit a simple and well-correlated fit, allowing for the establishment of stable mathematical models using traditional methods. However, when the monitoring section is affected by backwater, the flow pattern changes, resulting in a backwater-type non-uniform flow in the reservoir area. This leads to multiple different and complex relationship series between the index velocity and the average velocity of the section. The same index velocity value may correspond to multiple different average velocity values, forming multiple discrete clusters of relationship points. Traditional linear fitting methods struggle to accurately describe this complex relationship, which changes in real time with factors such as backwater intensity and water level difference, exhibiting significant nonlinearity and dynamic characteristics. This results in a significant decrease in the accuracy of the average velocity calculation, failing to meet the precision requirements of hydrological monitoring. Summary of the Invention
[0004] Therefore, it is necessary for the present invention to provide a method for constructing a real-time streaming model based on machine learning in order to solve at least one of the above-mentioned technical problems.
[0005] To achieve the above objectives, a method for constructing a real-time streaming model based on machine learning includes the following steps: Step S1: Obtain the historical raw feature dataset and raw traffic dataset; perform data quality unification processing on the raw feature dataset and raw traffic dataset to eliminate abnormal disturbances and restore time continuity, thus obtaining the feature dataset and traffic dataset. Step S2: Identify top-load flow condition samples in the feature dataset based on the flow dataset and calculate their proportion. Calculate the importance score of each feature in the feature dataset based on the proportion. Filter the feature dataset according to the importance score to form a filtered feature set. Step S3: Calculate the overburden level value based on the filtered feature set and determine the sample training weights. Extract the flow value corresponding to the non-overburden working condition from the flow dataset and calculate the water level flow constraint coefficient. Step S4: Using the filtered feature set as input and the flow dataset as target, establish a support vector machine model that integrates the training weights of the fused samples and the water level and flow constraint coefficients. Optimize the model parameters using a particle swarm optimization algorithm with weighted fitness for different working conditions to obtain the trained model. Step S5: Input the real-time collected hydrological features into the trained model to obtain the predicted flow rate, and adjust the filtering coefficient according to the real-time water level change to filter the predicted flow rate to obtain the final flow rate and output it.
[0006] This invention introduces a machine learning modeling framework oriented towards backwater hydrodynamic conditions, enabling stable, continuous, and physically consistent real-time estimation of station flow under complex hydrological scenarios. Based on multi-source hydrological observation data, this method first performs unified data quality processing on historical feature data and flow data. Through outlier identification, segmented interpolation, and smoothing, it significantly reduces the adverse effects of observational noise and missing data on model training, making the training samples more consistent with the objective characteristics of hydrological processes in terms of temporal continuity and statistical stability. Furthermore, by constructing a backwater flow condition identification mechanism based on the deviation between measured flow and natural flow, it can identify backwater conditions using differentiated criteria at different river scales. This allows the model to clearly distinguish between complex conditions influenced by downstream water levels and conventional hydrodynamic conditions during the training phase, effectively avoiding model bias caused by mixing samples from different conditions.
[0007] By incorporating the proportion of backwater conditions into the feature importance assessment process, feature selection is achieved based on information gain and correlation analysis. This ensures that the model input retains key hydrological features strongly correlated with flow rate while adaptively adjusting the feature weight structure according to the degree of backwater impact, thereby improving the model's sensitivity and interpretability to key control factors. During model construction, quantifying the degree of backwater and introducing sample training weights allows the model to pay greater attention to samples with significant backwater impact during the parameter learning phase, enhancing its fitting ability under complex hydrodynamic conditions. Simultaneously, by introducing a constraint term for the consistency of water level-flow changes into the objective function, hydrophysical mechanisms and statistical learning processes are organically integrated, ensuring that the model's prediction results maintain a reasonable trend over time, avoiding abrupt changes or reversals that contradict hydrodynamic laws. In the parameter optimization phase, a particle swarm optimization algorithm with weighted fitness for different flow conditions is employed to comprehensively balance prediction bias and flow continuity under different flow conditions, achieving global optimization of the model's hyperparameters and ensuring stable performance under both normal and backwater conditions. In the real-time flow projection phase, a dynamic filtering mechanism driven by water level changes is introduced into the prediction result output process. This enables the system to maintain continuous flow output when the hydrological process is stable and to respond promptly to hydrodynamic changes when the water level changes rapidly, thus balancing output stability and real-time performance. Overall, this method effectively solves the problems of insufficient accuracy, poor physical consistency, and large output fluctuations in traditional flow projection models under backwater conditions, significantly improving the reliability and engineering application value of flow projection results in complex hydrological environments. Attached Figure Description
[0008] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is a flowchart illustrating the steps of the real-time streaming model construction method based on machine learning of the present invention. Figure 2 This is a schematic diagram of the overall process of a real-time streaming model construction method based on machine learning according to an embodiment of the present invention. Figure 3 This is a schematic diagram illustrating the influence of upstream and downstream hydrodynamic conditions on the flow measurement at this station according to an embodiment of the present invention. Detailed Implementation
[0009] The technical method of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0010] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.
[0011] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0012] To achieve the above objectives, please refer to Figures 1 to 3 This invention provides a method for constructing a real-time streaming model based on machine learning, the method comprising the following steps: Step S1: Obtain the historical raw feature dataset and raw traffic dataset; perform data quality unification processing on the raw feature dataset and raw traffic dataset to eliminate abnormal disturbances and restore time continuity, thus obtaining the feature dataset and traffic dataset. Step S2: Identify top-load flow condition samples in the feature dataset based on the flow dataset and calculate their proportion. Calculate the importance score of each feature in the feature dataset based on the proportion. Filter the feature dataset according to the importance score to form a filtered feature set. Step S3: Calculate the backwater level value based on the filtered feature set and determine the sample training weights. Extract the flow rate value corresponding to the non-backwater condition from the flow data set and calculate the water level flow constraint coefficient. Step S4: Using the filtered feature set as input and the flow dataset as target, establish a support vector machine model that integrates the training weights of the fused samples and the water level and flow constraint coefficients. Optimize the model parameters using a particle swarm optimization algorithm with weighted fitness for different working conditions to obtain the trained model. Step S5: Input the real-time collected hydrological features into the trained model to obtain the predicted flow rate, and adjust the filtering coefficient according to the real-time water level change to filter the predicted flow rate to obtain the final flow rate and output it.
[0013] See Figure 2This paper illustrates the overall process of flow prediction based on hydrodynamic feature screening and adaptive filtering according to the present invention. The process includes a training sample construction stage, a model training stage, and an online prediction and adaptive filtering stage, with each stage forming a complete technical closed loop under the constraints of hydrodynamic mechanisms.
[0014] During the training sample construction phase, upstream and downstream water levels, the water level at the local station, hydraulic radius, flow area, and historical flow data were first collected, along with flow velocity data obtained using the ultrasonic time-of-flight method. Outlier processing was then performed on the raw data to eliminate the impact of measurement errors and sudden disturbances on subsequent modeling.
[0015] After data preprocessing, based on the correlation between water level changes and flow conditions, backwater flow condition samples are identified, and the proportion of backwater samples in the overall sample is statistically analyzed. This step enables the training samples to cover different hydrodynamic conditions, especially complex and unsteady flow scenarios such as backwater and water inflow.
[0016] Subsequently, the importance of all candidate features was evaluated, and the hydrodynamic features most relevant to flow changes were selected to form a selected feature set. Based on this, feature engineering was carried out to construct a training sample set.
[0017] During the model training phase, the constructed training samples are input into the SVR model for training. The model parameters are optimized through the objective function. Once the model output flow value meets the accuracy requirements, the trained prediction model is obtained.
[0018] During the online prediction phase, ultrasonic time-of-flight flow velocity data and filtered hydrodynamic characteristic data are collected in real time and input into the trained model to obtain the predicted flow rate at the current moment.
[0019] Because the prediction model is prone to short-term fluctuations under conditions of drastic changes in water flow or rapid fluctuations in water level, this invention further introduces an adaptive filtering mechanism. This mechanism does not use fixed filtering parameters, but rather characterizes the water level change based on the difference between the current water level at the test section and the water level at the previous moment. The water level change is then input into a Sigmoid function for nonlinear mapping, resulting in dynamic filtering coefficients located within the preset minimum and maximum filtering coefficient range.
[0020] The predicted flow rate is weighted and fused with the filtered flow rate from the previous moment using the dynamic filtering coefficient to obtain the final high-precision flow rate output value.
[0021] Through the above process, the filtering intensity can be automatically adjusted according to the hydrodynamic state: it enhances the stabilizing effect of historical flow when the water level changes gently, and enhances the response capability of predicted flow when the water level changes drastically, thereby significantly improving the stability and accuracy of flow calculation under complex hydrological conditions.
[0022] Furthermore, step S1 includes the following steps: Step S11: Obtain the ultrasonic time-of-flight flow velocity, test section water level, hydraulic radius, flow area, upstream water level and downstream water level of the station over the years to form the original feature dataset; In one embodiment, various monitoring data reflecting the hydrodynamic characteristics of the cross-section and the upstream-downstream water level relationship are acquired from the historical hydrological monitoring system, automatic flow measurement equipment, and related hydrological data of the target station at a unified time scale to construct an original feature dataset. The data includes: average flow velocity data of the cross-section measured using the ultrasonic time-of-flight method; water level data corresponding to the test cross-section; hydraulic radius data calculated based on the cross-sectional geometric parameters; flow area data corresponding to different water level conditions; and water level data from upstream and downstream stations used to characterize inflow and outflow conditions. All of the above-mentioned feature data are organized according to the same time benchmark and stored in time series format to form a multidimensional original feature dataset.
[0023] For example, the flow velocity, water level, and upstream and downstream water level data of a certain station over many years can be summarized on a daily or fixed time period basis, and the flow area and hydraulic radius corresponding to each time period can be calculated based on the cross-sectional measurement results, thereby forming a characteristic sample set covering the complete hydrological process.
[0024] Step S12: Obtain flow data from historical hydrological data of the station and interpolate non-hourly times to form the original flow dataset; In one embodiment, flow data consistent with the time range is obtained from historical hydrological measurement data, flow yearbooks, or scheduling results of the monitoring station, serving as the target data for model training and validation. To address situations where the recording time of some flow data is inconsistent with the time step used by the model, time interpolation is performed on flow records for non-hourly or non-standard time periods to ensure consistency between the flow data and feature data on the time scale, thus forming the original flow dataset.
[0025] For example, when flow data is recorded in a non-standard time period, and the model uses daily or fixed time periods as the time step, flow data from two adjacent standard time periods can be interpolated to obtain flow estimates at the corresponding time scale, thereby establishing a correspondence with hydrological characteristic data of the same time period.
[0026] Step S13: For each data point in the original feature dataset and the original traffic dataset, calculate the mean and standard deviation of the 10 data points before and after it. If the absolute difference between the data point and the mean is greater than three times the standard deviation, the data point is removed.
[0027] In one embodiment, to remove abnormal data points caused by monitoring anomalies, communication failures, or occasional interference, anomaly detection is performed on each time sample in the original feature dataset and the original traffic dataset. Specifically, a local time window is formed by selecting data from 10 time points before and after the current data point, and the mean and standard deviation of the data within this window are calculated. When the absolute difference between the current data point and the mean is greater than three times the standard deviation, the data point is identified as an outlier and removed.
[0028] For example, if a water level sample changes steadily over 30 consecutive time points, but the water level value at that moment shows a significant sudden increase or decrease exceeding three times the standard deviation, then the water level data is judged as abnormal and removed from the dataset.
[0029] Furthermore, step S1 also includes the following steps: Step S14: For the data after removing outliers, linear interpolation is used when the missing period is no more than three hours, and cubic spline interpolation is used when the missing period is more than three hours. In one embodiment, the missing time series segments formed after outlier removal are imputed to restore data continuity. When the continuous missing period does not exceed three hours, linear imputation is used, which involves linear estimation based on adjacent valid data points before and after the missing segment. When the continuous missing period exceeds three hours, cubic spline imputation is used to smoothly fit the data change trend within the longer missing segment, thereby completing the missing data.
[0030] For example, when a certain flow velocity or water level characteristic is missing for only a short period of time, it can be directly interpolated linearly using data from the preceding and following time periods; however, when equipment failure causes data to be missing for several consecutive hours, the trend of change within that time period can be reconstructed using a cubic spline function.
[0031] Step S15: Smooth the interpolated data using a moving average method with a window radius of 3 to obtain the feature dataset and the traffic dataset.
[0032] In one embodiment, the feature dataset and traffic dataset after interpolation are further smoothed to reduce the interference of high-frequency fluctuations on model training. Specifically, a moving average method with a window radius of 3 is used for each time series. That is, for each time point, the average value of its own data and the data of the three time points before and after it is calculated, and the average value is used as the smoothed data result to finally obtain the feature dataset and traffic dataset for subsequent modeling.
[0033] For example, for cross-sectional flow velocity data of a certain period, the flow velocity values of multiple adjacent periods before and after it can be averaged, thereby reducing the impact of short-term fluctuations on model parameter learning.
[0034] It should be noted that the smoothing process improves the stability of the data and the quality that can be used for machine learning modeling while maintaining the overall hydrological trend without distortion.
[0035] Furthermore, step S2 includes the following steps: Step S21: When the downstream water level in front of the dam is lower than the preset threshold, it is determined to be a non-backflow state, and the corresponding sample in the flow data is determined to be a non-backflow flow sample. In one embodiment, the downstream upstream water level data of each sample in the flow dataset at a given time is compared with a pre-set upstream water level threshold to identify basic samples that are clearly unaffected by downstream water level backflow. When the downstream upstream water level at a certain time is lower than the preset threshold, it indicates that the downstream control water level at that time has a weak or negligible impact on the flow conditions at the station section. Therefore, the sample corresponding to that time is identified as a non-backflow state, and the corresponding sample in the flow dataset is marked as a non-backflow flow condition sample.
[0036] For example, based on historical operating experience or design water level data, the downstream dam front water level threshold can be set to a safe water level below the dam front warning water level or normal storage water level. When the real-time or historical water level is lower than this value, it can be directly determined to be in a non-backwater state.
[0037] It should be noted that the judgment in this step is only used to quickly remove samples that are obviously not affected by the top support. It is a preliminary screening and does not involve the quantitative calculation of the intensity of the top support effect.
[0038] Step S22: When the downstream water level in front of the dam is not lower than the preset threshold, it is determined to be in a backwater state. The natural flow value of the sample is calculated based on the upstream inflow and the runoff in the interval during the same period. In one embodiment, when it is determined that the downstream water level in front of the dam is not lower than a preset threshold at a certain moment, it is considered that the downstream water level at that moment may cause backflow or backwater effect on the station section, and thus the sample is identified as a candidate sample for backwater state. Based on this, in order to eliminate the interference of downstream backwater on the measured flow, it is necessary to construct a reference flow value for the sample under natural hydraulic conditions, i.e., the natural flow value.
[0039] Specifically, based on the upstream inflow data at the same time and the runoff data between the monitoring station and the upstream control station, the upstream inflow is superimposed and corrected to calculate the natural flow value when it is not affected by the downstream water level.
[0040] For example, the measured flow rate at the upstream control station can be used as the base inflow, and combined with the interval rainfall runoff or tributary inflow, and linearly or proportionally superimposed to obtain the natural flow rate value corresponding to that moment.
[0041] Step S23: Extract the measured flow value of each sample from the flow dataset, calculate the difference between the measured flow value and the natural flow value, and divide by the natural flow value to obtain the flow deviation ratio; In one embodiment, the measured flow value corresponding to each sample is extracted one by one from the flow dataset, and the measured flow value is compared and analyzed with the natural flow value calculated in step S22. Specifically, by calculating the difference between the measured flow value and the natural flow value, and dividing the difference by the natural flow value, the flow deviation ratio of the sample is obtained, which is used to reflect the degree of relative change of the measured flow relative to the flow under natural conditions.
[0042] For example, if the measured flow rate at a certain moment is 950 m³ / s, while the calculated natural flow rate is 1000 m³ / s, then the flow rate deviation ratio of this sample is (950−1000) / 1000=−5%.
[0043] Step S24: When the station belongs to a major river station, the flow deviation ratio is less than or equal to -6. This sample was identified as a top-support flow condition. In one embodiment, when the target station is classified as a major river station, the sample operating conditions are determined based on the flow deviation ratio; when the flow deviation ratio is less than or equal to -6%, the sample is determined to be a backwater flow sample and is marked accordingly. This determination criterion is used to identify flow attenuation caused by downstream water level rise in large-scale rivers.
[0044] For example, if the flow deviation ratio of a certain main stream control station is -7% within a certain time period, then the sample of that time period is determined to be a backflow condition sample.
[0045] Step S26: When the station is a medium-sized river station, the flow deviation ratio is less than or equal to -10. This sample was identified as a top-support flow condition. In one embodiment, when the target station is classified as a medium-sized river station, the operating condition is determined based on the flow deviation ratio obtained in step S21; when the flow deviation ratio is less than or equal to −10%, the corresponding sample is determined as a backwater condition sample. This threshold is used to adapt to the typical flow response characteristics of medium-scale rivers under backwater conditions.
[0046] For example, when the flow deviation ratio of a medium-sized river station reaches -11% within a certain time period, the sample is marked as a backwater flow condition sample.
[0047] Step S26: When the station is a small river station, the flow deviation ratio is less than or equal to -15. This sample was identified as a top-support flow condition. In one embodiment, when the target station is classified as a small river station, a more stringent determination of backflow conditions is made based on the proportion of flow deviation; when the proportion of flow deviation is less than or equal to -15%, the sample is determined to be a backflow condition sample. This determination method is used to avoid misidentifying slight flow changes in small rivers caused by local disturbances or measurement fluctuations as backflow phenomena.
[0048] For example, a sample is only considered to be a backflow condition sample when the flow deviation ratio of a small river monitoring station in a mountainous area is -16% within a certain time period.
[0049] Step S27: Count the total number of top-support flow condition samples and divide by the total number of samples to obtain the percentage value.
[0050] In one embodiment, after determining the backflow condition for all samples, the number of samples marked as having a backflow condition is counted, and this number is divided by the total number of samples to obtain the proportion of samples with a backflow condition in the overall sample. This proportion is used to characterize the frequency of backflow phenomena in historical data and serves as a weighting adjustment parameter in subsequent feature importance calculations.
[0051] For example, when the total number of samples is 12,000, and the top-support flow condition samples are 3,600, the proportion is 0.30.
[0052] Furthermore, step S2 also includes the following steps: Step S28: Establish a mapping relationship between each feature in the feature dataset and the traffic dataset, calculate the information gain score of the feature, and calculate the absolute value of the Pearson correlation coefficient between the feature and the traffic dataset as the correlation coefficient score. In one embodiment, for each feature variable in the feature dataset, a correspondence is established with the traffic dataset to evaluate the explanatory power of the feature for traffic changes. Specifically, the information gain score of the feature relative to the traffic dataset is calculated to reflect the feature's contribution to reducing traffic uncertainty; simultaneously, the Pearson correlation coefficient between the feature and the traffic dataset is calculated, and its absolute value is taken as the correlation coefficient score.
[0053] For example, the information gain value and correlation coefficient score between cross-sectional water level, hydraulic radius and flow rate can be calculated separately for subsequent feature importance evaluation.
[0054] Step S29: Add the product of the information gain score and the proportion value to the product of the correlation coefficient score and the proportion value minus one, to obtain the importance score; In one embodiment, the information gain score of each feature is multiplied by its proportion value, and the correlation coefficient score of the feature is multiplied by one minus the proportion value. The two results are then added together to obtain the overall importance score of the feature. By introducing the proportion value, adaptive adjustment of the focus of feature evaluation under different frequencies of top-down occurrence is achieved.
[0055] For example, when the information gain score of a feature is 0.55, the correlation coefficient score is 0.75, and the proportion is 0.4, then the importance score of that feature is 0.55×0.4+0.75×0.6.
[0056] Step S210: Set the ultrasonic time difference flow velocity feature, test section water level feature, and downstream first station water level feature in the feature dataset as mandatory retention features. Select several features from the remaining features according to their importance scores and merge them with the mandatory retention features to form a filtered feature set.
[0057] In one embodiment, the ultrasonic time-of-flight velocity features, the water level features at the test section, and the water level features at the first downstream station in the feature dataset are pre-set as mandatory retention features to ensure that the model input always includes core variables reflecting the hydrodynamic state of the section and the downstream backwater relationship. Subsequently, the remaining features, excluding the mandatory retention features, are sorted from high to low according to their importance scores, and several of the top-ranked features are selected and merged with the mandatory retention features to form a filtered feature set for subsequent modeling.
[0058] For example, among the remaining features, features with higher importance scores, such as the water level of the first upstream station and the hydraulic radius, are selected first and then combined with the three mandatory retention features to form the model input feature set.
[0059] See Figure 3 This diagram illustrates the hydrodynamic state of the station under the influence of downstream control structures when it is in a backwater backing condition.
[0060] like Figure 3 As shown, the river flow direction is from the upstream station to the downstream station. Upstream stations, this station, and downstream stations are sequentially located along the river. Control structures such as sluice gates and dams are located at the downstream station. Under normal operating conditions, the river surface exhibits a stable gradient, and the flow velocity changes gently along the river. The water level and velocity measured by the water level gauge and ultrasonic current meter accurately reflect the natural flow conditions.
[0061] When the downstream dam is raised or operational scheduling causes the downstream water level to rise, the water level at the downstream station rises from the normal level to the backwater level, forming a significant backwater backwater phenomenon, resulting in a backwater area between this station and the downstream station. Affected by this backwater area, the water surface line at this station rises overall, with the water level increasing by approximately H=0.5 m compared to normal operating conditions, while the water surface slope decreases significantly.
[0062] Under this backwater backwater condition, the kinetic energy of the water flow is suppressed, and the flow velocity near this station shows a significant decrease, for example, from about 2.5 m / s under natural conditions to about 1.8 m / s, a decrease of about 28%. Although the upstream inflow conditions have not changed significantly, the downstream backwater effect has altered the local hydrodynamic structure, resulting in a systematic difference between the measured flow velocity at this station and the equivalent flow velocity under natural flow conditions.
[0063] The aforementioned rise in water level and decrease in flow velocity are not caused by changes in the flow rate itself, but by the non-constant hydrodynamic conditions resulting from the backwater effect caused by downstream water level fluctuations. This makes it easy to underestimate the flow rate or lead to calculation errors when relying solely on a single flow velocity or water level parameter. This invention introduces upstream and downstream water level information, hydraulic parameters, and flow characteristics, and combines this with a machine learning-based flow propulsion model to model and correct the backwater effect, enabling it to accurately reflect the true water flow capacity of the station even in the presence of backwater.
[0064] Furthermore, step S3, which calculates the degree of support based on the filtered feature set and determines the sample training weights, includes: The downstream first station water level value and the test section water level value of each training sample are extracted from the filtered feature set; The difference between the water level at the first downstream station and the water level at the test section is calculated and divided by the pre-obtained water level difference between the two stations under natural conditions to obtain the backwater level value. The preset value range of the top support sensitivity coefficient is 0.5-2.0. The product of the top support degree value and the top support sensitivity coefficient plus one is used as the sample training weight.
[0065] In one embodiment, from the aforementioned filtered feature set, for each training sample, the downstream first station water level value and the test section water level value at the corresponding time are read, wherein the downstream first station water level is used to reflect the downstream control boundary conditions, and the test section water level is used to reflect the actual water level state of the current station. Subsequently, for the same training sample, the difference between the downstream first station water level value and the test section water level value is calculated, and this difference is divided by the water level difference between the downstream first station water level value and the current station water level under natural conditions, thereby obtaining a dimensionless backwater degree value, which is used to characterize the strength of the backwater influence of the downstream water level on the current test section flow conditions.
[0066] For example, if the water level at the test section in a certain training sample is 10.0m and the water level at the first station downstream is 10.3m, then the difference between the two is 0.3m. Dividing this difference by the water level at the test section of 10.0m, we get the backwater level value as 0.03.
[0067] After obtaining the support level value, a support sensitivity coefficient is pre-set, with a value range limited to 0.5 to 2.0, to adjust the influence weight of the support level on the model training process. For each training sample, its corresponding support level value is multiplied by the selected support sensitivity coefficient, and one is added to the product to obtain the sample training weight, so that samples with more significant support levels have higher weights in model training.
[0068] Furthermore, step S3, which involves extracting the flow rate value corresponding to the non-top-support condition from the flow rate dataset and calculating the water level-flow constraint coefficient, includes: Samples with a support level value less than 0.8 from all training samples are selected and marked as samples with weak support influence. Extract the flow value corresponding to the weak backwater impact sample from the flow dataset, calculate the difference between the flow at the current time and the flow at the previous time to obtain the flow change, and calculate the difference between the water level at the current time and the water level at the previous time to obtain the water level change. The water level-discharge constraint coefficient is obtained by summing the products of the flow rate changes and water level changes of all samples affected by weak backwater and dividing them by the sum of the squares of the water level changes.
[0069] In one embodiment, based on the calculated backwater level values, all training samples are screened, and samples with backwater level values less than 0.8 are selected and uniformly labeled as weak backwater impact samples. This ensures that the selected samples are basically unaffected by downstream water level backwater and can reflect a relatively natural water level-flow response relationship. Subsequently, the flow values corresponding to the weak backwater impact samples at each time point are extracted from the flow dataset, and the corresponding water level values are extracted simultaneously.
[0070] For each sample affected by weak backwater, the difference between the current flow rate and the previous flow rate is calculated to obtain the flow rate change; simultaneously, the difference between the current water level and the previous water level is calculated to obtain the water level change. Next, for all samples affected by weak backwater, the product of the flow rate change and the water level change is calculated, and the results are summed. Simultaneously, the water level change for all samples affected by weak backwater is squared, and the squared results are summed. Finally, the sum of the products of the flow rate change and the water level change is divided by the sum of the squares of the water level changes to obtain the water level-flow constraint coefficient. This coefficient characterizes the overall constraint relationship between water level change and flow rate change under non-backwater conditions.
[0071] For example, if in several samples with weak backwater influence, the water level rises by 0.1m at one moment, corresponding to an increase in flow rate of 20m³ / s, and the water level rises by 0.05m at another moment, corresponding to an increase in flow rate of 9m³ / s, then by summarizing and calculating the changes in multiple similar samples, a stable water level-flow constraint coefficient can be obtained to reflect the water level-flow coupling characteristics of the station under natural conditions.
[0072] It should be noted that by selecting only samples with weak backwater effects to calculate the water level-discharge constraint coefficient, the interference of the backwater effect on the water level-discharge relationship can be effectively avoided, thereby improving the physical rationality and stability of the constraint coefficient in the subsequent model training and prediction process.
[0073] Furthermore, step S4 includes the following steps: Step S41: Divide the filtered feature set into a training set and a validation set in a 7:3 ratio. Use the training set as the model input and the corresponding traffic dataset as the model training target to build a support vector machine model, where the support vector machine model is represented as: In the formula, For the filtered feature set, A nonlinear mapping function maps the input to a high-dimensional space. Let be the weight vector to be solved. The bias term to be solved. The flow rate predicted by the model. This is the matrix transpose symbol; In one embodiment, the filtered feature set obtained in the aforementioned steps is divided into a training set and a validation set in a 7:3 ratio according to the number of samples. The training set is used for learning the model parameters, and the validation set is used to evaluate the model's generalization performance on samples not used in the training. Subsequently, the feature vectors in the training set are used as the input to the support vector machine model, and the traffic dataset corresponding one-to-one with the training samples is used as the model's training target output, thereby constructing a traffic prediction support vector machine model.
[0074] In the model construction process, the regression form of support vector machines is adopted, and its mathematical expression is: ,in, This represents the filtered feature vector corresponding to a single training sample. This represents the result of mapping the input features to a high-dimensional feature space using a pre-defined nonlinear mapping function. This represents the weight vector to be solved in a high-dimensional space. This represents the bias term to be solved. This represents the model's predicted flow rate for the corresponding input sample. This form transforms hydrological features that are difficult to linearly fit in the original feature space into approximately linear relationships in a high-dimensional space, thereby improving the accuracy of flow rate prediction.
[0075] For example, for a certain training sample, its input features include ultrasonic time-of-flight flow velocity, water level at the test section and water level at the first downstream station, etc. After nonlinear mapping, they are input into the support vector machine model. The model outputs the predicted flow value at the corresponding time, which is used to calculate the error with the measured flow.
[0076] Step S42: Construct an objective function that integrates the training weights of the sample samples and the water level and flow constraint coefficients in the support vector machine model; In one embodiment, after defining the basic structure of the support vector machine (SVM) model, its objective function is improved by incorporating the sample training weights and water level / flow constraint coefficients obtained in the preceding steps. Specifically, based on the traditional SVM regression model's loss function, which focuses on prediction error, corresponding sample training weights are introduced into the error terms of different training samples. This ensures that samples with higher backwater levels and larger weights contribute more to the overall error in the objective function, thereby guiding the model to pay more attention to the flow prediction accuracy under backwater conditions during training.
[0077] Meanwhile, a water level-flow constraint coefficient is introduced into the objective function to constrain the trend of the model prediction results at adjacent time points, so that the predicted flow rate change and water level change satisfy the water level-flow coupling relationship obtained statistically under non-top backing conditions, thereby enhancing the physical rationality of the model prediction results.
[0078] For example, when a training sample has a large training weight, the deviation between its predicted flow and the measured flow will be amplified in the objective function, thereby prompting the model parameters to prioritize reducing the prediction error of this type of sample during the optimization process; while the water level and flow constraint coefficient is used to limit the model output from drastic fluctuations that do not conform to the water level change pattern at adjacent time points.
[0079] It should be noted that the introduction of sample training weights and water level / flow constraint coefficients are both reflected at the objective function level and do not change the basic structural form of the support vector machine model.
[0080] Step S43: Solve the objective function using the particle swarm optimization algorithm with weighted fitness for different work conditions, and substitute the optimal hyperparameter combination obtained from the search into the support vector machine model to obtain the optimal model parameters, thus obtaining the trained model.
[0081] In one embodiment, after constructing the objective function for the fusion sample training weights and water level / flow constraint coefficients, a particle swarm optimization algorithm with weighted fitness for different working conditions is used to globally optimize the hyperparameters of the support vector machine model. Specifically, the penalty factor and kernel function parameters in the support vector machine model are used as variables to be optimized in the particle swarm optimization algorithm. Each particle represents a set of candidate hyperparameter combinations, and the search space is iteratively updated.
[0082] In each iteration, the fitness function is weighted according to the differences in the importance of samples under different working conditions. This allows the particle to simultaneously consider the predictive performance of samples under conditions of strong support and samples with weak support when evaluating its performance, thus avoiding excessive bias of model parameters towards a particular type of working condition. The particle updates its hyperparameters based on its own historical best position and the group's best position, gradually approaching the optimal hyperparameter combination that minimizes the objective function.
[0083] Once the particle swarm optimization algorithm meets the preset convergence condition or reaches the maximum number of iterations, the optimal hyperparameter combination obtained from the search is substituted into the support vector machine model, and the model parameters w and b are finally solved to obtain the trained traffic prediction model.
[0084] For example, after multiple iterations, the particle swarm optimization algorithm determines that a certain set of penalty factors and kernel function parameters has the minimum weighted error on the validation set. Then, this set of parameters is used to train the support vector machine model to obtain the final optimal model.
[0085] Of particular importance, step S43 includes the following steps: Step S431: Initially set the particle swarm size to twenty to thirty particles. The position of each particle represents a set of hyperparameter candidate values, including data fitting penalty coefficient, physical constraint penalty coefficient and insensitivity coefficient. In one embodiment, before executing the particle swarm optimization algorithm, the size of the particle swarm is initially set to twenty to thirty particles to ensure both search diversity and computational efficiency. Each particle corresponds to a set of candidate hyperparameter values for the support vector machine model. These candidate hyperparameter values include at least a data fitting penalty coefficient to control the model's fitting complexity, a physical constraint penalty coefficient to constrain the model to satisfy the water level-flow rate physical relationship, and an insensitivity coefficient used in support vector machine regression to define the error tolerance interval. In this way, the model hyperparameter optimization problem is mapped to a particle search problem in a continuous parameter space.
[0086] For example, the position of a particle can be represented as [data fitting penalty coefficient = 10, physical constraint penalty coefficient = 5, insensitivity coefficient = 0.05], corresponding to a specific combination of model hyperparameters.
[0087] Step S432: Randomly initialize the position of each particle within the search range of each hyperparameter. For the hyperparameter combination represented by the current particle position, solve for the corresponding model parameters according to the objective function. Calculate the predicted flow rate for each sample using the model parameter validation set. In one embodiment, after setting the particle swarm size, the position of each particle is randomly initialized within the search range corresponding to each hyperparameter, so that the initial particles are uniformly distributed in the parameter space. Subsequently, for each particle's current position, the combination of hyperparameters is substituted into the aforementioned constructed support vector machine objective function to obtain the corresponding model parameters. These model parameters are then used to predict the flow rate of samples in the validation set, obtaining the predicted flow rate value for each sample in the validation set.
[0088] For example, for a certain particle, its initial combination of penalty coefficients is substituted into the model training to obtain a set of weight vectors and bias terms, and based on this, a predicted flow sequence is output for all samples in the validation set.
[0089] Step S433: Include samples in the validation set with a backflow intensity value less than 0.8 in the normal flow validation set, and include samples with a value greater than or equal to 0.8 in the backflow flow validation set. In one embodiment, after obtaining the predicted flow rates of each sample in the validation set, the validation samples are divided into different operating conditions based on the calculated backwater level values. Specifically, samples with backwater level values less than 0.8 are classified as the normal flow condition validation set, and samples with backwater level values greater than or equal to 0.8 are classified as the backwater flow condition validation set, thereby explicitly distinguishing the model's predictive performance under different hydrodynamic conditions during the validation phase.
[0090] For example, when the top pressure level value corresponding to a certain verification sample is 0.015, it is classified into the normal flow verification set; when the top pressure level value is 0.035, it is classified into the top pressure flow verification set.
[0091] Step S434: Calculate the average absolute difference between the measured flow and the predicted flow for each sample in the normal flow validation set and the top backflow validation set, respectively, to obtain the average error of the normal flow and the average error of the top backflow. In one embodiment, after the operating conditions of the validation samples are divided, the model prediction error is calculated for the normal flow validation set and the undercurrent flow validation set, respectively. Specifically, for each type of validation set, the absolute difference between the measured flow and the predicted flow for each sample is calculated, and the absolute differences of all samples in that type of validation set are averaged to obtain the average error for the normal flow condition and the average error for the undercurrent flow condition, respectively.
[0092] For example, in the top-down flow condition verification set, if there are several samples and their absolute prediction errors are several values, then the average of these values is taken as the average error of the top-down flow condition.
[0093] Step S435: Calculate the average absolute difference between the predicted flow increment and the measured flow increment for each sample in the validation set to obtain the flow continuity error.
[0094] In one embodiment, to evaluate the predictive smoothness and continuity of the model in the time dimension, a flow continuity error is further calculated within the validation set. Specifically, for samples at adjacent time points in the validation set, the change in predicted flow between adjacent time points and the change in measured flow between adjacent time points are calculated respectively, and the absolute value of the difference between the two is taken. Then, the difference corresponding to all validation samples is averaged to obtain the flow continuity error.
[0095] For example, if the predicted flow rate increases by 5 at a certain moment compared to the previous moment, but the actual flow rate only increases by 3, then the continuity error corresponding to that moment is 2.
[0096] Of particular importance, step S43 also includes the following steps: Step S436: Multiply the average error under normal flow conditions by 0.25, the average error under top backflow conditions by 0.65, and the flow continuity error by 0.10, and add the three together to obtain the fitness function value of the particle; In one embodiment, after obtaining the average error under normal flow conditions, the average error under undercurrent flow conditions, and the flow continuity error, the three types of errors are weighted and summed according to preset weights to construct the fitness function value of the particle. Specifically, the average error under normal flow conditions is multiplied by 0.25, the average error under undercurrent flow conditions is multiplied by 0.65, and the flow continuity error is multiplied by 0.10. The sum of these three values yields the fitness function value corresponding to the particle, which is used to comprehensively evaluate the overall performance of the model under this hyperparameter combination.
[0097] For example, when the three types of errors of a certain particle are 2, 3 and 1 respectively, its fitness function value is 2×0.25+3×0.65+1×0.10.
[0098] Step S437: Update the particle's velocity and position based on the fitness function value, where the particle's position represents the candidate hyperparameter value and the particle's velocity represents the adjustment amount of the hyperparameter. In one embodiment, after obtaining the fitness function value of each particle, the particle's velocity and position are updated according to the update rules of the particle swarm optimization algorithm. The particle's position corresponds to a set of candidate hyperparameter values, and the particle's velocity represents the adjustment magnitude of the hyperparameters in the current iteration relative to the previous round. During the update process, the particle comprehensively considers its own historical best position and the group's historical best position to adjust its current position, thereby guiding the particle to gradually move towards a parameter region with better fitness.
[0099] For example, if the fitness of a particle’s current parameter combination is better than its historical record, then its individual best position is updated and it moves toward that direction in the next iteration.
[0100] Step S438: Preset the maximum number of iterations. When the maximum number of iterations is reached, output the position of the particle with the smallest fitness function value as the optimal hyperparameter combination. In one embodiment, a maximum number of iterations is preset in the particle swarm optimization (PSO) algorithm. When the number of iterations reaches this maximum value, the particle update process stops, and the particle with the smallest fitness function value is selected from the current particle swarm. The position corresponding to this particle is the optimal hyperparameter combination. This method avoids infinite iteration of the PSO algorithm while ensuring that the search process is completed within a reasonable time.
[0101] For example, when the maximum number of iterations is set to 50, the hyperparameter combination corresponding to the globally optimal particle is selected after the 50th iteration.
[0102] Step S439: Using the optimal hyperparameter combination, the weight vector and bias term to be solved are obtained from the objective function, thus forming the trained model.
[0103] In one embodiment, after determining the optimal hyperparameter combination, this hyperparameter combination is substituted into the aforementioned constructed support vector machine objective function to finally solve the model, obtaining the corresponding weight vector and bias term, thereby forming a trained traffic prediction model. This model can be directly used for subsequent traffic prediction tasks and maintains high prediction accuracy and stability under mixed conditions of top-support and non-top-support operation.
[0104] For example, by retraining the support vector machine model using the optimal penalty coefficient and insensitivity coefficient, a prediction model can be obtained for final engineering applications.
[0105] Furthermore, the formula for the objective function in step S42 is as follows: In the formula, For training weights of the samples, The penalty coefficient for data fitting. As slack variables, This is the physical constraint penalty coefficient. To predict flow for the model, For water level, This is the water level-discharge constraint coefficient. This represents the actual traffic volume. The insensitivity coefficient.
[0106] In one embodiment, after completing the sample training weights and water level and flow rate constraint coefficient After the calculations, both are incorporated into the objective function of the support vector machine regression model to form a joint optimization objective that takes into account data fitting accuracy, the variability of top support conditions, and the physical continuity of water level and flow rate. Specifically, the following objective function is constructed, and the model parameters are solved by minimizing this objective function: The objective function consists of three parts: the first part is the L2 norm squared term of the weight vector, which is used to constrain the model complexity; the second part is the weighted empirical risk term, which is used to measure the model's fitting error to the training samples; and the third part is the physical constraint term for the consistency of water level-flow change, which is used to ensure the physical rationality of the model's prediction results in the time series.
[0107] The first term in the objective function is half the squared L2 norm of the weight vector w. This term constrains the complexity of the support vector machine model and prevents overfitting during training. During model solving, minimizing this term makes the regression function mapped to the high-dimensional feature space as smooth as possible, thereby improving the model's generalization ability to unknown samples.
[0108] For example, this term penalizes the increase in weights when the model attempts to fit a small number of outliers with excessively large weight values, thereby suppressing the disorderly growth of model complexity.
[0109] The second term of the objective function is a weighted empirical risk term with sample training weights, where The penalty coefficient for data fitting. Let i be the training weights corresponding to the i-th sample. and This is a slack variable introduced in support vector machine regression. It measures the degree of deviation between the model's predicted flow and the actual flow beyond the insensitivity interval ε, and distinguishes the importance of different samples through sample training weights.
[0110] During model training, when a training sample has a large value of top-supporting degree, its corresponding sample training weight λᵢ is larger, which makes the error penalty of the sample in the objective function increase accordingly, thereby prompting the model to pay more attention to the prediction accuracy under top-supporting flow conditions.
[0111] For example, for the same prediction error, samples with higher support will have a larger penalty value in this term, which in turn affects the direction of model parameter updates.
[0112] The third term of the objective function is the water level-flow physical constraint term, where This is the physical constraint penalty coefficient. and , respectively, are the predicted values of the flow rate by the model at adjacent times i and i−1. and These are the water level values at the corresponding times. This is the water level-discharge constraint coefficient obtained from the aforementioned calculation. This term constrains the relationship between the predicted change in flow rate and the change in water level to maintain an approximately linear proportional relationship, ensuring that the model's prediction results conform to the hydrophysical laws under non-backwater conditions.
[0113] During model training, when the predicted flow rate trend deviates significantly from the flow rate trend derived from water level changes, this term will generate a large penalty value, thereby guiding the model parameters to adjust in a direction that satisfies physical consistency.
[0114] For example, this term will increase significantly when the water level rises at adjacent times, while the model predicts a significant decrease in flow, thereby suppressing such predictions that do not conform to hydrological mechanisms.
[0115] While minimizing the objective function, standard inequality constraints are introduced into the support vector machine regression model to define the insensitive interval for prediction error. These constraints are as follows: as well as To ensure that the prediction error is within Within a certain range, the loss function is not considered; however, when the error exceeds a certain range... At that time, through slack variables and Punishment should be imposed; at the same time, it is required and All values are non-negative to ensure the rationality of the constraints.
[0116] For example, when the error between the predicted traffic and the actual traffic of a training sample is less than ε, the sample will not contribute to the empirical risk term.
[0117] Based on a comprehensive consideration of model complexity control, data fitting error weighted according to different operating conditions, and the physical continuity constraint of water level-flow rate, the weight vector w and bias term b in the support vector machine model are jointly solved by minimizing the above objective function, thereby obtaining a flow prediction model that balances prediction accuracy and physical rationality. This objective function then serves as the core calculation basis for fitness evaluation in the particle swarm optimization algorithm, guiding the global optimization search of hyperparameters.
[0118] Furthermore, step S5 includes the following steps: Step S51: Input the real-time collected ultrasonic time-of-flight flow velocity, test section water level, hydraulic radius, flow area, upstream water level, and downstream water level data into the training model. The model outputs the current flow prediction value as the predicted flow. In one embodiment, after the model completes offline training and enters the online operation phase, multi-source hydrological monitoring data at the monitoring station cross-section are collected in real time, and the data is organized into a model input vector according to the feature order and data format of the training phase. The real-time collected data includes hydrological features such as ultrasonic time-of-flight velocity, water level at the test cross-section, hydraulic radius, flow area, water levels at upstream stations, and water levels at downstream stations. This input vector is input into the trained support vector machine model, which outputs the predicted flow rate at the current moment based on the learned nonlinear mapping relationship, and this output is used as the predicted flow rate at the current moment.
[0119] For example, at a certain real-time operating moment, after the system receives monitoring values such as cross-sectional flow velocity, water level, and upstream and downstream water levels, it combines them into a feature vector and inputs it into the model. The model then outputs a corresponding flow rate value as the prediction result for that moment.
[0120] Step S52: Calculate the difference between the water level at the current test section and the water level at the previous test section, and use it as the water level change. In one embodiment, to characterize the changing trend of the current hydrological process, the water level value of the current test section is extracted from real-time water level data and compared with the water level value of the test section stored at the previous time. The difference between the two is used to calculate the water level change. The water level change is used to reflect the magnitude of water level rise and fall at the station section within adjacent time steps.
[0121] For example, if the water level at the current test section is 5.20 m and the water level at the previous time was 5.15 m, then the water level change is 0.05 m, indicating that the water level is rising.
[0122] Step S53: Set the minimum filter coefficient to 0.3 and the maximum filter coefficient to 0.7, input the water level change into the Sigmoid function for mapping, and obtain the dynamic filter coefficient within the range of the minimum filter coefficient to the maximum filter coefficient. In one embodiment, to adaptively adjust the smoothing intensity of the predicted flow rate according to the degree of water level change, a minimum filter coefficient and a maximum filter coefficient are preset, wherein the minimum filter coefficient is 0.3 and the maximum filter coefficient is 0.7. The water level change obtained in step S52 is input into the Sigmoid function for nonlinear mapping, so that the output result falls between the minimum filter coefficient and the maximum filter coefficient, thereby obtaining the dynamic filter coefficient corresponding to the current time.
[0123] For example, when the water level change is small, the output value of the Sigmoid function approaches the minimum filter coefficient, making the filtering process smoother; when the water level change is large, the output value gradually approaches the maximum filter coefficient, making the predicted flow rate more sensitive to the current change.
[0124] Step S54: Multiply the predicted flow rate by the dynamic filtering coefficient, add the result of multiplying the filtered flow rate of the previous moment by one minus the dynamic filtering coefficient, and obtain the final flow rate of the current moment and output it.
[0125] In one embodiment, the predicted flow rate at the current moment and the filtered flow rate at the previous moment are weighted and fused according to the dynamic filtering coefficient. Specifically, the predicted flow rate at the current moment is multiplied by the dynamic filtering coefficient, and the filtered flow rate at the previous moment is multiplied by one and the dynamic filtering coefficient is subtracted. The two are then added together to obtain the final output flow rate at the current moment, and this final output flow rate is output as the real-time streaming result.
[0126] For example, when the dynamic filtering coefficient is 0.6, the final output flow is determined by 60% of the current predicted flow and 40% of the output flow from the previous moment, thereby suppressing short-term fluctuations while ensuring response speed.
[0127] Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the application are intended to be included within the invention.
[0128] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.
Claims
1. A method for constructing a real-time streaming model based on machine learning, characterized in that, Includes the following steps: Step S1: Obtain the historical raw feature dataset and raw traffic dataset; The original feature dataset and the original traffic dataset are subjected to unified data quality processing to eliminate abnormal disturbances and restore time continuity, resulting in the feature dataset and the traffic dataset. Step S2: Identify top-load flow condition samples in the feature dataset based on the flow dataset and calculate their proportion. Calculate the importance score of each feature in the feature dataset based on the proportion. The feature dataset is filtered based on its importance score to form a filtered feature set; Step S3: Calculate the backwater level value based on the filtered feature set and determine the sample training weights. Extract the flow rate value corresponding to the non-backwater condition from the flow data set and calculate the water level flow constraint coefficient. Step S4: Using the filtered feature set as input and the flow dataset as target, establish a support vector machine model that integrates the training weights of the fused samples and the water level and flow constraint coefficients. Optimize the model parameters using a particle swarm optimization algorithm with weighted fitness for different working conditions to obtain the trained model. Step S5: Input the real-time collected hydrological features into the trained model to obtain the predicted flow rate, and adjust the filtering coefficient according to the real-time water level change to filter the predicted flow rate to obtain the final flow rate and output it.
2. The method for constructing a real-time streaming model based on machine learning according to claim 1, characterized in that, Step S1 includes the following steps: Step S11: Obtain the ultrasonic time-of-flight flow velocity, test section water level, hydraulic radius, flow area, upstream water level and downstream water level of the station over the years to form the original feature dataset; Step S12: Obtain flow data from historical hydrological data of the station and interpolate non-hourly times to form the original flow dataset; Step S13: For each data point in the original feature dataset and the original traffic dataset, calculate the mean and standard deviation of the 10 data points before and after it. If the absolute difference between the data point and the mean is greater than three times the standard deviation, the data point is removed.
3. The method for constructing a real-time streaming model based on machine learning according to claim 2, characterized in that, Step S1 also includes the following steps: Step S14: For the data after removing outliers, linear interpolation is used when the missing period is no more than three hours, and cubic spline interpolation is used when the missing period is more than three hours. Step S15: Smooth the interpolated data using a moving average method with a window radius of 3 to obtain the feature dataset and the traffic dataset.
4. The method for constructing a real-time streaming model based on machine learning according to claim 3, characterized in that, Step S2 includes the following steps: Step S21: When the downstream water level in front of the dam is lower than the preset threshold, it is determined to be a non-backflow state, and the corresponding sample in the flow data is determined to be a non-backflow flow sample. Step S22: When the downstream water level in front of the dam is not lower than the preset threshold, it is determined to be in a backwater state. The natural flow value of the sample is calculated based on the upstream inflow and the runoff in the interval during the same period. Step S23: Extract the measured flow value of each sample from the flow dataset, calculate the difference between the measured flow value and the natural flow value, and divide by the natural flow value to obtain the flow deviation ratio; Step S24: When the station belongs to a major river station, the flow deviation ratio is less than or equal to -6. This sample was identified as a top-support flow condition. Step S25: When the station is a medium-sized river station, the flow deviation ratio is less than or equal to -10. This sample was identified as a top-support flow condition. Step S26: When the station is a small river station, the flow deviation ratio is less than or equal to -15. This sample was identified as a top-support flow condition. Step S27: Count the total number of top-support flow condition samples and divide by the total number of samples to obtain the percentage value.
5. The method for constructing a real-time streaming model based on machine learning according to claim 4, characterized in that, Step S2 also includes the following steps: Step S28: Establish a mapping relationship between each feature in the feature dataset and the traffic dataset, calculate the information gain score of the feature, and calculate the absolute value of the Pearson correlation coefficient between the feature and the traffic dataset as the correlation coefficient score. Step S29: Add the product of the information gain score and the proportion value to the product of the correlation coefficient score and the proportion value minus one, to obtain the importance score; Step S210: Set the ultrasonic time difference flow velocity feature, test section water level feature, and downstream first station water level feature in the feature dataset as mandatory retention features. Select several features from the remaining features according to their importance scores and merge them with the mandatory retention features to form a filtered feature set.
6. The method for constructing a real-time streaming model based on machine learning according to claim 5, characterized in that, Step S3, which calculates the support level value based on the filtered feature set and determines the sample training weights, includes: The downstream first station water level value and the test section water level value of each training sample are extracted from the filtered feature set; The difference between the water level at the first downstream station and the water level at the test section is calculated and divided by the pre-obtained water level difference between the two stations under natural conditions to obtain the backwater level value. The preset value range of the top support sensitivity coefficient is 0.5-2.
0. The product of the top support degree value and the top support sensitivity coefficient plus one is used as the sample training weight.
7. The method for constructing a real-time streaming model based on machine learning according to claim 6, characterized in that, Step S3, which involves extracting the flow rate values corresponding to the non-top-back condition from the flow rate dataset and calculating the water level-flow constraint coefficient, includes: Samples with a support level value less than 0.8 from all training samples are selected and marked as samples with weak support influence. Extract the flow value corresponding to the weak backwater impact sample from the flow dataset, calculate the difference between the flow at the current time and the flow at the previous time to obtain the flow change, and calculate the difference between the water level at the current time and the water level at the previous time to obtain the water level change. The water level-discharge constraint coefficient is obtained by summing the products of the flow rate changes and water level changes of all samples affected by weak backwater and dividing them by the sum of the squares of the water level changes.
8. The method for constructing a real-time streaming model based on machine learning according to claim 7, characterized in that, Step S4 includes the following steps: Step S41: Divide the filtered feature set into a training set and a validation set in a 7:3 ratio. Use the training set as the model input and the corresponding traffic dataset as the model training target to build a support vector machine model, where the support vector machine model is represented as: In the formula, For the filtered feature set, A nonlinear mapping function maps the input to a high-dimensional space. Let be the weight vector to be solved. The bias term to be solved. The flow rate predicted by the model. This is the matrix transpose symbol; Step S42: Construct an objective function that integrates the training weights of the sample samples and the water level and flow constraint coefficients in the support vector machine model; Step S43: Solve the objective function using the particle swarm optimization algorithm with weighted fitness for different work conditions, and substitute the optimal hyperparameter combination obtained from the search into the support vector machine model to obtain the optimal model parameters, thus obtaining the trained model.
9. The method for constructing a real-time streaming model based on machine learning according to claim 8, characterized in that, The formula for the objective function in step S42 is as follows: In the formula, For training weights of the samples, The penalty coefficient for data fitting. As slack variables, This is the physical constraint penalty coefficient. To predict flow for the model, For water level, This is the water level-discharge constraint coefficient. This represents the actual traffic volume. The insensitivity coefficient.
10. The method for constructing a real-time streaming model based on machine learning according to claim 9, characterized in that, Step S5 includes the following steps: Step S51: Input the real-time collected ultrasonic time-of-flight flow velocity, test section water level, hydraulic radius, flow area, upstream water level, and downstream water level data into the training model. The model outputs the current flow prediction value as the predicted flow. Step S52: Calculate the difference between the water level at the current test section and the water level at the previous test section, and use it as the water level change. Step S53: Set the minimum filter coefficient to 0.3 and the maximum filter coefficient to 0.7, input the water level change into the Sigmoid function for mapping, and obtain the dynamic filter coefficient within the range of the minimum filter coefficient to the maximum filter coefficient. Step S54: Multiply the predicted flow rate by the dynamic filtering coefficient, add the result of multiplying the filtered flow rate of the previous moment by one minus the dynamic filtering coefficient, and obtain the final flow rate of the current moment and output it.