Method and system for creating false tide prediction model based on river and lake confluence water level
By collecting and processing hydrological data, extracting significant factors, and constructing a Logistic regression model, the problem of unstable false tide identification in the confluence of rivers and lakes was solved, enabling accurate and rapid prediction of false tides and supporting flood control scheduling and water resource management.
Patent Information
- Application Number
- CN202511339926.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2025-10-28
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies struggle to identify and predict false tides in the confluence of rivers and lakes in a timely manner. Traditional methods are inefficient and susceptible to measurement errors and extreme weather conditions, leading to unstable identification results.
By collecting hydrological data, calculating continuous water level fluctuations, extracting significant factors and inputting them into a binary logistic regression model, a prediction model is generated. Combined with smoothing processing, this enables accurate identification and rapid prediction of false tides.
It enables timely and reliable prediction of false tides on a minute-level timescale, improving the stability and applicability of the judgment and providing timely support for flood control scheduling and water resource allocation.
Smart Images

Figure CN120850252A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of false tide prediction technology, and in particular to a method and system for creating a false tide prediction model based on the water level of rivers and lakes confluence. Background Technology
[0002] In the hydrological processes of river-lake confluence areas, a phenomenon known as "false tides" often occurs due to the combined effects of upstream water flow, changes in water level differences, tides, wind, and waves. False tides are not directly caused by actual tides but are triggered by complex hydrodynamic conditions. Their water level fluctuations are similar to those of real tides, easily interfering with water level analysis and flood control scheduling decisions. In existing hydrological monitoring systems, water level data collection relies on fixed hydrological observation stations, typically collecting water level sequences on a minute or hourly basis, and then averaging them daily for trend analysis. However, in river-lake confluence areas, this daily average-based analysis method struggles to identify false tides in a timely manner because they are short-lived, frequently changing, and exhibit subtle and irregular amplitudes. Furthermore, traditional empirical methods rely heavily on manual analysis of water level curve morphology or single threshold judgments, which are not only inefficient but also susceptible to measurement errors, inconsistent sampling periods, and extreme weather conditions, leading to unstable false tide identification results. Summary of the Invention
[0003] Therefore, it is necessary for the present invention to provide a method and system for creating a false tide prediction model based on the confluence of rivers and lakes, in order to solve at least one of the above-mentioned technical problems.
[0004] To achieve the above objectives, a method for creating a false tide prediction model based on river and lake confluence water levels includes the following steps: Step S1: Collect hydrological data, calculate continuous water level fluctuations, and determine fluctuation discrimination data; convert the fluctuation discrimination data into a false tide binary classification variable; Step S2: Extract the initial factors affecting the water level of the confluence of rivers and lakes, perform correlation tests in combination with fluctuation discrimination data, eliminate non-significant factors, and form a set of significant factors; Step S3: Input the significant factor set and the pseudo-tide binary categorical variable into the binary logistic regression model for training to generate a prediction model; Step S4: Calculate the water level difference between Jianli Hydrological Station and the target water level station, and use the prediction model to determine the false tide judgment result; obtain the water level sequence of the target water level station, and perform smoothing processing to generate smooth water level data; determine the real-time false tide prediction result based on the false tide judgment result and the smooth water level data.
[0005] Preferably, the present invention also provides a system for creating a false tide prediction model based on the confluence of rivers and lakes, used to execute the above-described method for creating a false tide prediction model based on the confluence of rivers and lakes, wherein the system for creating a false tide prediction model based on the confluence of rivers and lakes includes: The false tide label generation module is used to collect hydrological data, calculate continuous water level fluctuations, determine fluctuation discrimination data, and convert the fluctuation discrimination data into false tide binary classification variables. The significant factor screening module is used to extract preliminary factors that affect the water level of the confluence of rivers and lakes, perform correlation tests in combination with fluctuation discrimination data, eliminate non-significant factors, and form a significant factor set. The prediction model building module is used to input the significant factor set and the pseudo-tide binary categorical variable into the binary logistic regression model for training and to generate the prediction model. The real-time false tide prediction module is used to calculate the water level difference between Jianli Hydrological Station and the target water level station, determine the false tide judgment result using the prediction model, obtain the water level sequence of the target water level station, and perform smoothing processing to generate smooth water level data; and determine the real-time false tide prediction result based on the false tide judgment result and the smooth water level data.
[0006] This invention enables accurate identification and rapid prediction of false tides in the complex hydrodynamic environment of river and lake confluence areas. By introducing continuous water level fluctuation characteristics to quantify the false tide state, it effectively avoids missed and false judgments caused by a single water level threshold. The combination of multi-factor screening and significance testing reduces the interference of redundant data on the analysis results, improving the stability and applicability of the judgment. Mathematical calculation of false tide probability gives the prediction results clear numerical interpretation, making it easy to apply directly to hydrological assessment. Real-time updates and noise suppression processing can reduce the impact of short-term abnormal fluctuations while ensuring the sensitivity of the judgment, thus providing more timely and reliable data support for flood control scheduling, shipping management, and water resource allocation on a minute-level time scale. Attached Figure Description
[0007] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is a schematic diagram illustrating the steps of creating a false tide prediction model based on the confluence of rivers and lakes according to the present invention. Figure 2 This is a water level fluctuation diagram from an embodiment of the present invention; Figure 3 This is a box plot of daily water level fluctuation values in an embodiment of the present invention; Figure 4 This is a graph showing the relationship between the proportion of "false tides" occurring during high water levels at Jianli Station in an embodiment of the present invention and the difference in water level. Figure 5This is a box plot of the error values of the Fast Fourier Transform method and the Local Weighted Regression algorithm in the embodiments of the present invention. Detailed Implementation
[0008] The technical method of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0009] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.
[0010] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0011] To achieve the above objectives, please refer to Figures 1 to 5 This invention provides a method for creating a false tide prediction model based on the confluence of rivers and lakes, the method comprising the following steps: Step S1: Collect hydrological data, calculate continuous water level fluctuations, and determine fluctuation discrimination data; convert the fluctuation discrimination data into a false tide binary classification variable; Step S2: Extract the initial factors affecting the water level of the confluence of rivers and lakes, perform correlation tests in combination with fluctuation discrimination data, eliminate non-significant factors, and form a set of significant factors; Step S3: Input the significant factor set and the pseudo-tide binary categorical variable into the binary logistic regression model for training to generate a prediction model; Step S4: Calculate the water level difference between Jianli Hydrological Station and the target water level station, and use the prediction model to determine the false tide judgment result; obtain the water level sequence of the target water level station, and perform smoothing processing to generate smooth water level data; determine the real-time false tide prediction result based on the false tide judgment result and the smooth water level data.
[0012] Preferably, step S1 includes the following steps: Step S11: Call the water level recording equipment and flow monitoring device deployed at the target water level station and upstream and downstream related hydrological stations in the confluence of rivers and lakes to collect continuous high water period raw hydrological data at 5-minute intervals to form raw dataset; Step S12: Extract the continuous water level sequence for 24 hours from the original dataset, calculate the water level fluctuation for each time period, sort them by the magnitude of the fluctuation, and extract the fluctuation value corresponding to the 10th percentile as the water level fluctuation discrimination data for the day. Step S13: Compare the fluctuation discrimination data with the preset fluctuation threshold. If the fluctuation discrimination data is greater than the preset fluctuation threshold, mark it as 1. If the fluctuation discrimination data is less than or equal to the preset fluctuation threshold, mark it as 0. Generate a classification dataset. Step S14: Assign the labeling results of the classification dataset to the pseudo-tide binary classification variable attributes to generate the pseudo-tide variable set.
[0013] In this embodiment of the invention, water level recording devices and flow monitoring devices are invoked at the target water level station and its upstream and downstream associated hydrological stations deployed in the confluence of rivers and lakes. Each monitoring point automatically records the instantaneous water level and instantaneous flow rate during the high water period at fixed intervals of 5 minutes according to a unified time control program. The data is synchronized to the data acquisition terminal via a wired transmission link. The data terminal stores the data in daily order using timestamps as index fields to form an original dataset. From the original dataset, a 24-hour water level record sequence is extracted for each day. The water level difference between two adjacent time periods is used to calculate the water level variation for each time period. All variation values are expressed as absolute values and arranged in ascending order. The variation value located at the 10th percentile in the sequence is selected as the absolute value. The data is used to identify water level fluctuations on a given day. Percentiles are calculated based on the total number of fluctuation values for that day, with corresponding amplitude values used for location determination. The fluctuation data is compared to a preset fluctuation threshold of 0.05m. If the fluctuation data is strictly greater than 0.05m, a value of 1 is assigned to the corresponding time period in the classification dataset; otherwise, a value of 0 is assigned. The classification dataset is structured with daily units, and fields representing time periods and classification labels. The labeling results are set as binary categorical variables for false tides, with a value of 1 indicating the presence of false tides and a value of 0 indicating the absence of false tides. This variable is stored in the data storage structure along with the date and station code, forming a false tide variable set.
[0014] Preferably, step S2 includes the following steps: Step S21: Call the daily average flow monitoring devices and water level recording devices of the upstream station, the target water level station and the downstream station to collect the daily average flow and daily average water level information for the corresponding time period to form a station dataset; Step S22: Calculate the flow ratio of each station, the water level difference between adjacent stations, and the daily rise and fall rate of each station using the station dataset, and summarize them into a preliminary factor set; Step S23: Perform a Pearson correlation test on the initial factor set and volatility discrimination data; Step S24: Remove non-significant factors with a two-sided significance level greater than or equal to the threshold, and retain significant factors to form a significant factor set.
[0015] In this embodiment of the invention, daily average flow monitoring devices and water level recording devices deployed at upstream stations, target water level stations, and downstream stations in the confluence of rivers and lakes are invoked. The monitoring devices at each station collect instantaneous flow and water level values at 5-minute intervals throughout the daily collection period according to a unified time scheduling program. The daily average flow and water level are calculated using a weighted average method. After aggregation, a station dataset is constructed based on station code, date, and measurement index fields. Based on the station dataset, the flow percentage of each station is calculated by dividing the daily average flow value of each station by the sum of the daily average flow values of the three stations on the same day, retaining four decimal places. Simultaneously, the water level difference between adjacent stations is calculated by the daily average water level difference between the upstream and downstream stations, the upstream and target water level stations, and the target and downstream stations, expressed as absolute values. Finally, the daily rise and fall rate of each station is calculated by (the daily rise and fall rate of the station on that day). The daily average water level value minus the previous day's daily average water level value is divided by the previous day's daily average water level value and expressed as a percentage. All calculation results are merged by station and date to form a preliminary factor set containing three types of indicators: flow rate ratio, water level difference, and daily fluctuation rate. Each indicator value sequence in the preliminary factor set is paired with the fluctuation discrimination data sequence extracted in the previous step. After matching by day, the Pearson correlation coefficient r value and its two-sided significance level p value are calculated in the bivariate correlation test tool. The sample size is no less than 30 days to ensure statistical stability. Indicators with a two-sided significance level p value greater than or equal to 0.05 are screened out from the test results. Indicators with a p value less than 0.05 are identified as significant factors and included in the significant factor set. The significant factor set is stored in tabular form, including factor name, correlation coefficient r value, significance level p value, and corresponding station and date index.
[0016] Preferably, step S3 includes the following steps: Step S31: Set the dependent variable of the binary logistic regression model to a pseudo-tide binary categorical variable and the independent variable to a set of significant factors; Step S32: Calculate the regression coefficients and intercepts of each independent variable to generate regression coefficients; Step S33: Verify the goodness of fit of the regression model, and output the significance level value and verification conclusion; Step S34: Based on the significance level and the verification conclusion, select regression models that meet the preset requirements for goodness of fit, and construct a prediction model.
[0017] In this embodiment of the invention, the dependent variable in the regression analysis is set as a binary categorical variable for false tides. This variable is derived from the false tide variable set, with a value of 1 indicating the presence of false tides and a value of 0 indicating the absence of false tides. The independent variable is set as a significant factor set. Factor values are matched one by one with the dependent variable according to the correspondence between date and station, ensuring that each observation sample contains a complete set of factor values and corresponding false tide markers. A numerical calculation tool is invoked in the data processing terminal to estimate the parameters of the sample data according to the formal expression of binary logistic regression. The maximum likelihood estimation method is used to calculate the regression coefficient and intercept value of each independent variable. The coefficient results are retained to four decimal places and recorded uniformly in the regression parameter table. The fields of this table include the independent variable name, regression coefficient value, and intercept value. The regression parameters are calculated and substituted into the binary logistic regression expression to generate predicted values for each sample. The overall significance level p-value of the model is calculated using the likelihood ratio test. Simultaneously, the Hosmer-Lemeshow test is used to determine the goodness of fit of the model. The test results include the significance level, chi-square statistic, and degrees of freedom. The verification conclusion is recorded in either "pass" or "fail" label format. From the test results, regression parameter combinations with a significance level p-value less than 0.05 and a Hosmer-Lemeshow test conclusion of "pass" are selected and marked as regression results that meet the preset goodness of fit requirements. This result, along with its corresponding regression coefficients and intercept values, is stored as a predicted parameter set.
[0018] Preferably, the formulas for calculating the regression coefficients and intercepts of each independent variable in step S32 are as follows: ; in, This represents the probability of a false tide occurring, with a value ranging from 0 to 1. The intercept of the regression equation represents the sum of all independent variables. When the value is 0, it represents the baseline value of the logarithmic probability. The number of independent variables represents the number of factors in the significant factor set. For the first The regression coefficients of the independent variables represent the logarithmic change in the ratio of the probability of a false tide occurring to the probability of it not occurring when the factor increases by one unit. For the first The numerical values of the independent variables represent the specific observed values of the significant factor set.
[0019] In this embodiment of the invention, the regression coefficients and intercepts of each independent variable are calculated numerically using the formalized expression of binary logistic regression, the mathematical form of which is: ,in This represents the probability of a false tide occurring, and its value is limited to 0 to 1. The intercept of the regression equation is the sum of all independent variables in the significant factor set. When the value is 0, The baseline value of the natural logarithm representing the ratio of the probability of a false tide occurring to the probability of it not occurring under this condition; This represents the number of factors in a significant factor set, corresponding to the number of variables in that significant factor set. Indicates the The regression coefficients of the independent variables, in units of log odds / unit value of the independent variable, reflect the logarithmic increase or decrease in the probability of false tides occurring versus not occurring when the value of the factor increases by one unit. Indicates the The values of each independent variable in the observed sample correspond to the actual observation results of the significant factor set under specific date and station conditions. During the calculation, the dependent variable and independent variable are first paired one by one according to the observed sample to establish a set containing... Initial value estimation Initial value and The initial value calculation matrix is then used to calculate the parameters using the maximum likelihood principle. and The solution is obtained through iterative steps, and each iteration updates the solution. and The estimated value is obtained until the log-likelihood function reaches the convergence condition, which is limited to the absolute value of the parameter change between two consecutive iterations being less than 1. Ultimately, the converged... With each Numerical values are recorded in the regression parameter table and stored according to fields such as independent variable name, regression coefficient, intercept, corresponding unit, and significance level.
[0020] Preferably, step S33, verifying the goodness of fit of the regression model, includes: The regression models and their corresponding regression coefficients were divided into 10 groups; Within each group, the square of the difference between the observed frequency and the predicted frequency is calculated and divided by the predicted frequency. The results of each group are then summed to generate statistical data. The significance level value is calculated by comparing the statistical data with the degrees of freedom. Compare the significance level value with the set threshold to determine the verification conclusion; When the significance level is greater than or equal to the set threshold, the corresponding regression model is identified as having a good fit. When the significance level is less than a set threshold, the corresponding regression model is identified as underfitting.
[0021] In this embodiment of the invention, the regression expression and its corresponding regression coefficients are first divided into 10 groups according to the ascending order of predicted probability values. The number of observed samples in each group is kept as balanced as possible to reduce the impact of uneven distribution on the statistical test. Within each group, the observed frequency of false tides and the predicted frequency of false tides calculated based on the regression coefficients are counted. The square of (observed frequency minus predicted frequency) is divided by the predicted frequency to obtain the deviation ratio of that group. The deviation ratios of the 10 groups are accumulated sequentially to form the statistical data corresponding to the regression result. Then, this statistical data is combined with the number of groups minus the regression coefficients. The degrees of freedom, obtained by subtracting 1 from the number of process parameters, are paired for comparison. The cumulative distribution function of the chi-square distribution is used to calculate the significance level. The degrees of freedom must be greater than zero for subsequent tests. The calculated significance level is compared with a preset threshold of 0.05. When the significance level is greater than or equal to 0.05, the corresponding regression result is marked as "good fit," indicating that the difference between the predicted probability and the actual observed distribution is not statistically significant. When the significance level is less than 0.05, the corresponding regression result is marked as "underfit," indicating that the difference between the predicted probability and the actual observed distribution is statistically significant.
[0022] Preferably, step S34, which involves selecting regression models that meet preset requirements for goodness of fit based on significance level values and validation conclusions, includes: Compare the significance level value with the set threshold. When the significance level value is greater than or equal to the set threshold, the goodness of fit of the regression model is marked as satisfactory; otherwise, it is marked as unsatisfactory. Extract the regression coefficients and intercept parameters of the regression model with a good fit, and denote them as parameter data; The parameter data and model structure information are encapsulated into a callable format to generate a prediction model.
[0023] In this embodiment of the invention, firstly, the goodness-of-fit test result table is called to read the significance level value and verification conclusion corresponding to each group of regression results. The significance level value is compared with a preset threshold of 0.05. When the significance level value is greater than or equal to 0.05 and the verification conclusion is "good fit", the goodness-of-fit of the regression results is marked as "meets the standard". When the significance level value is less than 0.05 or the verification conclusion is "underfit", the goodness-of-fit is marked as "not meeting the standard". Subsequently, the corresponding regression coefficients and intercept values are extracted from the records marked as "meets the standard" to form a parameter dataset. The fields of the dataset include independent variable names, regression coefficient values, intercept values, parameter units, and applicable station ranges. The above parameter dataset and the structured description of the regression equation are written into a callable storage format. This format includes a fixed expression structure, parameter index order, and calling interface description, and is stored as a file in the parameter library of the prediction calculation terminal. Finally, the resulting callable format record is the prediction calculation unit of the river and lake confluence water level false tide prediction method. It can directly reference its regression coefficients and intercept parameters in subsequent applications, and calculate the probability of false tide occurrence by inputting the observed values of the significant factor set.
[0024] Preferably, step S4, which uses a prediction model to determine the false tide determination result, includes: Input the water level difference value into the prediction model to calculate the probability of false tides; Compare the probability of false tide occurrence with the threshold of the dividing point, and output the false tide status indicator; When the water level difference is greater than 2.5 meters, a false tide is indicated as occurring. When the water level difference is greater than 2.75 meters, a false tide is indicated as inevitable.
[0025] In this embodiment of the invention, the prediction calculation unit constructed and stored in the parameter library in step S34 is first invoked, and the water level difference value measured on that day is substituted into the regression expression of the calculation unit as an input variable, according to... Read the intercept in sequence and the regression coefficients of the corresponding independent variables Multiply each value by the water level difference and other significant factors, sum them up, and then add the intercept. The results are then converted into the probability of a false tide occurrence using an exponential transformation and probability conversion formula, with the probability value limited to between 0 and 1. This probability value is then compared with a set threshold. When the probability value is greater than or equal to the threshold, the false tide status is marked as "false tide occurred"; when the probability value is lower than the threshold, the false tide status is marked as "false tide did not occur". Based on this, a direct judgment condition based on water level difference is introduced. When the measured water level difference is greater than 2.5 meters, regardless of the predicted probability value, the false tide status is marked as "false tide occurred". When the measured water level difference is greater than 2.75 meters, the false tide status is directly marked as "certain to occur". This status, along with the date, station code, and the input water level difference value, is written into the false tide judgment result table, serving as the final judgment result data output by the river-lake confluence water level false tide prediction method.
[0026] Preferably, step S4, which involves obtaining the water level sequence of the target water level station and performing smoothing processing, includes: Water level sequences at the target water level station are collected at preset intervals to generate sequence data; The sequence data is imported into a Fast Fourier Transform filter to generate FFT data; Import the sequence data into a locally weighted regression filter to generate regression data; Compare the error distributions of the FFT data and the regression data, and select the set of results with the smallest absolute error value as the preferred data; The preferred data is identified as smooth water level data.
[0027] In this embodiment of the invention, the instantaneous water level value is first automatically recorded by the water level recording device of the target water level station at preset time intervals (in units of 5 minutes). The water level records within a complete acquisition cycle are then organized into water level sequence data in chronological order. The sequence data includes two fields: a timestamp and the corresponding water level value. Subsequently, the water level sequence data is imported into a data analysis terminal with Fast Fourier Transform (FFT) processing capabilities. Frequency domain filtering is used to remove high-frequency components above a set cutoff frequency, retaining only low-frequency components to eliminate short-term fluctuations, generating a water level dataset after FFT processing. Finally, the same water level sequence data is imported into a smoothing system with local weighted regression capabilities. The calculation module performs a weighted average calculation on each sampling point according to a preset window width and weight function to generate a water level dataset after local weighted regression processing. In order to judge the merits of the two smoothing methods, the difference between the FFT dataset and the regression dataset and the water level value of the original water level sequence at the same time point is calculated, and the absolute value of the difference is taken as the single-point error. The distribution characteristics of the single-point error in the whole sequence (including mean, standard deviation and maximum value) are statistically analyzed, and the dataset with the smallest mean absolute error in the whole sequence is selected as the preferred data. Finally, the preferred data is identified as smoothed water level data and stored in the smoothed water level database with a one-to-one correspondence with the corresponding timestamp.
[0028] Most importantly, the real-time false tide prediction results are determined based on the false tide determination results and smoothed water level data, including: The false tide determination results are mapped to the corresponding smoothed water level data to generate prediction labels, thus generating label data; Calculate real-time prediction probabilities using identification data; Compare the real-time predicted probability with the set threshold to determine the real-time false tide status indicator; The real-time false tide status identifier and the corresponding smoothed water level data are encapsulated into a real-time false tide prediction result.
[0029] In this embodiment of the invention, firstly, the false tide determination result table and the smoothed water level database are called. Using timestamps and station codes as matching conditions, the false tide determination results are mapped one-to-one with the corresponding smoothed water level values. A prediction identifier field is established in the mapping record. The prediction identifier uses a value of 1 to indicate the occurrence of a false tide and a value of 0 to indicate the absence of a false tide, thus forming an identifier dataset containing timestamps, station codes, smoothed water level values, and prediction identifiers. Subsequently, the real-time prediction probability is calculated using the identifier dataset in the data processing terminal. The calculation method involves statistically analyzing records with a prediction identifier of 1 within a given time window (the window length is set to 1 hour). The proportion of the number of records recorded to the total number of records is used as the real-time prediction probability within the time window, with the prediction probability value limited to between 0 and 1. After calculating the real-time prediction probability, this value is compared with a preset probability threshold of 0.5. When the real-time prediction probability is greater than or equal to the threshold, the real-time false tide status is set to "false tide occurred"; when the real-time prediction probability is less than the threshold, the real-time false tide status is set to "false tide did not occur". Finally, the real-time false tide status is combined with the smoothed water level value corresponding to the same timestamp and encapsulated into a real-time false tide prediction result record, and the record is written into the real-time prediction result table in chronological order.
[0030] Most importantly, the specific formula for calculating the real-time prediction probability is as follows: ; in, This represents the predicted probability of a false tide occurring, with a value ranging from 0 to 1. For the intercept parameter, For regression coefficients, This refers to the water level difference.
[0031] In this embodiment of the invention, the water level difference data generated in step S4 is first used as input value and introduced into the real-time prediction calculation stage. The unit of water level difference is limited to meters and accurate to two decimal places. Simultaneously, the intercept parameter α and regression coefficient β output from step S3 are introduced. Both are derived from historical hydrological data training, and their values remain fixed within the prediction period, retaining four decimal places. The real-time prediction calculation uses the formula... ,in, Taking the natural constant as 2.71828, This represents the daily average water level difference between the upstream hydrological station and the target water level station at the current moment. In the specific calculation process, firstly, electronic water level gauges and data acquisition terminals are used to synchronously collect upstream and downstream water level data at 5-minute intervals. Then, the central data processing unit calculates the upstream and downstream water level differences in daily groups, forming a water level difference dataset. Subsequently, a high-precision floating-point processor is used to correlate each group of water level differences with regression coefficients. Multiply by, then multiply by the intercept parameter The sums yield the logarithmic probability value; this logarithmic probability value is used as the exponent in the exponentiation operation within the processor, and is calculated using the built-in exponentiation function module. The exponent value is taken, 1 is added, and the reciprocal is taken. The output value is the predicted probability of a false tide. The predicted probability is limited to a value between 0 and 1, stored in double-precision floating-point format, and transmitted in real time to the hydrological analysis terminal for synchronous comparison with the smoothed water level data. When the predicted probability P is greater than or equal to 0.5, the judgment module outputs that the false tide status has occurred; otherwise, it outputs that it has not occurred. At the same time, it is displayed in real time in a graphical manner on the hydrological monitoring screen.
[0032] Preferably, the present invention also provides a system for creating a false tide prediction model based on the confluence of rivers and lakes, used to execute the above-described method for creating a false tide prediction model based on the confluence of rivers and lakes, wherein the system for creating a false tide prediction model based on the confluence of rivers and lakes includes: The false tide label generation module is used to collect hydrological data, calculate continuous water level fluctuations, determine fluctuation discrimination data, and convert the fluctuation discrimination data into a false tide binary classification variable. The significant factor screening module is used to extract preliminary factors that affect the water level of the confluence of rivers and lakes, combine them with fluctuation discrimination data to perform correlation tests, eliminate non-significant factors, and form a significant factor set. The prediction model building module is used to input the significant factor set and the pseudo-tide binary categorical variable into the binary logistic regression model for training and to generate the prediction model. The real-time false tide prediction module is used to calculate the water level difference between Jianli Hydrological Station and the target water level station, determine the false tide judgment result using the prediction model, obtain the water level sequence of the target water level station, and perform smoothing processing to generate smooth water level data; and determine the real-time false tide prediction result based on the false tide judgment result and the smooth water level data.
[0033] Therefore, the embodiments should be considered as exemplary and non-limiting in all respects, and the scope of the invention is not limited by the foregoing description. Thus, all changes falling within the meaning and scope of the equivalents of the application are intended to be included within the scope of the invention.
[0034] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.
Claims
1. A method for creating a false tide prediction model based on the confluence of rivers and lakes, characterized in that, Includes the following steps: Step S1: Collect hydrological data, calculate continuous water level fluctuations, and determine fluctuation discrimination data; Convert the fluctuation discrimination data into a binary classification variable for false tides; Step S2: Extract the initial factors affecting the water level of the confluence of rivers and lakes, perform correlation tests in combination with fluctuation discrimination data, eliminate non-significant factors, and form a set of significant factors; Step S3: Input the significant factor set and the pseudo-tide binary categorical variable into the binary logistic regression model for training to generate a prediction model; Step S4: Calculate the water level difference between Jianli Hydrological Station and the target water level station, and use the prediction model to determine the false tide judgment result; Obtain the water level sequence of the target water level station and smooth it to generate smoothed water level data; determine the real-time false tide prediction result based on the false tide determination result and the smoothed water level data.
2. The method for creating a false tide prediction model based on the confluence of rivers and lakes according to claim 1, characterized in that, Step S1 includes the following steps: Step S11: Call the water level recording equipment and flow monitoring device deployed at the target water level station and upstream and downstream related hydrological stations in the confluence of rivers and lakes to collect continuous high water period raw hydrological data at 5-minute intervals to form raw dataset; Step S12: Extract the continuous water level sequence for 24 hours from the original dataset, calculate the water level fluctuation for each time period, sort them by the magnitude of the fluctuation, and extract the fluctuation value corresponding to the 10th percentile as the water level fluctuation discrimination data for the day. Step S13: Compare the fluctuation discrimination data with the preset fluctuation threshold. If the fluctuation discrimination data is greater than the preset fluctuation threshold, mark it as 1. If the fluctuation discrimination data is less than or equal to the preset fluctuation threshold, mark it as 0. Generate a classification dataset. Step S14: Assign the labeling results of the classification dataset to the pseudo-tide binary classification variable attributes to generate the pseudo-tide variable set.
3. The method for creating a false tide prediction model based on the confluence of rivers and lakes according to claim 1, characterized in that, Step S2 includes the following steps: Step S21: Call the daily average flow monitoring devices and water level recording devices of the upstream station, the target water level station and the downstream station to collect the daily average flow and daily average water level information for the corresponding time period to form a station dataset; Step S22: Calculate the flow ratio of each station, the water level difference between adjacent stations, and the daily rise and fall rate of each station using the station dataset, and summarize them into a preliminary factor set; Step S23: Perform a Pearson correlation test on the initial factor set and volatility discrimination data; Step S24: Remove non-significant factors with a two-sided significance level greater than or equal to the threshold, and retain significant factors to form a significant factor set.
4. The method for creating a false tide prediction model based on the confluence of rivers and lakes according to claim 1, characterized in that, Step S3 includes the following steps: Step S31: Set the dependent variable of the binary logistic regression model to a pseudo-tide binary categorical variable and the independent variable to a set of significant factors; Step S32: Calculate the regression coefficients and intercepts of each independent variable to generate regression coefficients; Step S33: Verify the goodness of fit of the regression model, and output the significance level value and verification conclusion; Step S34: Based on the significance level and the verification conclusion, select regression models that meet the preset requirements for goodness of fit, and construct a prediction model.
5. The method for creating a false tide prediction model based on river and lake confluence water levels according to claim 4, characterized in that, The specific formulas for calculating the regression coefficients and intercepts of each independent variable in step S32 are as follows: ; in, This represents the probability of a false tide occurring, with a value ranging from 0 to 1. The intercept of the regression equation represents the sum of all independent variables. When the value is 0, it represents the baseline value of the logarithmic probability. The number of independent variables represents the number of factors in the significant factor set. For the first The regression coefficients of the independent variables represent the logarithmic change in the ratio of the probability of a false tide occurring to the probability of it not occurring when the factor increases by one unit. For the first The numerical values of the independent variables represent the specific observed values of the significant factor set.
6. The method for creating a false tide prediction model based on the confluence of rivers and lakes according to claim 4, characterized in that, Step S33, verifying the goodness of fit of the regression model, includes: The regression models and their corresponding regression coefficients were divided into 10 groups; Within each group, the square of the difference between the observed frequency and the predicted frequency is calculated and divided by the predicted frequency. The results of each group are then summed to generate statistical data. The significance level value is calculated by comparing the statistical data with the degrees of freedom. Compare the significance level value with the set threshold to determine the verification conclusion; When the significance level is greater than or equal to the set threshold, the corresponding regression model is identified as having a good fit. When the significance level is less than a set threshold, the corresponding regression model is identified as underfitting.
7. The method for creating a false tide prediction model based on the confluence of rivers and lakes according to claim 4, characterized in that, In step S34, regression models that meet the preset requirements for good fit are selected based on the significance level and validation results. Compare the significance level value with the set threshold. When the significance level value is greater than or equal to the set threshold, the goodness of fit of the regression model is marked as satisfactory; otherwise, it is marked as unsatisfactory. Extract the regression coefficients and intercept parameters of the regression model with a good fit, and denote them as parameter data; The parameter data and model structure information are encapsulated into a callable format to generate a prediction model.
8. The method for creating a false tide prediction model based on the confluence of rivers and lakes according to claim 1, characterized in that, Step S4, which uses a prediction model to determine the false tide determination result, includes: Input the water level difference value into the prediction model to calculate the probability of false tides; Compare the probability of false tide occurrence with the threshold of the dividing point, and output the false tide status indicator; When the water level difference is greater than 2.5 meters, a false tide is indicated as occurring. When the water level difference is greater than 2.75 meters, a false tide is indicated as inevitable.
9. The method for creating a false tide prediction model based on the confluence of rivers and lakes according to claim 1, characterized in that, Step S4 involves obtaining the water level sequence of the target water level station and performing smoothing processing, including: Water level sequences at the target water level station are collected at preset intervals to generate sequence data; The sequence data is imported into a Fast Fourier Transform filter to generate FFT data; Import the sequence data into a locally weighted regression filter to generate regression data; Compare the error distributions of the FFT data and the regression data, and select the set of results with the smallest absolute error value as the preferred data; The preferred data is identified as smooth water level data.
10. A system for creating a false tide prediction model based on river and lake confluence water levels, characterized in that, The system for creating a false tide prediction model based on river and lake confluence water levels, as described in claim 1, is used to execute the method for creating such a model as described in claim 1. The false tide label generation module is used to collect hydrological data, calculate continuous water level fluctuations, determine fluctuation discrimination data, and convert the fluctuation discrimination data into a false tide binary classification variable. The significant factor screening module is used to extract preliminary factors that affect the water level of the confluence of rivers and lakes, combine them with fluctuation discrimination data to perform correlation tests, eliminate non-significant factors, and form a significant factor set. The prediction model building module is used to input the significant factor set and the pseudo-tide binary categorical variable into the binary logistic regression model for training and to generate the prediction model. The real-time false tide prediction module is used to calculate the water level difference between Jianli Hydrological Station and the target water level station, determine the false tide judgment result using the prediction model, obtain the water level sequence of the target water level station, and perform smoothing processing to generate smooth water level data; and determine the real-time false tide prediction result based on the false tide judgment result and the smooth water level data.
Citation Information
Patent Citations
Feature selection decomposing method applied to river water level forecasting data
CN107992447A
Real-time calculation method for tide-bound water level in intelligent channel design
CN115329604A
Tidal river reach water level and flow relation analysis method based on tidal level harmonic analysis
CN119885966A
A METHOD FOR FORECASTING STORM RISES IN WATER LEVELS FOR MARINE ESTUARIAL SECTIONS OF RIVER SEASONS
RU2011144628A