A Method for Calculating Watershed Flood Response Time Based on a Univariate Optimized DMCA Model
By combining the univariate optimization DMCA model with Gaussian mixture clustering and the NSGA-II algorithm, the problems of data dependence and noise identification in the calculation of watershed flood response time in existing technologies are solved, and efficient and accurate calculation of watershed flood response time is achieved.
Patent Information
- Application Number
- CN202310104274.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-13
- Publication Date
- 2026-03-06
- Estimated Expiration
- 2043-02-13
AI Technical Summary
Existing technologies require a large amount of data to calculate the response time of a watershed flood, resulting in high costs for data collection and analysis. Furthermore, they struggle to accurately identify inflection points in the hydrological process line under different flood magnitudes, leading to errors in the estimation of the peak time of flood events.
A univariate optimization DMCA model, combined with a Gaussian mixture clustering model and the NSGA-II algorithm, was used to calculate the watershed flood response time by performing cluster analysis and detrended moving average cross-correlation analysis on rainfall-runoff data, thus avoiding hydrological curve separation and parameter estimation.
It enables accurate calculation of watershed flood response time without assumptions or parameter estimation under flood conditions of different magnitudes, exhibits strong robustness and noise resistance, and reduces data acquisition and analysis costs.
Smart Images

Figure CN116167513B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of watershed flood prediction and management technology, and in particular to a method for calculating watershed flood response time based on a univariate optimized DMCA model. Background Technology
[0002] The rapid response time of a watershed to new rainfall inputs of different magnitudes is one of the key time variables in hydrology. It indicates the time required for rainfall input entering the watershed at a specific time at different locations within the watershed to respond to the outflow process. Accurate calculation of the watershed flood response time is a necessary condition for watershed-wide hydrological simulation and hydrographic mapping. Uncertainty in the results may lead to errors in the estimation of peak time during flood events. Rainstorm and flood disasters cause disruptions to urban and rural transportation and public services, thus increasing emergency rescue time. Therefore, accurate calculation of the response time of watershed floods of different magnitudes is helpful for the construction of a watershed flood early warning system and provides scientific support for refined watershed emergency management. Currently, methods for calculating watershed response time can be summarized into three categories: first, using hydrochemical information such as isotopes in rainfall and runoff, and utilizing changes in relevant tracers in input and output fluxes to calculate the average runoff response time; second, directly calculating the watershed response time T from rainfall and runoff time series. r The methods employed include: 1) directly estimating the time from the end of rainfall to the inflection point of the descending branch of the hydrological process line using hydrological curves; and 2) using a recursive single-parameter digital filter to separate the hydrological curves into direct runoff and baseflow, ensuring that the baseflow passes through the inflection point of the total rainfall process line. However, these methods require a large amount of data, resulting in high costs for data acquisition and analysis. Furthermore, calculating the response time for floods of different magnitudes requires estimating multiple parameters to separate complex hydrological curves. Additionally, when the temporal resolution of the data is high, noise in the signal makes it difficult to automatically identify inflection points in the hydrological process line. Summary of the Invention
[0003] To address the shortcomings of existing technologies, this invention aims to propose a method for calculating watershed flood response time based on a univariate optimized DMCA model. It employs a GMM clustering model to perform cluster analysis on floods of different magnitudes, requiring no assumptions about the rainfall-runoff conversion process and eliminating the need for parameter estimation methods to separate hydrological curves, thus enabling the calculation of watershed flood response time.
[0004] This invention is achieved using the following technical solution:
[0005] A method for calculating the response time of a watershed flood based on a univariate optimized DMCA model, comprising the following steps:
[0006] Step 1: Collect basic geographic data, historical rainfall data, and river flow process data of the watershed to be processed from rain gauges;
[0007] Step 2: Data preprocessing. The basic geographic data, historical rainfall data, and river flow process data of the watershed collected in Step 1 are normalized using the normalization method.
[0008] Step 3: Perform cluster analysis calculations for various flood data using the Gaussian Mixture Model (GMM), and establish the expectation-maximization clustering model of the GMM. This step specifically includes the following processing:
[0009] Initialize the Gaussian distribution parameter α for each cluster component. m μ m , ∑ m Establish a Gaussian mixture clustering model (GMM) for each cluster component;
[0010] The maximum likelihood function of the parameters is constructed using the Expectation-Maximization (EM) algorithm. Each iteration consists of two steps:
[0011] Step 3-1: Solve for the optimal sample parameters by maximizing the log-likelihood function of the mixture model, as shown in the following equation:
[0012]
[0013] In the formula: x = {x1, x2, ..., x} k} represents the flow process dataset from the Gaussian Mixture Clustering (GMM) model, p(x i Let θ) represent the model probability function, where θ represents a single unknown parameter or a parameter vector consisting of several unknown parameters, i represents the Gaussian distribution index, and z i The dataset represents a rainfall event, z = {z1, z2, ..., z...}. i , ..., z N The i-th value in};
[0014] Step 3-2: Using the expected value maximization algorithm, the log-likelihood function of the maximized model distribution of the maximum likelihood estimation is solved iteratively to obtain the maximum parameter values of the parameter vector θ, as shown in the following formula:
[0015]
[0016] In the formula, x i The data set represents the flow process, x = {x1, x2, ..., x}. i , ..., x N The i-th value in}, z i The dataset represents a rainfall event, z = {z1, z2, ..., z...}. i , ..., z N The i-th value in};
[0017] The prior probabilities of a Gaussian Mixture Clustering (GMM) model are defined as follows:
[0018]
[0019] In the formula: z = {z1, z2, ..., z} N} represents the rainfall process data, p() represents the probability function, and m represents the m-th sub-model of the Gaussian mixture clustering (GMM) model;
[0020] According to Bayes' theorem, maximum a posteriori probability estimation The definition is as follows:
[0021]
[0022] In the formula: α m μ represents the weight of the m-th Gaussian distribution in the Gaussian mixture model. m Let represent the mean of the m-th Gaussian distribution component, j represent the j-th sample, and m represent the m-th sub-model of the Gaussian Mixture Cluster Analysis (GMM) model.
[0023] To find the optimal parameter θ, the maximum value of the log-likelihood function parameter θ is defined as:
[0024] θ j+1 =argmax θ L(θ,θ j )
[0025] Calculate the Gaussian distribution model parameters for each cluster component in the new iteration.
[0026] The above process is iterated continuously until θ is reached. j With θ j+1 The value is infinitely close;
[0027] Step 4: Based on the Gaussian clustering analysis results, construct a univariate optimization-DMCA model for each flood group during different rainfall events. The specific process is as follows:
[0028] Step 4-1: Construct the rainfall R for each flood event. t and traffic Q t The cumulative time series;
[0029] Step 4-2: Using the Detrended Moving Average Cross-Correlation Analysis (DMCA) method, the trend term is removed by moving averages to estimate the cross-correlation between nonlinear time series.
[0030] The correlation coefficient of the DMCA model is calculated as follows:
[0031]
[0032] In the formula: ρ represents the mean square value of the binary fluctuation of the flow and rainfall series, λ represents the moving average window length of the model parameter variables; DMCA (λ) represents the correlation coefficient of the model; F R (λ) represents the root mean square value of the rainfall sequence; F Q (λ) represents the root mean square value of the flow process sequence.
[0033] Step 5: Use the NSGA-II algorithm to optimize the parameters. The optimization result is to minimize the correlation coefficient of the DMCA model. Then, use a univariate optimization model to solve for the minimum moving window length in the DMCA model.
[0034] Step 6: Calculate the response time for different types of floods in the watershed.
[0035] Compared with the prior art, the present invention can achieve the following beneficial technical effects:
[0036] 1) No assumptions are required regarding the rainfall-runoff transition: including the absence of hydrological curve separation and the selection of specific rainfall events;
[0037] 2) It has strong robustness to shorter time series data and remains effective even when there is noise and bias in the time series. Attached Figure Description
[0038] Figure 1 This is a flowchart of a watershed flood response time calculation method based on a univariate optimized DMCA model according to the present invention;
[0039] Figure 2 This is a schematic diagram of the relationship between rainfall and runoff. Detailed Implementation
[0040] The technical solution will now be clearly described in conjunction with the accompanying drawings and embodiments.
[0041] like Figure 1 The diagram shown is an overall flowchart of a watershed flood response time calculation method based on a univariate optimized DMCA model according to the present invention, which specifically includes the following steps:
[0042] Step 1: Collect basic geographic data, historical rainfall data, and river flow process data of the watershed to be processed from rain gauges. The basic geographic data of the watershed includes the watershed range, river system shape map layer, and high-precision DEM data of the watershed. The historical rainfall data includes hourly point rainfall data and latitude and longitude coordinates of rain gauges for typical floods (20-30 floods). The river flow process data includes hourly flow process data at the watershed outlet for typical floods (20-30 floods).
[0043] Step 2: Data preprocessing. The basic geographic data, historical rainfall data, and river flow process data of the watershed collected in Step 1 are normalized using the normalization method.
[0044] Step 3: Perform cluster analysis calculations for various flood data using the Gaussian Mixture Model (GMM), and establish the Expectation-Maximization (EM) clustering model of the Gaussian Mixture Model (GMM). This step specifically includes: assuming that the clustering factors such as peak discharge, rainfall duration, and maximum rainfall intensity of each flood event conform to a Gaussian mixture distribution, and initializing the Gaussian distribution parameter α for each cluster component. m μ m ,∑ m Establish a Gaussian mixture clustering (GMM) sub-model for each cluster component;
[0045] For a random variable X, a Gaussian model can be determined by its probability density function, which is mainly composed of the mean μ and the variance σ. 2 The parameter set describing the Gaussian mixture model is denoted as θ, i.e., α. m μ m ,∑ m (α m μ represents the weight of the m-th Gaussian distribution in the Gaussian mixture model. m Let ∑ represent the mean of the m-th Gaussian distribution component. m (This represents the covariance matrix of the m-th Gaussian distribution component). To obtain a high-quality clustering result, the clusters need to fit the distribution of the sample dataset infinitely. The optimal sample parameters are solved by maximizing the log-likelihood function of the mixture model, as shown in the following equation:
[0046]
[0047] Where: x1, x2, ..., x m Let p(x) represent a mixture model of m Gaussian-distributed flow process sample data. i Let θ) represent the model probability function, where θ represents a single unknown parameter or a parameter vector consisting of several unknown parameters, i represents the Gaussian distribution index, and z i The dataset represents a rainfall event, z = {z1, z2, ..., z...}. i , ..., z N The i-th value in};
[0048] Taking a flood peak discharge sample as an example, assume that all data (rainfall to discharge data) follow a mixture Gaussian distribution, and each Gaussian distribution forms a cluster. Let the random variable be X, X = {x1, x2, ..., x...} NFurthermore, the mixture model consists of M Gaussian distributions, and the Gaussian Mixture Clustering (GMM) model is shown in the following equation:
[0049]
[0050]
[0051] In the formula: θ represents a single unknown parameter or a parameter vector consisting of several unknown parameters; α m μ represents the weight of the m-th Gaussian distribution in the Gaussian mixture model, and is generally unknown; m Let ∑ represent the mean of the m-th Gaussian distribution component. m θ represents the covariance matrix of the m-th Gaussian distribution component; m Let m represent the parameter set of the m-th Gaussian distribution.
[0052] The maximum likelihood function of the parameters is constructed using the Expectation Maximum (EM) algorithm, with each iteration consisting of two steps:
[0053] Step 3-1: For the flow process sequence X = {x1, x2, ..., x...} N The Gaussian model can be determined through the probability density function, which is mainly composed of the mean μ and the variance σ. 2 The parameter set describing the Gaussian mixture model is denoted as θ, i.e., α. m μ m ,∑ m To obtain a high-quality clustering result, the clusters need to fit the distribution of the sample dataset infinitely. The optimal sample parameters are solved by maximizing the log-likelihood function of the mixture model, as shown in the following equation:
[0054]
[0055] Where: x1, x2, ..., x m Let p(x) represent a mixture model of m Gaussian-distributed flow process sample data. i Let θ) represent the model probability function, where θ represents a single unknown parameter or a parameter vector consisting of several unknown parameters, i represents the Gaussian distribution index, and z i The dataset represents a rainfall event, z = {z1, z2, ..., z...}. i , ..., z N The i-th value in};
[0056] Step 3-2: Use the EM algorithm, i.e., the expectation-maximization algorithm, to solve for the parameter θ of the maximum likelihood estimate through iteration. For the flow process sequence X = {x1, x2, ..., x...} N The rainfall process sequence Z = {Z1, Z2, ..., Z}N By maximizing the log-likelihood function of the model distribution, we can obtain the maximum parameter values:
[0057]
[0058] In the formula, x i The data set X represents the flow process dataset X = {x1, x2, ..., x...} N The i-th value in}, z i The dataset represents the rainfall process: Z = {Z1, Z2, ..., Zn} N The i-th value in};
[0059] The probability generated by the model itself, i.e., the prior probability, is defined as follows:
[0060]
[0061] In the formula: z = {z1, z2, ..., z} N} represents the rainfall event data, p() represents the probability function, and m represents the total number of Gaussian distributions;
[0062] According to Bayes' theorem, by selecting initial values for the parameters of the Gaussian mixture model, the m-th sub-model in the Gaussian mixture model is used to analyze the sample data X = {x1, x2, ..., x...}. N The influence of}, i.e., the maximum a posteriori probability estimate, can be defined as:
[0063]
[0064] In the formula: α m μ represents the weight of the m-th Gaussian distribution in the Gaussian mixture model. m Let represent the mean of the m-th Gaussian distribution component. This represents the m-th sub-model in the Gaussian Mixture Cluster Analysis (GMM) model for the sample data x. j The influence degree, i.e., the maximum a posteriori probability; j represents the j-th sample; m represents the m-th sub-model of the Gaussian mixture clustering (GMM) model;
[0065] As shown in equation (2), the value of the likelihood function changes with the change of θ. The purpose of maximum likelihood estimation is to fix the population sample data and solve for the optimal parameter θ.
[0066] The maximum value of the parameter θ of the log-likelihood function is defined as:
[0067] θ j+1 =argmax θ L(θ,θ j (7)
[0068] Calculate the Gaussian distribution model parameters for each cluster component in the new iteration. The formula is as follows:
[0069]
[0070]
[0071]
[0072] Where, α m μ represents the weight of the m-th Gaussian distribution in the Gaussian mixture model. m Let ∑ represent the mean of the m-th Gaussian distribution component. m Let x represent the covariance matrix of the m-th Gaussian distribution component. j The data set represents the flow process, x = {x1, x2, ..., x}. j , ..., x N The j-th value in} This represents the m-th sub-model in the Gaussian Mixture Cluster Analysis (GMM) model for the sample data x. j The influence degree, i.e., the maximum a posteriori probability; j represents the j-th sample; m represents the m-th sub-model of the Gaussian mixture clustering (GMM) model;
[0073] The above process is iterated continuously until θ is reached. j With θ i+1 The value is infinitely close.
[0074] Step 4: Based on the Gaussian clustering analysis results, i.e., different levels of rainfall processes, construct a univariate optimization-DMCA model for each group of floods, as follows:
[0075] Step 4-1: Construct the rainfall R for each flood event. t and traffic Q t The cumulative time series must have the same time scale and time step; the cumulative rainfall R within time t for each flood event. t and cumulative traffic Q t The formula is as follows;
[0076]
[0077]
[0078] In the formula: r l Let q be the rainfall record for the l-th time step within time t. l This is the flow record for the l-th time step within time t;
[0079] Step 4-2: Using the Detrended Moving Average Cross-Correlation Analysis (DMCA) method, the trend term is removed by moving averages to estimate the cross-correlation between nonlinear time series.
[0080] Calculate the fluctuation of the central moving average of each cumulative time series relative to the moving average window length λ of the DMCA model parameters, and then calculate the mean square value of the binary fluctuations of the cumulative rainfall and cumulative flow time series. As shown in the following formula:
[0081]
[0082]
[0083] In the formula: These are the cumulative rainfall and the central moving average of runoff, respectively, within the moving average window length λ of the model parameter variables at time t.
[0084]
[0085]
[0086] The mean square values of the two time series bivariate fluctuations, cumulative rainfall and cumulative flow, are calculated as follows:
[0087]
[0088] The correlation coefficient of the DMCA model is calculated as follows:
[0089]
[0090] Step 5: Use the NSGA-II algorithm to optimize the parameters. The optimization result is to minimize the correlation coefficient of the DMCA model. Then, use a univariate optimization model to solve for the minimum moving window length in the DMCA model.
[0091] Step 6: Calculate the response time for different types of floods in the watershed.
[0092] like Figure 2 The figure shows a schematic diagram of the relationship between rainfall and runoff. When the positive fluctuations in rainfall and the negative fluctuations in runoff cover the entire time span between the two centroids (as indicated by the arrow on the right), and because a fluctuation of one sign is to be generated within a specific duration, the moving average window length λ is twice the time between the two centroids (as indicated by the arrow on the left, which approaches 0).
[0093] Based on the embodiments of this invention, all other embodiments and technical substitutions of the embodiments obtained by those skilled in the art without departing from the spirit of this invention and without creative effort will fall within the scope of protection of this invention.
Claims
1. A method for calculating a flood response time of a watershed based on a univariate optimization DMCA model, characterized in that, The method comprises the following steps: Step 1, collecting basic geographic data, historical rainfall data and river flow process data of a to-be-processed basin from a rainfall station; Step 2, data preprocessing, the basic geographic data, historical rainfall data and river flow process data of the to-be-processed basin collected in the step 1 are normalized respectively by using a normalization method; Step 3, a Gaussian mixture clustering model GMM is used to complete clustering analysis calculation of various flood data, and an expectation maximization clustering model of the Gaussian mixture clustering model GMM is established; the step specifically comprises the following processing: Initializing Gaussian distribution parameters a, μ, ∑ of each cluster component m m m Establishing Gaussian mixture clustering analysis (GMM) model of each cluster component The expectation maximization algorithm EM is used to construct a maximum likelihood function of parameters, and each iteration comprises two steps: Step 3-1: the optimal sample parameter is solved by maximizing the logarithmic likelihood function of the mixed model, as shown in the following formula: where x = {x1, x2,..., x k} represents the flow process data set from the Gaussian Mixture Clustering Analysis (GMM) model, p(x i , θ) represents the model probability function, θ represents an unknown parameter or a parameter vector composed of several unknown parameters, i represents the number of Gaussian distributions, z i represents the i-th value in the rainfall process data set z = {z1, z2,..., z i ,..., z N}. Step 3-2: the maximum likelihood estimation of the maximization model distribution is solved by iteration by using the expectation maximization algorithm, and the maximum parameter value of the parameter vector θ is obtained, as shown in the following formula: where x i represents the i-th value in the flow process data set x = {x1, x2,..., x i ,..., x N} and z i represents the i-th value in the rainfall process data set z = {z1, z2,..., z i ,..., z N}. The prior probability of the Gaussian mixture clustering analysis GMM model is defined as follows: In the formula: z = {z1, z2, ..., z} N } represents the rainfall process data, p() represents the probability function, and m represents the m-th sub-model of the Gaussian mixture clustering (GMM) model; According to Bayes' theorem, the maximum a posteriori probability estimate is defined as follows: wherein: a m represents the weight of the mth Gaussian distribution in the Gaussian mixture model, μ m represents the mean of the mth Gaussian distribution component, j represents the jth sample; m represents the mth sub-model of the Gaussian mixture clustering analysis (GMM) model; The maximum value of the logarithmic likelihood function parameter θ is defined as the optimal parameter θ. θ j+1 = argmax θ L(θ, θ j ) computing the Gaussian distribution model parameters for each cluster component of a new round of iteration The above process is iterated until θ j is infinitely close to the value of θ j+1 . Step 4, according to the Gaussian clustering analysis result, a univariate optimization-DMCA model is constructed for each group of floods in different grades of rainfall process, and the specific process is as follows: Step 4-1, constructing the rainfall R of each event t and cumulative time series of flow Q t ; Step 4-2, the detrended moving average cross-correlation analysis method DMCA is used to remove the trend term by using a moving average line, and the cross-correlation between nonlinear time series is estimated; The correlation coefficient of the DMCA model is calculated, as shown in the following formula: where: denotes the mean square value of the binary fluctuations of the flow and rainfall series, λ denotes the model parameter variable moving average window length, and ρ DMCA denotes the correlation coefficient of the model, F R denotes the root mean square value of the rainfall process series, F Q denotes the root mean square value of the flow process series; Step 5, the NSGA-II algorithm is used for parameter optimization, and the optimization result is that the correlation coefficient of the DMCA model is minimum, and the minimum moving window length in the DMCA model is solved by using the univariate optimization model; Step 6, the response time of different types of floods in the basin is calculated.
2. The method of claim 1, wherein the DMCA model is based on a single variable optimization. The basin basic geographic data includes the basin range, water system shp layer and high-precision DEM data of the basin, the historical rainfall data includes hourly point rainfall data and longitude and latitude coordinates of a typical flood rainfall station, and the river flow process data includes hourly flow process data of the basin outlet in a typical flood.
3. The method of claim 1, wherein the DMCA model is based on a single variable optimization. Two time series have the same time scale T and time step i.
4. The method of claim 1, wherein the DMCA model is based on a single variable optimization. The Gaussian mixture clustering analysis GMM model of each clustering component established in the step 3 is specifically as follows: It is assumed that all rainfall and flow data are subject to a mixed Gaussian distribution, and each Gaussian distribution is a cluster, a random variable is X, and the mixed model is composed of M Gaussian distributions, and the Gaussian mixture model is as shown in the following formula: where θ represents an unknown parameter or a parameter vector consisting of several unknown parameters, α m represents the weight of the mth Gaussian distribution in the mixture Gaussian model, μ m represents the mean of the mth Gaussian distribution component, ∑ m represents the covariance matrix of the mth Gaussian distribution component, θ m represents the parameter set of the mth Gaussian distribution.
5. The method for calculating the watershed flood response time based on the univariate optimized DMCA model as described in claim 1, wherein the Gaussian distribution model parameters... The formula for the estimated value is as follows: wherein α m denotes the weight of the mth Gaussian distribution in the Gaussian Mixture Model, μ m denotes the mean of the mth Gaussian distribution component, ∑ m denotes the covariance matrix of the mth Gaussian distribution component, x j denotes the jth value in the flow process data set x = {x1, x2,..., x j ,..., x N} denotes the influence degree of the mth sub-model in the Gaussian Mixture Cluster Analysis GMM model on the sample data x j , i.e. the maximum a posteriori probability estimation, j represents the jth sample, and m represents the mth sub-model of the Gaussian Mixture Cluster Analysis GMM model.
6. The method for calculating the flood response time of a watershed based on the univariate optimization DMCA model according to claim 1, wherein, The cumulative rainfall R in each flood t time t and the cumulative flow Q t The formula is as follows; where: r l is the rainfall record for the lth equal time step within time t, q l is the flow record for the lth equal time step within time t.
7. The method of claim 1, wherein the DMCA model is based on a single variable optimization. In step 5, the DMCA model parameter moving average window length λ is taken as the decision variable, the upper limit of the decision variable is the dimension of the cumulative rainfall time series, the lower limit of the decision variable is 0, and the DMCA model correlation coefficient ρ DMCA (λ) is a univariate optimization objective, and the objective function is the correlation coefficient ρ DMCA (λ) is the minimum, a single-objective optimization model for maximizing the continuity of the decision variable is constructed, and the minimum moving average window length λ is solved.
8. The method for calculating the flood response time of a watershed based on the univariate optimization DMCA model according to claim 4, wherein, In step 4-2, the fluctuation of the moving average value of each cumulative time series relative to the center of the moving average window length λ of the DMCA model parameters is calculated, and then the mean square value of the binary fluctuation of the cumulative rainfall and the cumulative flow time series is calculated As shown in the following formula: where: respectively, are the cumulative rainfall and the center moving average of the runoff with the moving average window length λ and the time model parameter variable t. The mean square value of the binary fluctuation of the two time series, cumulative rainfall and cumulative flow, is calculated as shown in the following equation:
9. The method for calculating the flood response time of a watershed based on the univariate optimization DMCA model according to claim 8, wherein, When the rainfall positive fluctuation and the flow negative fluctuation cover the entire time span between the two centroids, the moving average window length λ is twice the time between the two centroids.
Citation Information
Patent Citations
Insulating tube bus fault diagnosis method based on LDA optimization multi-scale textural features
CN111626329A
Flood forecasting model and flood forecasting method based on deep learning framework
CN114154417A