Calculation method and device for time downscaling of river material flux, and electronic equipment
By using the random forest model to optimize the time scale of river material flux, the problem of insufficient time scale of river material flux in existing technologies is solved, and higher-precision river material flux prediction is achieved, supporting river pollution management and prevention.
Patent Information
- Application Number
- CN202410752080.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-12
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2044-06-12
AI Technical Summary
Existing technologies make it difficult to optimize the time scale of river material flux efficiently and accurately, resulting in a low time scale of river material flux data, which affects the accuracy of water environment management models.
Machine learning models, especially random forest models, are used to collect water quality, hydrological and meteorological data of rivers, establish nonlinear relationships between response variables and predictor variables, optimize the feature importance of predictor variables, and retrain the model to reduce the time scale of river material flux.
It improves the accuracy and robustness of the time scale of river material flux, reduces prediction errors, makes the prediction results closer to the measured values, and provides more accurate data to support river pollution management and prevention.
Smart Images

Figure CN118747279B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of numerical simulation and optimization of water environment, and specifically to a calculation method and device, and electronic equipment for time-scaling the material flux of a river, which is suitable for reducing the time scale of the material flux of a river. Background Art
[0002] With the rapid development of the natural environment and socioeconomic development in recent years, nutrient pollution in river basins has become increasingly severe and complex. Obtaining high-resolution material flux data has gradually become a key technical requirement and an important decision-making basis for controlling the total amount of pollutants in river water environments. However, due to high costs, the frequency of monitoring concentrations of substances such as nitrogen and phosphorus is relatively low, making it difficult to obtain water quality data synchronized with hydrological monitoring. Current surface water monitoring technical specifications in my country stipulate that water quality monitoring is conducted 4-12 times per year, while hydrological monitoring is conducted daily. This frequency difference is one of the main factors contributing to the low temporal scale of material flux data. Given the scarcity of water quality observational data globally, data on the drivers of material fluxes are increasingly being used to predict water quality fluxes. Previous studies have shown a significant correlation between river material fluxes and basin hydrometeorology. Improving the accuracy of river material fluxes based on data on these drivers has become an important method for predicting river material load intensity. Currently, statistical principles and time series analysis methods are widely used to fit material fluxes within a certain time range through data interpolation. However, this type of downscaling method often overestimates river material fluxes and has insufficient response capabilities to data variations caused by multi-factor interference, affecting model accuracy. Therefore, efficiently and accurately optimizing and reducing the time scale of river material fluxes is a key technical issue that needs to be urgently addressed in the field of water environment management. Summary of the Invention
[0003] The purpose of this application is to provide a calculation method, device, and electronic equipment for time downscaling of river material flux, which uses a machine learning model to optimize predictive variables, establishes a nonlinear relationship between predictive variables and flux, and solves the problem of insufficient time scale accuracy of river material flux by fitting high time resolution flux.
[0004] According to a first aspect of an embodiment of the present application, a method for calculating time downscaling of river material flux is provided, comprising:
[0005] Collect river water quality data, hydrological data and meteorological data;
[0006] Calculating the daily flux of river substances based on the water quality data and the hydrological data, and setting the daily flux of river substances as a response variable of the model;
[0007] Calculating monthly scale flux based on the daily scale flux, wherein the monthly scale flux, hydrological data and meteorological data are set as prediction variables of the model;
[0008] Based on the response variable and the predictor variable, establishing a random forest model of the response variable and the predictor variable;
[0009] Selecting predictor variables based on the predictor variable feature importance results in the random forest model;
[0010] The random forest model was retrained based on the preferred predictor variables, and the new predictor variables were input into the retrained random forest model to downscale the river flux to the daily scale.
[0011] Optionally, the river water quality data includes: substance concentration data at a river monitoring section;
[0012] The hydrological data include: daily runoff of a river monitoring section based on the time corresponding to the water quality data, and annual accumulated days based on the time corresponding to the water quality data collection time;
[0013] The meteorological data includes: daily rainfall, rainfall in the previous three days, rainfall in the previous seven days and daily surface temperature values based on the corresponding time of the water quality data.
[0014] Optionally, the daily scale flux of river substances is calculated based on the water quality data and hydrological data. The daily scale flux of river substances is the measured value of the daily flux of the river. The calculation formula is as follows:
[0015] L=C×Q×0.0864
[0016] Where L is the daily flux of river substances, in kg / d; C is the substance concentration in the monitoring section, in mg / L; Q is the runoff in the monitoring section, in m 3 / s.
[0017] Optionally, a monthly scale flux is calculated based on the daily scale flux, and the monthly scale flux, hydrological data, and meteorological data are set as prediction variables of the model, including:
[0018] The monthly flux is obtained by accumulating daily flux data;
[0019] The hydrological data after extraction and processing is the daily runoff and annual accumulation of the monitoring section;
[0020] The meteorological data after extraction and processing is daily rainfall data, the accumulated rainfall in the previous three days, the accumulated rainfall in the previous seven days and the daily surface temperature value;
[0021] Optionally, based on the response variable and the predictor variable, a random forest model of the response variable and the predictor variable is established, with the daily flux as the response variable, and the monthly flux, daily runoff, annual accumulation day, daily rainfall data, rainfall accumulation in the previous three days, rainfall accumulation in the previous seven days and daily surface temperature values as predictor variables, and the model is constructed using the random forest's function of fitting the response variable and calculating the importance of the predictor variable features.
[0022] Optionally, the predictor variables are selected based on the predictor variable feature importance results of the random forest model, including:
[0023] The feature importance of the predictor variable is calculated by the random forest model. In the random forest model, each splitting node of each decision tree will evaluate the importance of the feature according to the Gini coefficient of the node. The importance of the same feature in all decision trees is accumulated and averaged. The average value is called Gini importance and is also the feature importance of the predictor variable.
[0024] Selecting predictive variables based on the model performance data and the importance scores of the predictive variables in the random forest to improve the predictive ability of the random forest model;
[0025] Optionally, retrain the random forest model based on the preferred predictor variables, input the new monthly fluxes and related predictor variables into the retrained random forest model, and downscale the river flux to the daily scale, including:
[0026] The preferred predictor variables are re-entered into the random forest model to improve the model's analysis and fitting capabilities.
[0027] The new prediction variables include new monthly fluxes to be downscaled, and hydrological data and meteorological data corresponding to the monthly fluxes.
[0028] According to a second aspect of an embodiment of the present application, a device for time-scaling river material flux is provided, comprising:
[0029] Collection module, used to collect water quality data, hydrological data and meteorological data of river monitoring sections;
[0030] A calculation module, configured to calculate daily river flux data based on the hydrological data and water quality data, and calculate annual daily data based on the time span of the data;
[0031] A training module is used to establish a random forest model for the response variable and the predictor variable, including calculating the importance of the random forest model using a training set with daily flux as the response variable and monthly flux, daily runoff, annual accumulation day, daily rainfall data, rainfall accumulation in the previous three days, rainfall accumulation in the previous seven days, and daily surface temperature as the predictor variables;
[0032] Modeling module, which selects random forest predictor variables based on importance ranking, retrains the random forest model, and establishes a river material flux downscaling model;
[0033] The prediction module is used to input the new monthly scale material flux and the corresponding preferred prediction variables into the model to estimate the predicted value of the daily scale river material flux.
[0034] According to a third aspect of the embodiments of the present application, there is provided an electronic device, including:
[0035] one or more processors;
[0036] a memory for storing one or more programs;
[0037] When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in the first aspect.
[0038] According to a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which computer instructions are stored. When the instructions are executed by a processor, the steps of the method described in the first aspect are implemented.
[0039] The technical solutions provided by the embodiments of the present application may have the following beneficial effects:
[0040] This application presents one of the primary methods for downscaling monthly river material fluxes to daily fluxes. Machine learning methods have been widely applied in numerous watershed environments, particularly in addressing the limited temporal resolution of fluxes in large rivers, both domestically and internationally. The proposed method has low data type requirements, a wide range of data acquisition methods, and simple data preprocessing. The proposed method has broad applicability across a variety of river water environment scenarios and pollutant types.
[0041] This application uses random forest importance analysis to screen feature data, which can effectively identify the environmental effect factors that drive changes in different substances and analyze the contribution capabilities of multiple factors, thereby achieving accurate analysis of material flux, optimizing variables to improve the accuracy and noise resistance of the overall model, and more comprehensively understanding the impact of environmental characteristics on material flux.
[0042] This application proposes a method for calculating the time scale of river material flux, overcoming the problem of poor fit between predictions and measured data in current research. This method uses machine learning data processing technology and statistical modeling to screen variable data based on the complexity and spatiotemporal variability of river environments to improve the accuracy and robustness of random forest predictions. By deeply analyzing data characteristics and optimizing model parameters, prediction errors are effectively reduced, making the prediction results closer to the measured values. This method provides important scientific support for the field of river pollution control and prevention, and provides an accurate data foundation and a complete methodological framework for relevant decision-making.
[0043] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0045] Figure 1 The present invention is a flowchart showing a method for calculating time downscaling of river material flux according to an exemplary embodiment.
[0046] Figure 2 The figure shows a random forest feature importance ranking diagram according to an exemplary embodiment.
[0047] Figure 3 2 is a graph showing the performance of river flux downscaling prediction based on random forest according to an exemplary embodiment.
[0048] Figure 4 The present invention is a block diagram of a device for downscaling monthly-scale material flux of a river to daily flux according to an exemplary embodiment. DETAILED DESCRIPTION
[0049] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0050] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. As used in this application and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0051] Figure 1 FIG. 1 is a flow chart showing a method for calculating time downscaling of river material flux according to an exemplary embodiment. Figure 1 As shown, the method may include the following steps:
[0052] S1: Collect river water quality data, hydrological data and meteorological data;
[0053] Specifically, the water quality data selected for the example of the present invention is the daily total nitrogen concentration of the Dongjiang Boluo monitoring section from 2019 to 2021, with a total of 1096 monitoring samples. The hydrological data selected is the daily runoff of the Dongjiang Boluo monitoring section, and the meteorological data selected is the 24-hour daily surface air temperature value and the 24-hour daily rainfall of the monitoring section meteorological station from 2019 to 2021. After collecting the hydrological, water quality and meteorological data of the monitoring section, the collected data are cleaned, outliers are removed and missing values are interpolated to obtain complete paired data. The average daily temperature value, daily rainfall, multi-day rainfall (three days and seven days) and annual accumulated days of the Boluo monitoring section are calculated. In order to test the overall prediction ability of the method, the Dongjiang data set is divided into a training set and a test set. After the training set data is used to establish the model, the test set data is called to verify the fitting accuracy.
[0054] S2: Calculating the daily flux of river substances based on the water quality data and the hydrological data, and setting the daily flux of river substances as the response variable of the model;
[0055] Specifically, according to the total nitrogen concentration and daily runoff data of the Boluo monitoring section, the daily total nitrogen flux of the monitoring section is calculated;
[0056] Specifically, the daily total nitrogen flux calculation formula is as follows:
[0057] L=C×Q×0.0864
[0058] Where L is the daily flux of river substances, in kg / d; C is the substance concentration in the monitoring section, in mg / L; Q is the runoff in the monitoring section, in m 3 / s, the total nitrogen daily flux obtained is called the measured value, which is set as the response variable (Y) in the model;
[0059] S3: Calculate the monthly scale flux based on the daily scale flux, and set the monthly scale flux, hydrological data and meteorological data as the prediction variables of the model.
[0060] Specifically, the daily-scale flux data were accumulated to obtain the monthly-scale flux data, and the prediction variables and response variables were constructed as the input files of the random forest model for the flux time downscaling of the Boluo monitoring section of the Dongjiang River.
[0061] S4: Based on the response variable and the predictor variable, establishing a random forest model of the response variable and the predictor variable;
[0062] Specifically, the method includes the following sub-steps:
[0063] S41: Using the monthly flux (X1), daily runoff (X2), annual accumulated days (X3), daily rainfall data (X4), cumulative rainfall over the previous three days (X6), cumulative rainfall over the previous seven days (X7), and daily surface temperature (X8) of the Dongjiang Boluo monitoring section as predictor variables, the input file is stripped of censored data rows, and basic summary information is collected, including the minimum value, first quartile, median, mean, third quartile, and maximum value, as well as the characteristics and distribution of the fast data. The characteristics of the input file data are shown in Table 1:
[0064] Table 1 Basic sample statistical characteristic values
[0065]
[0066]
[0067] S42: The constructed random forest model will randomly extract a subset of the data set to construct a decision tree. For each node split, a subset of features will be randomly selected to find the optimal split point, and multiple decision trees will be repeatedly constructed.
[0068] S5: Predictor variables are selected based on the feature importance results of the predictor variables in the random forest model;
[0069] Specifically, the importance of the feature of the predictor variable at each split node of each decision tree in the random forest will be evaluated according to the Gini coefficient of the node. The importance of the same feature in all decision trees is accumulated and averaged, and the Gini importance is also the feature importance of the predictor variable.
[0070] The smaller the average value of the mean square residual of the random forest model results, the better the model fit; the higher the variable explanation percentage value, the stronger the model's explanatory ability for the target variable. The prediction variables selected for the Dongjiang Boluo monitoring section can explain 85.95% of the variance in the training set data in the random forest model, which is a relatively high degree of explanation, indicating that the model has a good fitting effect. Therefore, the selected prediction flux can well predict the daily scale flux of total nitrogen in the Boluo monitoring section.
[0071] S6: retraining the random forest model based on the preferred predictor variables, inputting the new predictor variables into the retrained random forest model, and downscaling the river flux to the daily scale; specifically, retraining the random forest model, calling the trained random forest to fit the test set daily scale total nitrogen flux data, this step may include the following sub-steps:
[0072] S61: The optimized prediction variables are re-output into the random forest, and the daily flux of the Dongjiang Boluo monitoring section of the training set is fitted to become the predicted value. A linear equation of the predicted value and the measured value is constructed to judge the simulation performance. The performance results of the model for the predicted value and the measured value of the test set are as follows: Figure 3 In (a), the linear correlation between the predicted value and the measured value of the training set is 98.08%, indicating a good fitting effect. The trained random forest model has a strong downscaling prediction effect on the total nitrogen flux of the training set of the Boluo monitoring section.
[0073] S62: The data in the test set are used as new predictor variables to input into the downscaling model of the Dongjiang Boluo monitoring section to obtain the daily flux prediction value, which is then compared with the measured value in the test set to judge the downscaling capability of the model. Figure 3 In (b), the p-values of the coefficient estimates of the time downscaling model for the total nitrogen flux at the Boluo monitoring section of the Dongjiang River are very small (<2 -16 ), indicating that it is statistically significant; the model determination coefficient (R-squared) is 0.922, indicating that the model can explain 92.2% of the variation; the model F statistic is used to test the significance of the model as a whole. In the trained random forest model, the F statistic value is 3711, and the corresponding p value is very small (<2.2 -16 ), indicating that the model is significant overall. As can be seen from the above examples, the river flux downscaling method proposed in this application has reasonable calculation results, high accuracy of the output results, can simulate the daily accumulation of flux, and can downscale monthly river flux to daily flux. This method has greater application potential and practical value.
[0074] Corresponding to the aforementioned embodiment of the river material flux downscaling method, the present application also provides an embodiment of a river material flux downscaling device.
[0075] Figure 4FIG. 1 is a block diagram of a river material flux downscaling device according to an exemplary embodiment. Figure 4 , the device comprises:
[0076] Collection module 1, used to collect water quality data, hydrological data and meteorological data of river monitoring sections;
[0077] Calculation module 2, used to calculate the daily flux data of the river based on the hydrological data and water quality data, and calculate the annual daily data based on the time span of the data;
[0078] Training module 3 is used to establish a random forest model for the response variable and the predictor variable, including calculating the importance of the random forest model using a training set with daily flux as the response variable and monthly flux, daily runoff, annual accumulation day, daily rainfall data, rainfall accumulation in the previous three days, rainfall accumulation in the previous seven days, and daily surface temperature as the predictor variables;
[0079] Modeling module 4, for optimizing random forest predictor variables according to the importance ranking, retraining the random forest model, and establishing a river material flux downscaling model;
[0080] The prediction module 5 is used to input the new monthly scale material flux and the corresponding preferred prediction variables into the model to estimate the predicted value of the daily scale river material flux.
[0081] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0082] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present application scheme. A person of ordinary skill in the art can understand and implement it without paying any creative work.
[0083] Accordingly, the present application also provides an electronic device, comprising: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the river material flux downscaling method as described above.
[0084] Accordingly, the present application also provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the above-mentioned river material flux downscaling method.
[0085] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the contents disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered merely as exemplary, and the true scope and spirit of the present application are indicated by the claims.
[0086] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. A method for calculating the temporal downscaling of river material flux, characterized in that: include: Collect river water quality data, hydrological data and meteorological data; Calculating the daily flux of river substances based on the water quality data and the hydrological data, and setting the daily flux of river substances as a response variable of the model; Calculating monthly scale flux based on the daily scale flux, wherein the monthly scale flux, hydrological data and meteorological data are set as prediction variables of the model; Based on the response variable and the predictor variable, establishing a random forest model of the response variable and the predictor variable; Selecting predictor variables based on the predictor variable feature importance results in the random forest model; The random forest model was retrained based on the preferred predictor variables, and the new predictor variables were input into the retrained random forest model to downscale the river flux to the daily scale; The daily flux of river substances is calculated by the following formula: ; Where, L is the daily flux of river material, in kg / d, C is the substance concentration at the monitoring section, in mg / L; Q The runoff volume of the monitoring section is in m 3 / s; Calculating a monthly scale flux based on the daily scale flux, wherein the monthly scale flux is obtained by accumulating the daily scale flux; A random forest model of response variables and predictor variables was established, with daily flux as the response variable and monthly flux, daily runoff, annual accumulation, daily rainfall data, rainfall accumulation in the previous three days, rainfall accumulation in the previous seven days, and daily surface temperature as the predictor variables. Predictor variables are selected based on the predictor variable feature importance results in the random forest model; including: The importance of the predictor variable features is calculated using a random forest model. In the random forest model, each splitting node of each decision tree will evaluate the importance of the feature based on the Gini coefficient of the node. The importance of the same feature in all decision trees is accumulated and averaged. The average value is called the Gini importance and is also the feature importance of the predictor variable. Based on the performance data of the random forest model and the importance scores of the predictor variables in the random forest, the predictor variables are selected to improve the prediction ability of the random forest model; The random forest model is retrained based on the preferred predictor variables to obtain the downscaling model and the performance of the downscaling model. The new predictor variables include the new monthly flux to be downscaled, and the hydrological data and meteorological data corresponding to the monthly flux. The model result outputs the predicted daily flux.
2. The method according to claim 1, characterized in that The water quality data is the material concentration data of the river monitoring section; the hydrological data includes the daily runoff and annual accumulated runoff of the monitoring section; the meteorological data includes daily rainfall, accumulated rainfall in the previous three days, accumulated rainfall in the previous seven days and daily surface temperature values.
3. A device for temporal downscaling of river material flux, characterized in that: For implementing the calculation method according to any one of claims 1 to 2, the device comprises: Collection module, used to collect water quality data, hydrological data and meteorological data of river monitoring sections; A calculation module, used to calculate the daily flux data of the river based on the hydrological data and water quality data, and calculate the annual cumulative daily data based on the data time; A training module is used to establish a random forest model for the response variable and the predictor variable, where the daily flux is the response variable, and the monthly flux, daily runoff, annual accumulated days, daily rainfall data, rainfall accumulation in the previous three days, rainfall accumulation in the previous seven days, and daily surface temperature are used as the training set for predictor variables to calculate the importance of the random forest model; Modeling module, which selects random forest predictor variables based on importance ranking, retrains the random forest model, and establishes a river material flux downscaling model; The prediction module is used to input the new monthly scale material flux and the corresponding preferred prediction variables into the model to estimate the predicted value of the daily scale river material flux.
4. An electronic device, characterized in that: include: one or more processors; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 2.
5. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the instruction is executed by a processor, the steps of the method according to any one of claims 1 to 2 are implemented.
Citation Information
Patent Citations
Random forest-based multifactor remote sensing surface temperature space downscaling method
CN107748736A
Chemical fertilizer application amount prediction method and device based on machine learning, and electronic equipment
CN117609901A