Basin hydrological forecasting method and system based on big data
By using big data platform and edge computing nodes in basin hydrological forecasting, and combining random forest model and LSTM model for model fusion, the traditional method's shortcomings in real-time data processing and model fusion diversity are solved, and more efficient and accurate basin hydrological forecasting is achieved.
Patent Information
- Application Number
- CN202510051415.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-16
AI Technical Summary
Traditional basin hydrological forecasting methods have shortcomings in real-time data processing and model fusion diversity, and it is difficult to adjust predictions in a timely manner to deal with sudden hydrological changes, and a single model cannot fully utilize the information of multi-source data.
The basin hydrological forecasting method based on big data is adopted, and multi-source data is collected through the big data platform and preprocessed and feature extraction in real time at edge computing nodes. Combined with the random forest model and the long and short-term memory network model LSTM, the model is fusion using the stacking method to form a basin hydrological prediction model.
It improves the accuracy and stability of basin hydrological forecasts, can promptly reflect changes in basin hydrological status, and provides real-time information for decisions such as flood control and water resource scheduling.
Smart Images

Figure CN120011761A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of hydrological forecasting, and more specifically, to a basin hydrological forecasting method and system based on big data. Background Art
[0002] Traditional methods of basin hydrological forecasting mostly rely on simple statistical models or physical concept models. Statistical models often establish prediction models based on the statistical relationship of historical data, but it is difficult to fully consider the complex nonlinear relationships and spatiotemporal variation characteristics in the hydrological process. Although the physical concept model can describe the hydrological process from a physical mechanism, when faced with modern basin hydrological conditions with large data volumes and complex changes, the determination of model parameters and calculation efficiency become difficult. With the changes in the surrounding environment of the basin and the increasing interference of human activities on the hydrological system, traditional methods are difficult to capture these changes in real time and accurately, resulting in limited forecast accuracy. At the same time, the application of a single model cannot comprehensively utilize the information of multi-source data, and it is difficult to fully reflect the laws of hydrological changes in a complex basin environment. New methods are urgently needed to improve the accuracy and real-time performance of basin hydrological forecasts.
[0003] Prior art, such as a Chinese patent application with publication number "CN116643331A", discloses a method for hydrological forecasting based on big data of hydrological information of a regional river basin, the method comprising S1 data collection, including historical hydrological and meteorological data and precipitation ensemble forecast data in the river basin, screening and classifying the original data, and taking rainfall, reservoir inflow, and flow hydrological characteristics as input features of a long-short-term memory cyclic prediction model; S2 data analysis, based on a big data algorithm, generating a refined water system map of the target river basin according to DEM data, and analyzing the hydrological parameters of the target river basin; training and determining the structure of the LSTMC model by simulating the rainfall process; S3 establishing a target interval hydrological forecast model based on a deep learning algorithm; and S4 model experimental results and analysis.
[0004] The problem with the above-mentioned existing technology is that the method is insufficient in real-time data processing and diversity of model fusion; if faced with sudden hydrological changes, the model cannot quickly adjust the prediction due to the inability to obtain and process the latest data in time, which may cause the forecast results to lag and miss the best time to respond to disasters; this method is mainly based on the long-short-term memory cycle prediction model. Although it combines the random forest algorithm for feature extraction and model optimization, it has not formed a complex structure in which multiple models collaborate with each other to jointly improve the prediction effect. It is impossible to fully utilize the ability of different models to analyze data from different angles, which limits the improvement of the overall performance of the model. Summary of the invention
[0005] In order to solve the above technical problems, the present invention proposes a basin hydrological forecasting method and system based on big data.
[0006] The technical solution of the present invention is as follows:
[0007] The present invention proposes a basin hydrological forecasting method based on big data, comprising the following steps:
[0008] Step S1, collecting basin hydrological data through the big data platform, and transmitting it to the edge computing node deployed at the data collection end in real time, and preprocessing and feature extraction of the collected data;
[0009] Step S2, forming a data set with the basin hydrological related data after feature extraction, and dividing the data set into a test set and a validation set through spatiotemporal cross-validation;
[0010] Step S3, using the divided training sets to respectively construct and train a random forest model and a long short-term memory network model LSTM, to obtain a trained random forest model and a long short-term memory network model LSTM;
[0011] Step S4, using the stacking method to fuse the trained random forest model and the long short-term memory network model LSTM to obtain a watershed hydrological prediction model, and using the test set data to evaluate and optimize the watershed hydrological prediction model;
[0012] Step S5, the basin hydrological data collected by the big data platform is transmitted in real time to the edge computing node deployed at the data collection end for preprocessing and feature extraction, and the new feature data is integrated into the original data set, and the basin hydrological data is predicted in real time through the basin hydrological prediction model.
[0013] As a preferred implementation, the basin hydrological related data include: meteorological data, hydrological data, topographic data and soil data; wherein, the meteorological data include: precipitation, temperature, wind speed, wind direction, humidity; the hydrological data include: water level, flow, evaporation; the topographic data include: terrain elevation, slope, slope direction, land use type covering the basin; the soil data include: soil type, texture, porosity.
[0014] As a preferred implementation, the data set is divided into a test set and a validation set by spatiotemporal cross validation, specifically comprising the following steps:
[0015] Time division: According to the time characteristics of the basin hydrological data, the time series is divided into k non-overlapping time periods T1, T2, …, T k ;
[0016] Spatial division: According to the spatial characteristics of the watershed, the watershed space is divided into m non-overlapping sub-areas R1, R2, ..., R m ;
[0017] Verification process: In each verification, a time period T is selected i ,i∈k and a subregion R j ,j∈m is used as the validation set, and the rest of the time and space data is used as the training set. By traversing all time and space combinations, the data can be fully utilized and verified in both time and space dimensions.
[0018] As a preferred implementation, the random forest model outputs the following results:
[0019]
[0020] in:
[0021]
[0022] Where: is the output result of the random forest model; N is the number of decision trees in the random forest model; is the predicted value of sample x by the nth decision tree; w n is the voting weight of the th decision tree; Acc n is the prediction accuracy of the nth decision tree on the training set; Acc o is the prediction accuracy of the oth decision tree on the training set.
[0023] As a preferred implementation, the long short-term memory network model LSTM adopts a multi-modal fusion LSTM model of bidirectional attention LSTM, hierarchical LSTM and skip connection, wherein:
[0024] The bidirectional attention LSTM includes:
[0025] Forward LSTM calculation:
[0026] Forget Gate
[0027]
[0028] Input Gate
[0029]
[0030] Candidate memory cells
[0031]
[0032] Memory Unit
[0033]
[0034] Output Gate
[0035]
[0036] Hidden State
[0037]
[0038] Where: and They are the weight matrices of the forget gate, input gate, candidate memory unit, and output gate in the forward calculation respectively; and are the bias vectors of the forget gate, input gate, candidate memory unit, and output gate in the forward calculation; σ is the sigmoid function; x t is the input data at the time step; ⊙ is the element multiplication; tanh() is the hyperbolic tangent function;
[0039] Backward LSTM calculation:
[0040] Forget Gate
[0041]
[0042] Input Gate
[0043]
[0044] Candidate memory cells
[0045]
[0046] Memory Unit
[0047]
[0048] Output Gate
[0049]
[0050] Hidden State
[0051]
[0052] Where: and They are the weight matrices of the forget gate, input gate, candidate memory unit, and output gate in the backward calculation respectively; and They are the bias vectors of the forget gate, input gate, candidate memory unit, and output gate in the backward calculation respectively;
[0053] The attention weight calculation formula is as follows:
[0054]
[0055] Where: are the weights of the forward and backward attention mechanisms, respectively; are the weight matrices of the forward and backward attention mechanisms respectively;
[0056] The hidden state of the forward LSTM and the hidden state of the backward LSTM are merged. The specific formula is as follows:
[0057]
[0058] Where: h t is the hidden state after fusion;
[0059] Hierarchical LSTM with skip connections includes:
[0060] The bottom LSTM and high-level LSTM in the hierarchical LSTM are based on the same LSTM basic structure as the forward LSTM and backward LSTM. t As the input of the underlying LSTM, the output is the hidden state h of the underlying LSTM t1 ; will h t1 As the input of the high-level LSTM, the output is the hidden state h of the high-level LSTM t2 ; will h t1 and h t2 Perform a jump connection to obtain the final fused hidden state. The specific formula is as follows:
[0061]
[0062] The final fused hidden state is converted into a predicted value through the output layer. The specific formula is as follows:
[0063]
[0064] Where: Output the predicted value for the LSTM model; W y is the dimension-adapted weight matrix; b y is the bias vector of the output layer.
[0065] As a preferred implementation, the stacking method is used to fuse the trained random forest model and the long short-term memory network model LSTM to obtain a watershed hydrological prediction model, which includes:
[0066] The first layer model predicts:
[0067] Use the trained random forest and LSTM models to predict the training set and test set; the prediction results of the random forest model for the training set and test set are and The prediction results of the LSTM model for the training set and the test set are and
[0068] Second layer fusion model training:
[0069] Will and As new features, form a new training data set Will and As new features, form a new test data set
[0070] Second layer model training and prediction:
[0071] use The second layer model is trained with actual observations from the basin hydrological data; and the new test data set is used The final fusion prediction result is obtained. The specific formula is as follows:
[0072]
[0073] Where: is the fusion prediction result; θ0 is the intercept term; θ1 is the influence of the random forest model prediction result on the final prediction value, and θ2 is the influence of the LSTM model prediction result on the final prediction value.
[0074] As a preferred embodiment, the use The second-layer model is trained with the actual observation values in the basin hydrological data; the optimal parameters are obtained by solving the loss function, and the loss function expression is as follows:
[0075]
[0076] Where: is the actual observation value of the ath training sample; is the predicted value of the second layer model for the ath training sample.
[0077] On the other hand, the present invention also provides a basin hydrological forecasting system based on big data, comprising:
[0078] The data collection and preprocessing module collects basin hydrological data through the big data platform and transmits it to the edge computing node deployed at the data collection end in real time to preprocess and extract features of the collected data;
[0079] The data set division module constructs a data set from the basin hydrological data after feature extraction, and divides the data set into a test set and a validation set through spatiotemporal cross-validation;
[0080] The model training module uses the divided training sets to construct and train the random forest model and the long short-term memory network model LSTM, respectively, to obtain the trained random forest model and the long short-term memory network model LSTM;
[0081] The model fusion and optimization module uses the stacking method to fuse the trained random forest model and the long short-term memory network model LSTM to obtain the basin hydrological prediction model, and uses the test set data to evaluate and optimize the basin hydrological prediction model;
[0082] The real-time prediction module transmits the basin hydrological data collected by the big data platform to the edge computing node deployed at the data collection end for preprocessing and feature extraction, and integrates the new feature data into the original data set, and predicts the basin's hydrological data in real time through the basin hydrological prediction model.
[0083] On the other hand, the present invention further provides an electronic device having a computer program stored thereon, wherein when the computer program is executed by a processor, the computer program implements a basin hydrological forecasting method based on big data as described in any embodiment of the present invention.
[0084] On the other hand, the present invention also provides a computer-readable medium for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement a watershed hydrological forecasting method based on big data as described in any embodiment of the present invention.
[0085] The present invention has the following beneficial effects:
[0086] 1. Collect multi-source data through the big data platform, covering meteorology, hydrology, topography, soil and other aspects, and perform real-time preprocessing and feature extraction at the edge computing node, which can quickly process large amounts of data and provide rich and accurate information for the model.
[0087] 2. The Stacking method is used to fuse the random forest model and the long short-term memory network model LSTM, combining the advantages of both. The random forest can handle nonlinear relationships and screen important features, and the LSTM is good at processing time series data, which improves the accuracy and stability of model prediction.
[0088] 3. The newly collected data are processed in real time and integrated into the original data set. Real-time prediction through the optimized basin hydrological prediction model can timely reflect the changes in the basin's hydrological status and provide real-time information for decisions such as flood control and water resources scheduling. BRIEF DESCRIPTION OF THE DRAWINGS
[0089] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0090] Figure 1 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION
[0091] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0092] It should be understood that the step numbers used in this document are only for convenience of description and are not intended to limit the order in which the steps are executed.
[0093] It should be understood that the terms used in the present specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0094] The terms “include” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.
[0095] The term "and / or" means and includes any and all possible combinations of one or more of the associated listed items.
[0096] Embodiment 1:
[0097] In order to make the purpose, technical solutions and advantages of the present invention more clear, the following will be combined with the specific embodiments of the present application and refer to the attached Figure 1 , clearly and completely describe the technical solution of the present invention.
[0098] To solve the problems of the prior art, the present invention provides a basin hydrological forecasting method based on big data, comprising the following steps:
[0099] Step S1, collecting basin hydrological data through the big data platform, and transmitting it to the edge computing node deployed at the data collection end in real time, and preprocessing and feature extraction of the collected data;
[0100] The basin hydrological related data include: meteorological data, hydrological data, topographic data and soil data; among them, meteorological data include: precipitation, temperature, wind speed, wind direction, humidity; hydrological data include: water level, flow, evaporation; topographic data include: terrain elevation, slope, slope direction, land use type covering the basin; soil data include: soil type, texture, porosity.
[0101] The edge computing node selects suitable data acquisition devices and edge computing hardware, such as the Raspberry Pi series or a dedicated edge computing gateway. Configure sufficient memory and storage capacity to meet the needs of real-time data processing and temporary storage. At the same time, ensure that the device has a stable network connection interface, such as Ethernet, Wi-Fi or 4G / 5G communication module, to transmit data with the data center. The software framework uses a lightweight real-time data processing framework, such as Node-RED or OpenCV, to achieve real-time data acquisition, preprocessing and feature extraction.
[0102] Edge computing node preprocessing includes: data cleaning, data conversion, data aggregation and compression; the data cleaning steps are as follows: use filtering algorithms to remove random noise and interference signals in the data; identify and process outliers by setting thresholds, statistical methods (such as the 3σ principle) or machine learning-based anomaly detection algorithms; use interpolation methods such as linear interpolation, spline interpolation, etc., or fill in missing data based on mean filling, regression filling and other methods of similar data. Data conversion includes: converting data of different formats and encodings into a standard format, such as converting all sensor data into a unified JSON or CSV format to facilitate subsequent processing and analysis; using normalization or standardization methods, such as Min-Max normalization, Z-score standardization, etc., to map data to a specific interval or eliminate dimensional differences to make different indicators comparable; data aggregation: aggregating data collected from multiple data sources or the same data source at different times, such as summing, averaging, counting, etc., to reduce the amount of data and extract more meaningful information; data compression: using lossless compression algorithms, such as Huffman coding, LZ77 algorithm, etc., or lossy compression algorithms, such as JPEG, MP3 and other image and audio compression algorithms, to compress data and reduce the cost of data storage and transmission.
[0103] Step S2, forming a data set with the basin hydrological related data after feature extraction, and dividing the data set into a test set and a validation set through spatiotemporal cross-validation;
[0104] The method of dividing the data set into a test set and a validation set by spatiotemporal cross-validation specifically includes the following steps:
[0105] Time division: According to the time characteristics of the basin hydrological data, the time series is divided into k non-overlapping time periods T1, T2, …, T k ;
[0106] Spatial division: According to the spatial characteristics of the watershed, the watershed space is divided into m non-overlapping sub-areas R1, R2, ..., R m ;
[0107] Verification process: In each verification, a time period T is selected i ,i∈k and a subregion R j ,j∈m is used as the validation set, and the rest of the time and space data is used as the training set. By traversing all time and space combinations, the data can be fully utilized and verified in both time and space dimensions.
[0108] Training set: In each partition, the training set contains all spatiotemporal data except the currently selected validation set. This data is used to train the model so that the model learns the patterns and regularities in the data. For example, in the partition with T1 and R1 as the validation set, the training set contains T2-T k Data for all sub-areas within the time period, and R2-R within the T1 time period m Sub-region data.
[0109] Test set: The validation set in each partition is the test set, which is used to evaluate the performance of the model under specific spatiotemporal conditions. The data in the test set does not participate in the model training process, but is only used to test the prediction and generalization capabilities of the model after the model training is completed. For example, when T1 and R1 are used as test sets, the trained model is used to predict the data in the test set, and the prediction results are compared with the actual observations, and the relevant evaluation indicators (such as root mean square error, mean absolute error, etc.) are calculated to evaluate the performance of the model.
[0110] Step S3, using the divided training sets to respectively construct and train a random forest model and a long short-term memory network model LSTM, to obtain a trained random forest model and a long short-term memory network model LSTM;
[0111] In this embodiment, the random forest model and the long short-term memory network model LSTM are constructed and trained respectively using the divided training sets as follows:
[0112] Build and train a random forest model: Extract multiple subsets from the training set using the bootstrap sampling method and build a decision tree for each subset. In the process of building the decision tree, randomly select feature subsets for node splitting to increase the diversity between decision trees. Optimize the performance of the random forest model by adjusting hyperparameters such as the number of decision trees, maximum depth, and feature subset size. Use the training set data to train the random forest model so that it can learn the nonlinear relationships and feature patterns in the data.
[0113] Construct and train the LSTM network model: According to the characteristics of the hydrological time series data, design a suitable LSTM network structure, including determining the number of hidden layers, the number of neurons in each hidden layer, and the dimensions of input and output. Train the LSTM network through the back propagation algorithm and optimizer (such as the Adam optimizer), adjust the network's weights and biases, so that it can effectively capture the long-term dependencies and trend changes in the time series.
[0114] The random forest model outputs the following results:
[0115]
[0116] in:
[0117]
[0118] Where: is the output result of the random forest model; N is the number of decision trees in the random forest model; is the predicted value of sample x by the nth decision tree; w n is the voting weight of the th decision tree; Acc n is the prediction accuracy of the nth decision tree on the training set; Acc o is the prediction accuracy of the oth decision tree on the training set.
[0119] The long short-term memory network model LSTM adopts a multi-modal fusion LSTM model of bidirectional attention LSTM, hierarchical LSTM and skip connection; in the bidirectional LSTM structure, the forward LSTM captures the positive dependency of the sequence, and the backward LSTM captures the reverse dependency of the sequence. In the hierarchical LSTM structure, the bottom-level LSTM processes short-term detail information, and the high-level LSTM processes long-term trend information. By fusing the bottom-level and high-level information through skip connections, the model can capture both short-term fluctuations in the time series and long-term change trends, and more comprehensively describe the characteristics of the hydrological time series. Among them:
[0120] The bidirectional attention LSTM includes:
[0121] Forward LSTM calculation:
[0122] Forget Gate
[0123]
[0124] Input Gate
[0125]
[0126] Candidate memory cells
[0127]
[0128] Memory Unit
[0129]
[0130] Output Gate
[0131]
[0132] Hidden State
[0133]
[0134] Where: and They are the weight matrices of the forget gate, input gate, candidate memory unit, and output gate in the forward calculation respectively; and are the bias vectors of the forget gate, input gate, candidate memory unit, and output gate in the forward calculation; σ is the sigmoid function; x t is the input data at the time step; ⊙ is the element multiplication; tanh() is the hyperbolic tangent function;
[0135] Backward LSTM calculation:
[0136] Forget Gate
[0137]
[0138] Input Gate
[0139]
[0140] Candidate memory cells
[0141]
[0142] Memory Unit
[0143]
[0144] Output Gate
[0145]
[0146] Hidden State
[0147]
[0148] Where: and They are the weight matrices of the forget gate, input gate, candidate memory unit, and output gate in the backward calculation respectively; and They are the bias vectors of the forget gate, input gate, candidate memory unit, and output gate in the backward calculation respectively;
[0149] The attention weight calculation formula is as follows:
[0150]
[0151] Where: are the weights of the forward and backward attention mechanisms, respectively; are the weight matrices of the forward and backward attention mechanisms respectively;
[0152] The hidden state of the forward LSTM and the hidden state of the backward LSTM are merged. The specific formula is as follows:
[0153]
[0154] Where: h t is the hidden state after fusion;
[0155] Hierarchical LSTM with skip connections includes:
[0156] The bottom LSTM and high-level LSTM in the hierarchical LSTM are based on the same LSTM basic structure as the forward LSTM and backward LSTM. t As the input of the underlying LSTM, the output is the hidden state h of the underlying LSTM t1 ; will h t1 As the input of the high-level LSTM, the output is the hidden state h of the high-level LSTM t2 ; will h t1 and h t2 Perform a jump connection to obtain the final fused hidden state. The specific formula is as follows:
[0157]
[0158] The final fused hidden state is converted into a predicted value through the output layer. The specific formula is as follows:
[0159]
[0160] Where: Output the predicted value for the LSTM model; W y is the dimension-adapted weight matrix; b y is the bias vector of the output layer.
[0161] Step S4, using the stacking method to fuse the trained random forest model and the long short-term memory network model LSTM to obtain a watershed hydrological prediction model, and using the test set data to evaluate and optimize the watershed hydrological prediction model;
[0162] The stacking method is used to fuse the trained random forest model and the long short-term memory network model LSTM to obtain a watershed hydrological prediction model, which includes:
[0163] The first layer model predicts:
[0164] Use the trained random forest and LSTM models to predict the training set and test set; the prediction results of the random forest model for the training set and test set are and The prediction results of the LSTM model for the training set and the test set are and
[0165] Second layer fusion model training:
[0166] Will and As new features, form a new training data set Will and As new features, form a new test data set
[0167] Second layer model training and prediction:
[0168] use The second layer model is trained with actual observations from the basin hydrological data; and the new test data set is used The final fusion prediction result is obtained. The specific formula is as follows:
[0169]
[0170] Where: is the fusion prediction result; θ0 is the intercept term; θ1 is the influence of the random forest model prediction result on the final prediction value, and θ2 is the influence of the LSTM model prediction result on the final prediction value.
[0171] The use The second-layer model is trained with the actual observation values in the basin hydrological data; the optimal parameters are obtained by solving the loss function, and the loss function expression is as follows:
[0172]
[0173] Where: is the actual observation value of the ath training sample; is the predicted value of the second layer model for the ath training sample.
[0174] Step S5, the basin hydrological data collected by the big data platform is transmitted in real time to the edge computing node deployed at the data collection end for preprocessing and feature extraction, and the new feature data is integrated into the original data set, and the basin hydrological data is predicted in real time through the basin hydrological prediction model.
[0175] Embodiment 2:
[0176] This embodiment provides a watershed hydrological forecasting system based on big data, including:
[0177] The data collection and preprocessing module collects basin hydrological data through the big data platform and transmits it to the edge computing node deployed at the data collection end in real time to preprocess and extract features of the collected data;
[0178] The data set division module constructs a data set from the basin hydrological data after feature extraction, and divides the data set into a test set and a validation set through spatiotemporal cross-validation;
[0179] The model training module uses the divided training sets to construct and train the random forest model and the long short-term memory network model LSTM, respectively, to obtain the trained random forest model and the long short-term memory network model LSTM;
[0180] The model fusion and optimization module uses the stacking method to fuse the trained random forest model and the long short-term memory network model LSTM to obtain the basin hydrological prediction model, and uses the test set data to evaluate and optimize the basin hydrological prediction model;
[0181] The real-time prediction module transmits the basin hydrological data collected by the big data platform to the edge computing node deployed at the data collection end for preprocessing and feature extraction, and integrates the new feature data into the original data set, and predicts the basin's hydrological data in real time through the basin hydrological prediction model.
[0182] Embodiment three:
[0183] This embodiment provides an electronic device having a computer program stored thereon, wherein when the computer program is executed by a processor, a method for watershed hydrological forecasting based on big data as described in any embodiment of the present invention is implemented.
[0184] Embodiment 4:
[0185] This embodiment provides a computer-readable medium for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement a watershed hydrological forecasting method based on big data as described in any embodiment of the present invention.
[0186] In the embodiments of the present application, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone. Among them, A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b and c can be represented by: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, c can be single or multiple.
[0187] Those of ordinary skill in the art will appreciate that the various units and algorithm steps described in the embodiments disclosed herein can be implemented in a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0188] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0189] In several embodiments provided in the present application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory; hereinafter referred to as: ROM), random access memory (Random Access Memory; hereinafter referred to as: RAM), disk or optical disk, and other media that can store program codes.
[0190] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A basin hydrological forecasting method based on big data, characterized in that: The following steps are involved: Step S1, collecting basin hydrological data through the big data platform, and transmitting it to the edge computing node deployed at the data collection end in real time, and preprocessing and feature extraction of the collected data; Step S2, forming a data set with the basin hydrological related data after feature extraction, and dividing the data set into a test set and a validation set through spatiotemporal cross-validation; Step S3, using the divided training sets to respectively construct and train a random forest model and a long short-term memory network model LSTM, to obtain a trained random forest model and a long short-term memory network model LSTM; Step S4, using the stacking method to fuse the trained random forest model and the long short-term memory network model LSTM to obtain a watershed hydrological prediction model, and using the test set data to evaluate and optimize the watershed hydrological prediction model; Step S5, the basin hydrological data collected by the big data platform is transmitted in real time to the edge computing node deployed at the data collection end for preprocessing and feature extraction, and the new feature data is integrated into the original data set, and the basin hydrological data is predicted in real time through the basin hydrological prediction model.
2. The method for basin hydrological forecasting based on big data according to claim 1, characterized in that: The basin hydrological related data include: meteorological data, hydrological data, topographic data and soil data; among them, meteorological data include: precipitation, temperature, wind speed, wind direction, humidity; hydrological data include: water level, flow, evaporation; topographic data include: terrain elevation, slope, slope direction, land use type covering the basin; soil data include: soil type, texture, porosity.
3. The method for basin hydrological forecasting based on big data according to claim 1, characterized in that: The method of dividing the data set into a test set and a validation set by spatiotemporal cross-validation specifically includes the following steps: Time division: According to the time characteristics of the basin hydrological data, the time series is divided into k non-overlapping time periods T1, T2, …, T k ; Spatial division: According to the spatial characteristics of the watershed, the watershed space is divided into m non-overlapping sub-areas R1, R2, ..., R m ; Verification process: In each verification, a time period T is selected i ,i∈k and a subregion R j ,j∈m is used as the validation set, and the rest of the time and space data is used as the training set. By traversing all time and space combinations, the data can be fully utilized and verified in both time and space dimensions.
4. The method for basin hydrological forecasting based on big data according to claim 1, characterized in that: The random forest model outputs the following results: in: Where: is the output result of the random forest model; N is the number of decision trees in the random forest model; is the predicted value of sample x by the nth decision tree; w n is the voting weight of the th decision tree; Acc n is the prediction accuracy of the nth decision tree on the training set; Acc o is the prediction accuracy of the oth decision tree on the training set.
5. The method for basin hydrological forecasting based on big data according to claim 1, characterized in that: The long short-term memory network model LSTM adopts a multi-modal fusion LSTM model of bidirectional attention LSTM, hierarchical LSTM and skip connection, in which: The bidirectional attention LSTM includes: Forward LSTM calculation: Forget Gate Input Gate Candidate memory cells Memory Unit Output Gate Hidden State Where: and They are the weight matrices of the forget gate, input gate, candidate memory unit, and output gate in the forward calculation respectively; and are the bias vectors of the forget gate, input gate, candidate memory unit, and output gate in the forward calculation; σ is the sigmoid function; x t is the input data at the time step; ⊙ is the element multiplication; tanh() is the hyperbolic tangent function; Backward LSTM calculation: Forget Gate Input Gate Candidate memory cells Memory Unit Output Gate Hidden State Where: and They are the weight matrices of the forget gate, input gate, candidate memory unit, and output gate in the backward calculation respectively; and They are the bias vectors of the forget gate, input gate, candidate memory unit, and output gate in the backward calculation respectively; The attention weight calculation formula is as follows: Where: are the weights of the forward and backward attention mechanisms, respectively; are the weight matrices of the forward and backward attention mechanisms respectively; The hidden state of the forward LSTM and the hidden state of the backward LSTM are merged. The specific formula is as follows: Where: h t is the hidden state after fusion; Hierarchical LSTM with skip connections includes: The bottom LSTM and high-level LSTM in the hierarchical LSTM are based on the same LSTM basic structure as the forward LSTM and backward LSTM. t As the input of the underlying LSTM, the output is the hidden state h of the underlying LSTM t1 ; will h t1 As the input of the high-level LSTM, the output is the hidden state h of the high-level LSTM t2 ; will h t1 and h t2 Perform a jump connection to obtain the final fused hidden state. The specific formula is as follows: The final fused hidden state is converted into a predicted value through the output layer. The specific formula is as follows: Where: Output the predicted value for the LSTM model; W y is the dimension-adapted weight matrix; b y is the bias vector of the output layer.
6. The method for basin hydrological forecasting based on big data according to claim 1, characterized in that: The stacking method is used to fuse the trained random forest model and the long short-term memory network model LSTM to obtain a watershed hydrological prediction model, which includes: The first layer model predicts: Use the trained random forest and LSTM models to predict the training set and test set; the prediction results of the random forest model for the training set and test set are and The prediction results of the LSTM model for the training set and the test set are and Second layer fusion model training: Will and As new features, form a new training data set Will and As new features, form a new test data set Second layer model training and prediction: use The second layer model is trained with actual observations from the basin hydrological data; and the new test data set is used The final fusion prediction result is obtained. The specific formula is as follows: Where: is the fusion prediction result; θ0 is the intercept term; θ1 is the influence of the random forest model prediction result on the final prediction value, and θ2 is the influence of the LSTM model prediction result on the final prediction value.
7. The method for basin hydrological forecasting based on big data according to claim 6 is characterized by: The use The second-layer model is trained with the actual observation values in the basin hydrological data; the optimal parameters are obtained by solving the loss function, and the loss function expression is as follows: Where: is the actual observation value of the ath training sample; is the predicted value of the second layer model for the ath training sample.
8. A basin hydrological forecasting system based on big data, characterized in that: include: The data collection and preprocessing module collects basin hydrological data through the big data platform and transmits it to the edge computing node deployed at the data collection end in real time to preprocess and extract features of the collected data; The data set division module constructs a data set from the basin hydrological data after feature extraction, and divides the data set into a test set and a validation set through spatiotemporal cross-validation; The model training module uses the divided training sets to construct and train the random forest model and the long short-term memory network model LSTM, respectively, to obtain the trained random forest model and the long short-term memory network model LSTM; The model fusion and optimization module uses the stacking method to fuse the trained random forest model and the long short-term memory network model LSTM to obtain the basin hydrological prediction model, and uses the test set data to evaluate and optimize the basin hydrological prediction model; The real-time prediction module transmits the basin hydrological data collected by the big data platform to the edge computing node deployed at the data collection end for preprocessing and feature extraction, and integrates the new feature data into the original data set, and predicts the basin's hydrological data in real time through the basin hydrological prediction model.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method for watershed hydrological forecasting based on big data as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, a basin hydrological forecasting method based on big data as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Hydrological forecasting method based on hydrological information big data of regional drainage basin
CN116643331A
Cited By
Hydrological flow monitoring method and system based on dynamic topology adaptive technology
CN121067814A
Hydrological flow monitoring method and system based on dynamic topology adaptive technology
CN121067814B