Clustering short-term runoff prediction method based on encoder-decoder and attention mechanism

By clustering runoff data and combining the CA-AGRU model with an encoder-decoder and attention mechanism, the shortcomings of existing runoff prediction models in capturing complex correlations are addressed, achieving runoff prediction with higher accuracy and reliability.

CN120670882APending Publication Date: 2025-09-19HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510834699.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-09-19

Smart Images

  • Figure CN120670882A_ABST
    Figure CN120670882A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of hydrological prediction, and particularly discloses a clustering short-term runoff prediction method based on an encoder-decoder and an attention mechanism, and the method comprises the steps: carrying out the clustering of historical runoff data, obtaining a plurality of clusters, enabling each cluster to represent a sub-data set, and enabling the sub-data set to be a data set composed of the same kind of data; based on each sub-data set, training an attention time sequence model, and obtaining a trained attention time sequence model corresponding to each cluster; wherein the attention time sequence model is constructed through a cascade encoder, an attention module and a decoder, and the attention time sequence model is used for predicting the runoff data after the current time step based on the runoff data before the current time step. According to the method, the precision and reliability of runoff prediction can be effectively improved by avoiding mutual interference of different hydrological mechanisms and capturing key driving factors in similar data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of hydrological prediction technology, and more specifically, relates to a clustering short-term runoff prediction method based on encoder-decoder and attention mechanism. Background Art

[0002] Runoff forecasts can help better understand future trends in river flow, guiding hydrological forecasting, flood prevention and mitigation, and water resources planning. Furthermore, the precision and accuracy of runoff forecasts are directly related to the rational utilization and management of water resources, and are crucial for ensuring their security and stability. Due to the influence of natural meteorological conditions, river basin characteristics, and human activities, runoff series exhibit characteristics such as trends, randomness, and periodicity. Finding more accurate runoff prediction methods has always been a key technical issue in this field. Summary of the Invention

[0003] In view of the shortcomings of the existing technology, the purpose of this application is to improve the accuracy of runoff prediction.

[0004] To achieve the above objectives, in a first aspect, the present application provides a clustering short-term runoff prediction method based on an encoder-decoder and an attention mechanism, the method comprising: Cluster the historical runoff data to obtain multiple clusters, each cluster represents a sub-dataset, and the sub-dataset is a data set composed of similar data; Based on each sub-dataset, train the attention timing model and obtain the trained attention timing model corresponding to each cluster; Among them, the attention timing model is constructed by cascading encoders, attention modules and decoders. The attention timing model is used to predict the runoff data after the current time step based on the runoff data before the current time step.

[0005] It can be understood that this application proposes a new runoff prediction framework - Clustering Enhanced Attention Time Series Runoff Prediction Model (CA-AGRU), and applies this model to runoff prediction.

[0006] Specifically, by clustering historical runoff data, multiple clusters are obtained, each cluster represents a sub-dataset, which can characterize a hydrological mechanism. Then, based on each sub-dataset, an attention time series model is trained to obtain the trained attention time series model corresponding to each cluster, that is, the trained attention time series model corresponding to each hydrological mechanism is obtained, which can effectively avoid the mutual interference of different hydrological mechanisms; at the same time, the attention time series model is constructed by cascading encoders, attention modules and decoders, making full use of the advantages of attention mechanism and encoder-decoder, and can more accurately capture the key driving factors in similar data, so that the model can focus on the key features that have the greatest impact on the prediction results during the prediction process, thereby overcoming the limitation of existing models that are difficult to capture the complex correlations between runoff data.

[0007] Therefore, this application can effectively improve the accuracy and reliability of runoff prediction by combining the above-mentioned technical means of avoiding mutual interference between different hydrological mechanisms and the above-mentioned technical means of capturing key driving factors in similar data.

[0008] In a possible implementation, the encoder and the decoder use a gated recurrent neural network GRU.

[0009] In one possible implementation, the attention module is used to determine : ; in, It is The decoder in the attention temporal model corresponding to the cluster (referred to as the Class decoder) in the step attention context vector, For the The attention temporal model corresponding to the cluster (abbreviated as The encoder in the attention-based temporal model Step 1 for the decoder The attention weight of the step, For the The encoder in the attention temporal model corresponding to the cluster (referred to as the Class encoder) in the The hidden state of the step, For the The sequence length of the sub-dataset represented by the cluster.

[0010] In one possible implementation, It is determined by the following formula: ; in, For the The encoder of the attention temporal model corresponding to the cluster Step 1 for the decoder The similarity score of the step, For the The encoder of the attention temporal model corresponding to the cluster Step 1 for the decoder The similarity score of the step, is a variable.

[0011] In one possible implementation, It is determined by the following formula: ; in, is the rating vector, represents the hyperbolic tangent activation function, To query the transformation matrix, is the key transformation matrix, For the The decoder of the attention temporal model corresponding to the cluster is in the The hidden state at the previous time step.

[0012] In one possible implementation, the decoder is configured to 、 and ,Sure ; in, For the The decoder of the attention temporal model corresponding to the cluster is in the The output of the step (previous time step), For the The decoder of the attention temporal model corresponding to the cluster is in the The hidden state at the previous time step, For the The decoder of the attention temporal model corresponding to the cluster is in the Output of the current time step.

[0013] In a possible implementation, the decoder is specifically configured to determine : ; ; in, The formula that represents the gated recurrent neural network to determine the hidden state of the current time step, For the The decoder of the attention temporal model corresponding to the cluster is in the The hidden state of the step, is the weight matrix, is the prediction step length.

[0014] In a possible implementation, the clustering method is specifically K-means clustering.

[0015] In one possible implementation, after training the attention temporal model, the following steps are also included: Obtain new runoff data; Based on the Euclidean distance between the new runoff data and the cluster centers of each cluster, the target cluster to which the new runoff data belongs is determined by finding the minimum Euclidean distance; Input new runoff data into the attention time series model corresponding to the target cluster and obtain the runoff prediction result output by the attention time series model; Among them, the new runoff data is the runoff data before the current time step, and the runoff prediction result represents the runoff data after the current time step predicted by the attention time series model.

[0016] In a second aspect, the present application provides an electronic device comprising: at least one memory for storing programs; and at least one processor for executing the programs stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method described in the first aspect or any possible implementation of the first aspect.

[0017] It can be understood that the beneficial effects of the second aspect mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here.

[0018] In general, the above technical solutions conceived by this application have the following beneficial effects compared with the existing technologies: (1) This application proposes a CA-AGRU comprehensive prediction model, which combines the K-means clustering method and the encoder-decoder GRU network based on the attention mechanism. It has better adaptability to the nonlinear and non-stationary characteristics of runoff sequences, has stronger generalization ability, and can effectively improve the prediction accuracy.

[0019] (2) Experiments show that the CA-AGRU model proposed in this application has two major advantages over the baseline model: first, the CA-AGRU model is more accurate in predicting runoff, and the error between the model and the measured runoff is smaller; second, in the case of sudden changes in runoff, the CA-AGRU model proposed in this application can adaptively focus on key hydrological periods and show better prediction performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1This is an overall architecture diagram of clustered short-term runoff prediction based on encoder-decoder and attention mechanism provided by an embodiment of the present application; Figure 2 This is a workflow diagram of the CA-AGRU model provided in the embodiments of the present application; Figure 3 This is a comparison chart of the LSTM and CA-AGRU predictions of daily runoff at Dongjing Station in 2019 provided in an embodiment of the present application; Figure 4 This is a comparison chart of the daily runoff GRU and CA-AGRU predictions for Dongjing Station in 2019 provided in the embodiments of the present application; Figure 5 This is a comparison chart of the AGRU and CA-AGRU predictions of daily runoff at Dongjing Station in 2019 provided in the embodiments of the present application; Figure 6 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0021] In order to facilitate a clearer understanding of the various embodiments of the present application, some relevant background knowledge is first introduced as follows.

[0022] Deep learning technology has made significant progress in the field of runoff forecasting, greatly improving the accuracy and reliability of runoff forecasting. Currently, the application of deep learning in hydrological-runoff forecasting focuses on time series modeling algorithms. While these algorithms can capture the temporal dependencies of data, they often fail to fully consider the complex interactions and interferences between different hydrological mechanisms. Furthermore, while the introduction of clustering methods into deep learning models can, to a certain extent, classify data and uncover potential patterns in the data, the traditional combination of clustering and deep learning still faces problems such as insufficient feature extraction and loss of key information, making it difficult for the model to accurately capture the inherent laws of runoff data.

[0023] Therefore, this application proposes a clustered short-term runoff prediction model based on an encoder-decoder architecture and an attention mechanism. This model considers the non-stationary and multi-scale characteristics of runoff data and incorporates an attention mechanism into the modeling process, enabling the model to focus on the key features that have the greatest impact on the prediction results. This overcomes the limitation of existing models that struggle to capture the complex relationships between runoff data, providing a technical reference for short-term runoff prediction.

[0024] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0025] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0026] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more, for example, multiple processing units means two or more processing units, etc.; multiple elements means two or more elements, etc.

[0027] The embodiments of the present application are described below in conjunction with the drawings in the embodiments of the present application.

[0028] The embodiments of this application cluster historical runoff data to generate multiple stable subsets. They then train specialized runoff prediction models for different subsets. They also incorporate an attention mechanism and encoder-decoder structure into the prediction model to enhance the model's learning of spatiotemporal features and alleviate the vanishing gradient problem of deep learning models. The main components include: (1) overall architecture, (2) K-means clustering, (3) attention time series model (AGRU), (4) classification prediction, and (5) comparative experiments.

[0029] (1) Overall architecture; like Figure 1 As shown in the figure, the model consists of K-means clustering and the attention time series model AGRU. AGRU includes an encoder GRU, an attention module, and a decoder GRU. GRU (Gated Recurrent Unit) is a gated recurrent neural network.

[0030] (2) K-means clustering; The K-means clustering method is used to cluster the runoff data of the target basin. Input sequence ,in is the feature dimension, Represents the input sequence The tth data point in the t, t ranges from 1 to T, each All included dimensional features, including rainfall data from weather stations and runoff data from reservoirs; is the input sequence length.

[0031] K-means is an unsupervised clustering algorithm based on distance, which divides the data into non-overlapping clusters, so that each data point belongs to the cluster corresponding to the nearest cluster center. First, randomly select data points as the initial cluster center, and then for each point in the data set, calculate its The Euclidean distance between cluster centers is calculated and assigned to the cluster corresponding to the nearest cluster center.

[0032] ; in, Refers to data points No. dimensional eigenvalues, It is The centroid of the cluster, Refers to the center of mass No. dimensional eigenvalues.

[0033] According to the current cluster assignment, the center of each cluster is recalculated, that is, the mean of all points in the cluster is calculated as the new cluster center, as shown in the following formula: ; in, Refers to A cluster contains all the data points assigned to the cluster. Indicates the The number of data points in a cluster.

[0034] Repeat the above steps until the cluster center no longer changes or the maximum number of iterations is reached. Final output subsets.

[0035] ; in, It is The sub-dataset of the class, is the classification result.

[0036] (3) Attentional Temporal Model AGRU; The attention timing model AGRU proposed in this application includes an encoder GRU, an attention module and a decoder GRU.

[0037] (3.1) Encoder GRU; The encoder can encode the input sequence into a fixed-length context vector (context vector), which captures all the important information of the input sequence. This application uses GRU to implement this function. Sub-dataset of class Input is processed in the GRU-based encoder to obtain all hidden states : ; in, For the The class encoder is at time step The hidden state of For the The sequence length of the class sub-dataset.

[0038] (3.2) Attention module; In this application, after the encoder completes encoding, the final hidden state is not directly input into the decoder. Instead, it is transformed through the attention mechanism to generate the context variables required by the decoder. The attention mechanism dynamically assigns weights to different input elements, allowing the model to focus on the most relevant information while ignoring other less important information. As a result, the context vector generated by the attention mechanism can more accurately capture the key driving factors in similar data.

[0039] The attention mechanism takes a query and a set of key-value pairs as input, first calculates the weight coefficients of different value vectors using the correlation between the current query vector and different key vectors, and then performs a weighted summation of the value vectors according to the weight coefficients to obtain the final attention value. Step 1, calculate its context vector using the following formula: ; ; ; in, For the The encoder in the encoder-decoder class Step 1 for the decoder The similarity score of the step, is the score vector, which is a learnable parameter vector. represents the hyperbolic tangent activation function, To query the transformation matrix, is the key transformation matrix. For the The hidden state of the decoder at the previous time step, For the The class encoder is at time step The hidden state of . For the The encoder in the encoder-decoder class Step 1 for the decoder The attention weight of the step, It is a temporary variable with a value range of 1~ . It is Class Decoder The attention context vector of the step.

[0040] (3.3) Decoder GRU; The decoder generates an output sequence based on the context vector. At each time step, the decoder generates the output of the current time step based on the output of the previous time step, the hidden state of the previous time step, and the current context vector. In this application, the runoff value of the last time step in the input sequence is taken. As the initial input to the GRU-based decoder , the last hidden state output by the encoder As the initial state of the GRU-based decoder , after GRU calculation, the state of the decoder is updated , and get the Class decoder Output prediction of step : ; ; in, is the weight matrix, is the prediction step length, For the Class decoder The final hidden state of the step.

[0041] Loop through the above steps The final prediction output of the model can be obtained : .

[0042] (4) Classification prediction; After the training of N subsets is completed, N independent prediction models will be obtained. Figure 2 As shown, for the new input data First, calculate the Euclidean distance between it and the centroid of each category to determine the category n to which the new data belongs, and then call the corresponding nth model to perform runoff prediction and output the final runoff prediction value.

[0043] (5) Comparative experiments; The following comparative experimental analysis is conducted using the Beipanjiang River Basin as the experimental object. The Beipanjiang River is a major tributary of the Hongshui River, the upper source of the Xijiang River in the Pearl River Basin. It originates from the northwest foot of Ma Xiong Mountain in the Wumeng Mountains in Zhanyi District, Yunnan Province, and flows through Yunnan and Guizhou Provinces. It is 449 kilometers long, with a total drop of 1,985 meters and an average gradient of 4.42‰. The basin area is 26,557 square kilometers. The long-term average flow at the estuary is 390 cubic meters per second, and the long-term average runoff in the basin is 384.9 million cubic meters. This application selects the water inlet section of the Dongjing Hydropower Station as the main research object. The regional dataset in this application contains runoff data collected by 4 hydropower stations and rainfall data measured by 6 meteorological stations, including daily data for five years from January 1, 2015 to December 31, 2019. The four hydropower stations are all on the main stream of the Beipanjiang River. From upstream to downstream, they are Shannipo Hydropower Station, Guangzhao Hydropower Station, Mamaya Hydropower Station and Dongjing Hydropower Station. Six meteorological stations are distributed around the hydropower station. Rainfall falls to the surface and forms surface runoff, which then flows into rivers.

[0044] (5.1) Baseline model: This example selects three models for comparison, including the long short-term memory neural network (LSTM), the gated recurrent unit (GRU), and the independent attention temporal model (AGRU).

[0045] (5.2) Model hyperparameter setting: The algorithm program is written in Python. For the total data set of 5 years and each sub-data set after classification, the data of the first four years are used as the training set, and the data of the remaining year are used as the test set. The maximum number of iterations for the LSTM, GRU and AGRU models is 100, and the maximum number of iterations for the CA-AGRU model is 200. The number of neural network layers in all models is 3, the number of neurons in each layer is 64, the exit probability of each neural network layer is 20%, and all use the Adam optimizer, the loss function is the mean squared error (MSE), the learning rate is 0.001, and the batch size of each group of samples is 64. All experiments are repeated three times and the average results are taken. In this application, all models are implemented using the "pytorch1.9.0" framework in python and run in the python 3.8 environment.

[0046] (5.3) Forecast accuracy evaluation indicators: These metrics include the Nash-Sutcliffe Efficiency coefficient (NSE), the Root Mean Square Error (RMSE), and the Mean Absolute Error (MAE). All metrics are calculated for each forecast window and for the entire rolling dataset. Figure 3This is a comparison chart of the LSTM and CA-AGRU predictions of daily runoff at Dongjing Station in 2019 provided in the embodiment of this application. Figure 4 This is a comparison chart of the daily runoff GRU and CA-AGRU predictions for Dongjing Station in 2019 provided in the embodiment of this application. Figure 5 This is a comparison chart of the AGRU and CA-AGRU forecasts of daily runoff at Dongjing Station in 2019 provided in the embodiments of the present application.

[0047] Table 1 Prediction index results of various models in the daily runoff forecast process at Dongqing Station

[0048] (5.4) Result analysis: As shown in Table 1, the proposed CA-AGRU model achieves the best prediction performance compared to the baseline models, with an RMSE of 108.3034 m³ / s, a MAE of 66.31 m³ / s, and a NSE of 0.8172. Compared to the worst-performing baseline model, the LSTM model, CA-AGRU significantly reduces its RMSE and MAE by 17.84% and 10.79%, respectively, while improving its NSE by 8.80%. Compared to the best-performing baseline model, the AGRU model, CA-AGRU still reduces its RMSE and MAE by 11.31% and 3.83%, respectively, while improving its NSE by 4.96%. This result fully demonstrates the superiority of the CA-AGRU model in runoff prediction.

[0049] Furthermore, a comparison of the performance of the AGRU and GRU models reveals that the introduction of the encoder-decoder architecture and attention mechanism to the GRU reduces the model's RMSE and MAE by 1.96% and 2.62%, respectively, while improving the NSE by 0.94%. Although the improvements are relatively small, the AGRU can identify key hydrological periods in runoff sequences and enhance its ability to capture hydrological dynamics. When runoff undergoes sudden changes, the AGRU more accurately reflects the nonlinear characteristics of the runoff process than the GRU, resulting in better forecasting performance. This suggests that the synergistic effect of the attention mechanism and the encoder-decoder architecture can optimize forecasting performance in hydrological time series forecasting, but their individual contributions are still limited by the strong nonlinearity and stochasticity of the hydrological process itself.

[0050] Based on the method in the above embodiment, an embodiment of the present application provides an electronic device, Figure 6 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application, such as Figure 6As shown, the electronic device may include: a processor (Processor) 810, a communication interface (Communications Interface) 820, a memory (Memory) 830, and a communication bus 840. The processor 810, the communication interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 may call the logic instructions in the memory 830 to execute the method in the above embodiment.

[0051] In addition, the logic instructions in the aforementioned memory 830 can be implemented in the form of a software functional unit and, when sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application.

[0052] Based on the method in the above embodiment, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the method in the above embodiment.

[0053] Based on the method in the above embodiment, an embodiment of the present application provides a computer program product. When the computer program product runs on a processor, the processor executes the method in the above embodiment.

[0054] It is understood that the processor in the embodiments of the present application may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.

[0055] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and storage medium can be located in an ASIC.

[0056] The above embodiments can be implemented in whole or in part using software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product. The computer program product comprises one or more computer instructions. When loaded and executed on a computer, the computer program instructions fully or partially produce the processes or functions described in the embodiments of this application. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be magnetic media (e.g., floppy disk, hard disk, tape), optical media (e.g., DVD), or semiconductor media (e.g., solid-state drive (SSD)).

[0057] It will be understood that the various numerical numbers involved in the embodiments of the present application are merely distinctions for the convenience of description and are not intended to limit the scope of the embodiments of the present application.

[0058] It is easy for those skilled in the art to understand that the above is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the scope of protection of the present application.

Claims

1. A clustering short-term runoff prediction method based on encoder-decoder and attention mechanism, characterized in that: include: Cluster the historical runoff data to obtain multiple clusters, each cluster represents a sub-dataset, and the sub-dataset is a data set composed of similar data; Based on each sub-dataset, train the attention timing model and obtain the trained attention timing model corresponding to each cluster; Among them, the attention timing model is constructed by cascading encoders, attention modules and decoders. The attention timing model is used to predict the runoff data after the current time step based on the runoff data before the current time step.

2. The clustering short-term runoff prediction method based on encoder-decoder and attention mechanism according to claim 1 is characterized in that: The encoder and decoder use gated recurrent neural network GRU.

3. The clustering short-term runoff prediction method based on encoder-decoder and attention mechanism according to claim 2 is characterized in that: The attention module is used to determine the : ; in, It is The decoder of the attention temporal model corresponding to the cluster is in the step attention context vector, For the The encoder of the attention temporal model corresponding to the cluster Step 1 for the decoder The attention weight of the step, For the The encoder in the attention temporal model corresponding to the cluster is in the The hidden state of the step, For the The sequence length of the sub-dataset represented by the cluster.

4. The clustering short-term runoff prediction method based on encoder-decoder and attention mechanism according to claim 3 is characterized in that: It is determined by the following formula: ; in, For the The encoder of the attention temporal model corresponding to the cluster Step 1 for the decoder The similarity score of the step, For the The encoder of the attention temporal model corresponding to the cluster Step 1 for the decoder The similarity score of the step, is a variable.

5. The clustering short-term runoff prediction method based on encoder-decoder and attention mechanism according to claim 4 is characterized in that: It is determined by the following formula: ; in, is the rating vector, represents the hyperbolic tangent activation function, To query the transformation matrix, is the key transformation matrix, For the The decoder of the attention temporal model corresponding to the cluster is in the The hidden state of the step.

6. The clustering short-term runoff prediction method based on encoder-decoder and attention mechanism according to claim 3 is characterized in that: The decoder is used based on 、 and ,Sure ; in, For the The decoder of the attention temporal model corresponding to the cluster is in the The output of the step, For the The decoder of the attention temporal model corresponding to the cluster is in the The hidden state of the step, For the The decoder of the attention temporal model corresponding to the cluster is in the Step output.

7. The clustering short-term runoff prediction method based on encoder-decoder and attention mechanism according to claim 6 is characterized in that: The decoder is specifically used to determine the following formula : ; ; in, The formula that represents the gated recurrent neural network to determine the hidden state of the current time step, For the The decoder of the attention temporal model corresponding to the cluster is in the The hidden state of the step, is the weight matrix, is the prediction step length.

8. The clustering short-term runoff prediction method based on encoder-decoder and attention mechanism according to claim 1 is characterized in that: The specific clustering method is K-means clustering.

9. The clustering short-term runoff prediction method based on encoder-decoder and attention mechanism according to any one of claims 1 to 8, characterized in that: After training the attention temporal model, it also includes: Obtain new runoff data; Based on the Euclidean distance between the new runoff data and the cluster centers of each cluster, the target cluster to which the new runoff data belongs is determined by finding the minimum Euclidean distance; Input new runoff data into the attention time series model corresponding to the target cluster and obtain the runoff prediction result output by the attention time series model; Among them, the new runoff data is the runoff data before the current time step, and the runoff prediction result represents the runoff data after the current time step predicted by the attention time series model.

10. An electronic device, characterized in that: include: at least one memory for storing a computer program; At least one processor is used to execute the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method according to any one of claims 1 to 9.