Modeling Method, Device, Terminal and Storage Medium for Operating Characteristics of Pumped-Storage Units
Through the method of segmenting the data set and iterating the training and verification, a pumped storage unit operation characteristic model was constructed, which solved the problem of large model deviation in the existing technology and achieved a better fitting effect.
Patent Information
- Application Number
- CN202410375922.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-29
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-03-29
AI Technical Summary
In the prior art, the operating characteristic model of the pumped storage unit has a large deviation, the construction process is complicated, and the interpolation point is large deviation from the actual operation.
By acquiring multiple data sets, dividing them into training groups and verification groups, the basic model expresses the relationship between the operating characteristics of the turbine unit and the influencing factors, and through repeated iterative training and verification, a turbine unit operation characteristic model with a deviation of less than the characteristic threshold from multiple data sets is constructed.
It effectively prevents overfitting, can output accurate operation characteristics, and can also accurately fit data other than existing data, and the fitting effect is better than interpolation method.
Smart Images

Figure CN118364708B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pumped - storage energy, and in particular, to a method, device, terminal and storage medium for modeling the operating characteristics of a pumped - storage unit. Background Art
[0002] Pumped - storage energy is a kind of energy storage technology. That is, water is used as the energy storage medium, and through the mutual conversion of electrical energy and potential energy, the storage and management of electrical energy are realized. Electrical energy during low - load periods of the power grid is used to pump water to the upper reservoir, and then the water is released to the lower reservoir for power generation during peak - load periods of the power grid. The excess electrical energy during low - load periods of the power grid can be converted into high - value electrical energy during peak - load periods of the power grid. It is applicable to frequency modulation and phase modulation, stabilizing the cycle and voltage of the power system, and can also improve the efficiency of thermal power plants and nuclear power plants in the system. For new energy such as wind power generation and solar power generation, it can make up for the volatility brought by new energy, absorb the surplus electrical energy of new energy power generation, and its significance is even greater.
[0003] Pumped - storage units can be divided into fixed - speed pumped - storage units and variable - speed pumped - storage units from the operating mode. The variable - speed pumped - storage unit has the advantage of high conversion efficiency compared with the fixed - speed pumped - storage unit, so it is widely used.
[0004] When a variable - speed pumped - storage unit is used in combination with new energy, it is always expected that the pumped - storage unit can absorb the electrical energy of new energy power generation according to the volatility of new energy output, reduce the volatility of power supply in the power system, and achieve a high energy conversion efficiency during pumping - storage or power generation.
[0005] To make the pumped - storage unit meet the above requirements, the control strategy of the general pumped - storage unit is formulated in combination with the operating characteristics of the unit. Since the pumped - storage unit switches between the pump and turbine modes and is affected by many factors in the system, its operating characteristics are relatively complex. In the prior art, the characteristic curve formed by multiple experimental data points is usually planned through experimental data, and then the operating characteristics of the pumped - storage unit are determined by interpolation. This method usually uses a small amount of data in the construction process, but the construction process is complex, and the deviation between the interpolation points and the actual operation is large.
[0006] Based on this, it is necessary to develop and design a method for modeling the operating characteristics of a pumped - storage unit. Summary of the Invention
[0007] Embodiments of the present invention provide a method, device, terminal and storage medium for modeling the operating characteristics of a pumped - storage unit, which are used to solve the problem of large deviation in the operating characteristics model of the pumped - storage unit in the prior art.
[0008] In a first aspect, embodiments of the present invention provide a method for modeling the operating characteristics of a pumped - storage unit, including:
[0009] Obtain multiple datasets, where each dataset includes multiple characteristic data representing the operating characteristics of a water turbine unit and multiple factor data representing the factors affecting the operating characteristics of the water turbine unit;
[0010] Divide the multiple datasets into a training group and a validation group;
[0011] Determine a basic model based on the training group and the validation group, where the basic model is used to express the relationship between the factors affecting the operating characteristics of the water turbine unit and the operating characteristics of the water turbine unit;
[0012] By repeatedly iterating between training the basic model according to the training group and verifying the accuracy of the basic model according to the validation group, construct the basic model into a water turbine unit operating characteristic model with an output deviation less than a characteristic threshold from the multiple datasets.
[0013] In a possible implementation manner, the dividing the multiple datasets into a training group and a validation group includes:
[0014] Cluster the multiple datasets to obtain multiple classes, where the operating characteristic deviation between the multiple datasets in a class is less than a class threshold;
[0015] Determine multiple operating characteristic deviation sums corresponding to the multiple classes according to the multiple classes, where the operating characteristic deviation sum is determined according to the operating characteristic deviation between the center of a class and the class centers of other classes in the multiple classes, and the class center has the smallest sum of operating characteristic deviations from other datasets in the class;
[0016] Select the two classes with the largest operating characteristic deviation sums as two target classes;
[0017] Determine the sampling quantity according to the grouping ratio and the quantity of datasets in the two target classes;
[0018] Extract multiple datasets from the multiple classes other than the two target classes as multiple target datasets according to the sampling quantity;
[0019] Use the multiple target datasets and the datasets in the two target classes as the datasets of the validation group;
[0020] Use the datasets in the multiple datasets other than the validation group as the datasets of the training group.
[0021] In a possible implementation manner, the characteristic data includes flow rate and output, and the factor data includes: guide vane opening, water head, and rotational speed. The clustering the multiple datasets to obtain multiple classes includes:
[0022] Construct a flow vector based on the flow rate, the guide vane opening, the water head, and the rotational speed, and construct an output vector based on the output power, the guide vane opening, the water head, and the rotational speed;
[0023] Class target selection step: Randomly select one dataset from multiple datasets to be classified as the class target, where the datasets to be classified are the unclustered datasets among the multiple datasets;
[0024] Class dataset search step: Search for the class dataset from the multiple datasets to be classified according to the operating characteristic deviation, the class target, and the first formula, where the first formula is:
[0025]
[0026] where, is the flow similarity coefficient, is the th element of the flow vector of the class target, is the th element of the flow vector of the dataset to be classified, is the total number of elements of the flow vector or the output vector, is the output similarity coefficient, is the th element of the output vector of the class target, is the th element of the output vector of the dataset to be classified, is the operating characteristic deviation between the class target and the dataset to be classified, is the class threshold;
[0027] If the class dataset is found and the number of datasets to be classified is not zero, then use the found class datasets as the class targets respectively, and jump to the class dataset search step;
[0028] If the class dataset is not found, jump to the class target selection step.
[0029] In a possible implementation manner, the determining the basic model according to the training group and the verification group includes:
[0030] Build the framework of the basic model, where the framework includes an input function, multiple intermediate element functions, and two output functions. The multiple intermediate element functions are divided into multiple layers according to the data transmission rules, and the two output functions correspond to the flow characteristics and the output characteristics respectively;
[0031] Use multiple activation functions as the intermediate element functions and output functions of the framework respectively to obtain multiple models;
[0032] Obtain the validation residuals of the multiple models according to the training set and the validation set;
[0033] Select the model with the smallest validation residual as the basic model;
[0034] Among them, obtaining the validation residuals of the multiple models according to the training set and the validation set includes:
[0035] For the multiple models, perform the following steps:
[0036] Data input step: Input the factor data in the multiple data sets of the training set into the model to obtain multiple outputs of the model;
[0037] Determine the output residuals of the model according to the multiple outputs of the model and the multiple characteristic data in the multiple data sets of the training set;
[0038] If the output residual is greater than the residual threshold, adjust the parameters of the model through the backpropagation algorithm according to the output residual, and jump to the data input step, where the parameters of the model are used to adjust the fitting of the model to the training set data;
[0039] Otherwise, input the factor data in the multiple data sets of the validation set into the model to obtain multiple outputs of the model, and determine the validation residuals of the model according to the multiple outputs of the model and the multiple characteristic data in the multiple data sets of the validation set.
[0040] In a possible implementation manner, building the framework of the basic model includes:
[0041] Determine the total number of layers of the intermediate element function and the total number of intermediate element functions in each layer according to the total number of data sets in the training set and the second formula, where the second formula is:
[0042]
[0043] In the formula, is the total number of intermediate element functions in each layer, is the total number of layers of the intermediate element function, is the total number of data sets in the training set;
[0044] Build the framework of the basic model according to the third formula, the total number of layers of the intermediate element function, and the total number of intermediate element functions in each layer, where the third formula is:
[0045]
[0046] In the formula, is the layer and the Output of the intermediate element function of the individual is the activation function is the element of the weight array of the intermediate element function of the th layer and the th element, is the th output of the input function, is the flow characteristic output, is the output characteristic output, is the th element of the weight array of the flow characteristic output function, is the th element of the weight array of the output characteristic output function.
[0047] In a possible implementation manner, the basic model is constructed as a water turbine unit operation characteristic model with an output deviating from the multiple data sets by less than a characteristic threshold by repeatedly iterating between training the basic model according to the training set and verifying the accuracy of the basic model according to the verification set, including:
[0048] Basic model training step: Input the factor data in the multiple data sets of the training set into the basic model to obtain multiple outputs of the model;
[0049] Determine the training residuals of the basic model according to the multiple outputs of the basic model and the multiple characteristic data in the multiple data sets of the training set;
[0050] If the number of training times has not reached the training threshold, adjust the parameters of the basic model according to the training residuals through the backpropagation algorithm, and jump to the basic model training step, where the parameters of the basic model are used to adjust the fitting of the basic model to the multiple data sets of the training set;
[0051] Otherwise, input the factor data in the multiple data sets of the verification set into the basic model to obtain multiple outputs of the basic model; Determine the fitting deviation of the current training of the basic model according to the multiple outputs of the basic model and the multiple characteristic data in the multiple data sets of the verification set; Develop an adjustment strategy for the basic model according to the change situation of the current training fitting deviation and the previous training fitting deviation.
[0052] In a possible implementation manner, the basic model includes an input function, multiple intermediate element functions, and two output functions. The multiple intermediate element functions are divided into multiple layers according to the data transmission rules. Developing the adjustment strategy for the basic model according to the change situation of the current training fitting deviation and the previous training fitting deviation includes:
[0053] If the fitting deviation of the current training is greater than the fitting deviation threshold and the fitting deviation of the current training is less than the fitting deviation of the previous training, then jump to the basic model training step;
[0054] If the fitting deviation of the current training is greater than the fitting deviation threshold and the fitting deviation of the current training is not less than the fitting deviation of the previous training, then reduce the total number of layers of the multiple intermediate element functions and increase the total number of intermediate element functions in each layer or increase the total number of layers of the multiple intermediate element functions and reduce the total number of intermediate element functions in each layer, and jump to the basic model training step;
[0055] If the fitting deviation of the current training is less than the fitting deviation threshold, then use the basic model as the operating characteristic model of the water turbine unit.
[0056] In a second aspect, an embodiment of the present invention provides a modeling device for the operating characteristics of a pumped-storage unit, which is used to implement the method for modeling the operating characteristics of a pumped-storage unit described in the first aspect above or any possible implementation manner of the first aspect. The modeling device for the operating characteristics of a pumped-storage unit includes:
[0057] A data set acquisition module, configured to acquire a plurality of data sets, where the data set includes a plurality of characteristic data representing the operating characteristics of the water turbine unit and a plurality of factor data representing the factors affecting the operating characteristics of the water turbine unit;
[0058] A data set grouping module, configured to divide the plurality of data sets into a training group and a verification group;
[0059] A basic model determination module, configured to determine a basic model according to the training group and the verification group, where the basic model is used to express the relationship between the factors affecting the operating characteristics of the water turbine unit and the operating characteristics of the water turbine unit;
[0060] And,
[0061] A water turbine unit operating characteristic model construction module, configured to iteratively construct the basic model into a water turbine unit operating characteristic model with an output deviation less than the characteristic threshold from the plurality of data sets by repeatedly iterating between training the basic model according to the training group and verifying the accuracy of the basic model according to the verification group.
[0062] In a third aspect, an embodiment of the present invention provides a terminal, including a memory and a processor. A computer program that can run on the processor is stored in the memory. When the processor executes the computer program, the steps of the method described in the first aspect above or any possible implementation manner of the first aspect are implemented.
[0063] Fourthly, an embodiment of the present invention provides a computer-readable storage medium storing a computer program, which when executed by a processor, implements the steps of the method as described in the first aspect or any possible implementation manner of the first aspect above.
[0064] The beneficial effects of the embodiment of the present invention compared with the prior art are as follows:
[0065] An embodiment of the method for modeling the operating characteristics of a pumped-storage unit disclosed in the embodiment of the present invention first obtains a plurality of data sets, where the data sets include a plurality of characteristic data representing the operating characteristics of the water turbine unit and a plurality of factor data representing the factors affecting the operating characteristics of the water turbine unit; then divides the plurality of data sets into a training group and a verification group; then determines a basic model according to the training group and the verification group, where the basic model is used to express the relationship between the factors affecting the operating characteristics of the water turbine unit and the operating characteristics of the water turbine unit; finally, by repeatedly iterating between training the basic model according to the training group and verifying the accuracy of the basic model according to the verification group, the basic model is constructed into a water turbine unit operating characteristic model with an output deviation from the plurality of data sets less than a characteristic threshold.
[0066] In the embodiment of the present invention, by repeatedly iterating the basic model between the verification group and the training group during the training and verification processes, a water turbine unit operating characteristic model is constructed. Therefore, overfitting can be effectively prevented, and accurate operating characteristics can also be output for data outside the existing data, and the fitting effect is better than that of the interpolation method.
[0067] In terms of determining the training group and the verification group of the present invention, by constructing a flow vector and a processing vector and using a clustering method, the data sets are divided into multiple classes, and the target class with operating characteristics far from other classes is used as the verification group. The verification group obtained through the above process can better verify the ability of the basic model to express the characteristics of the water turbine unit, and a better basic model can be selected. Description of the Drawings
[0068] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.
[0069] Figure 1 is a flowchart of the method for modeling the operating characteristics of a pumped-storage unit provided by an embodiment of the present invention;
[0070] Figure 2It is the characteristic diagram of the rotational speed - flow rate - guide vane opening of the water turbine unit provided by the embodiment of the present invention;
[0071] Figure 3 It is the characteristic diagram of the flow rate - guide vane opening of the water turbine unit at a specific rotational speed provided by the embodiment of the present invention;
[0072] Figure 4 It is the schematic diagram of the basic model principle provided by the embodiment of the present invention;
[0073] Figure 5 It is the functional block diagram of the modeling device for the operating characteristics of the pumped - storage unit provided by the embodiment of the present invention;
[0074] Figure 6 It is the functional block diagram of the terminal provided by the embodiment of the present invention. Detailed implementation manners
[0075] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are presented to thoroughly understand the embodiments of the present invention. However, those skilled in the art should clearly understand that the present invention can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well - known systems, devices, and methods are omitted to avoid unnecessary details from interfering with the description of the present invention.
[0076] To make the purpose, technical solutions, and advantages of the present invention clearer, the following will be described through specific implementation manners in conjunction with the accompanying drawings.
[0077] The following details the embodiments of the present invention. This example is implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given. However, the protection scope of the present invention is not limited to the following embodiments.
[0078] Figure 1 It is the flowchart of the method for modeling the operating characteristics of the pumped - storage unit provided by the embodiment of the present invention.
[0079] As Figure 1 shown, it shows the implementation flowchart of the method for modeling the operating characteristics of the pumped - storage unit provided by the embodiment of the present invention, which is described in detail as follows:
[0080] In step 101, a plurality of data sets are obtained, where the data set includes a plurality of characteristic data representing the operating characteristics of the water turbine unit and a plurality of factor data representing the factors affecting the operating characteristics of the water turbine unit.
[0081] In step 102, the plurality of data sets are divided into a training group and a verification group.
[0082] In some embodiments, step 102 includes:
[0083] Cluster the multiple data sets to obtain multiple classes, where the deviation of the operating characteristics between multiple data sets in a class is less than the class threshold;
[0084] According to the multiple classes, determine the sum of multiple operating characteristic deviations corresponding to the multiple classes, where the sum of operating characteristic deviations is determined according to the operating characteristic deviation between the center of the class and the class centers of other classes in the multiple classes, and the sum of the operating characteristic deviations between the class center and other data sets in the class is the smallest;
[0085] Select two classes with the largest sum of operating characteristic deviations as two target classes;
[0086] Determine the sampling quantity according to the grouping ratio and the quantity of data sets in the two target classes;
[0087] According to the sampling quantity, extract multiple data sets from multiple classes other than the two target classes as multiple target data sets;
[0088] Use the multiple target data sets and the data sets in the two target classes as the data sets of the verification group;
[0089] Use the data sets other than the verification group in the multiple data sets as the data sets of the training group.
[0090] In some embodiments, the characteristic data includes flow rate and output, and the factor data includes: guide vane opening, water head, and rotational speed. The clustering of the multiple data sets to obtain multiple classes includes:
[0091] Construct a flow vector according to the flow rate, the guide vane opening, the water head, and the rotational speed, and construct an output vector according to the output, the guide vane opening, the water head, and the rotational speed;
[0092] Class target selection step: Randomly select a data set from multiple data sets to be classified as the class target, where the data sets to be classified are the data sets in the multiple data sets that have not been clustered;
[0093] Class data set search step: Search for class data sets from the multiple data sets to be classified according to the operating characteristic deviation, the class target, and the first formula, where the first formula is:
[0094]
[0095] where, is the flow similarity coefficient, is the th element of the flow vector of the class target, is the th element of the flow vector of the data set to be classified, is the total number of elements of the flow vector or output vector, is the output similarity coefficient, is the th element of the output vector of the class target, is the th element of the output vector of the dataset to be classified, is the deviation of the operating characteristics between the class target and the dataset to be classified, is the class threshold;
[0096] If a class dataset is found and the number of datasets to be classified is not zero, the found class datasets are respectively used as class targets, and jump to the class dataset search step;
[0097] If no class dataset is found, jump to the class target selection step.
[0098] Exemplarily, Figure 2 shows the characteristic diagram of the water turbine unit speed-flow-guide vane opening in the power generation mode. The present invention aims to establish a model of the operating characteristics of the water turbine unit through existing data, that is, to be able to express the relationship between these multi-dimensional data such as flow rate, guide vane opening, water head, speed, output, and flow rate.
[0099] Actually, the accuracy of the model is used to express the accuracy of fitting the existing data on the one hand, and on the other hand, it is more expected to be able to make accurate outputs for unknown data.
[0100] Figure 2 The edge part 201 of the surface in is far from the surface body, and in the fitting of the model, it is a part that is not easily fitted accurately, Figure 3 is based on Figure 2 The curve diagram of the guide vane opening and flow rate drawn at a certain specific speed in can be seen that the curve is constructed by multiple existing data 301. Due to the lack of data at the endpoints 302 of the curve, the fitting effect is often not as accurate as the interval data 303 between the existing data 301.
[0101] Therefore, in testing the expression of the operating characteristics of the model, using edge data combined with interval data for testing has a more prominent effect.
[0102] The present invention provides a method for grouping multiple datasets through a clustering method. Among them, the datasets far from the points used for fitting will be used as the verification group, and the remaining datasets will be used as the training group. In other words, the distance between the data of the verification group and the training group is widened as much as possible to achieve the effect of improving the resolution and fitting ability.
[0103] Therefore, after clustering, according to the centers of the classes, the two classes with the farthest sum of distances to the centers of other classes are used as the verification group. For the part that fails to reach the number of the verification group dataset (generally, the ratio of the number of verification group data to the number of training group data is configured as 1:4, and the number of datasets found through clustering usually cannot reach this number, and the accuracy of the fitted part also needs to be verified), random sampling is performed from the remaining classes. After reaching the preset number of the verification group dataset, the remaining datasets are used as the training group data for model training.
[0104] Actually, some operating characteristic models can express the relationships among flow rate, guide vane opening, water head, rotational speed, and output. By constructing a flow rate and output model, the general requirements of the water turbine unit can be met. In the present invention, through the two outputs of the model, the relationships among guide vane opening, water head, rotational speed, output, and flow rate are expressed.
[0105] Before clustering, first construct two vectors for each dataset, namely the flow rate vector and the output vector. The vectors can be obtained by arranging the data in a preset order.
[0106] Then, randomly find a clustering center from the unclustered datasets, and through the first formula:
[0107]
[0108] where, is the flow rate similarity coefficient, is the -th element of the flow rate vector of the class target, is the -th element of the flow rate vector of the dataset to be classified, is the total number of elements of the flow rate vector or the output vector, is the output similarity coefficient, is the -th element of the output vector of the class target, is the -th element of the output vector of the dataset to be classified, is the operating characteristic deviation between the class target and the dataset to be classified, is the class threshold, and perform in-class dataset search. When a dataset is found, the found dataset is used as the clustering center and the search operation is performed again. Otherwise, randomly specify a clustering center from the unclustered datasets again.
[0109] In step 103, according to the training group and the verification group, determine a basic model, where the basic model is used to express the relationship between the factors affecting the operating characteristics of the water turbine unit and the operating characteristics of the water turbine unit.
[0110] In some embodiments, step 103 includes:
[0111] Build the framework of the basic model, where the framework includes an input function, a plurality of intermediate element functions, and two output functions. The plurality of intermediate element functions are divided into multiple layers according to the data transfer rules, and the two output functions correspond to the flow characteristics and the output characteristics respectively;
[0112] Use a plurality of activation functions as the intermediate element functions and output functions of the framework respectively to obtain a plurality of models;
[0113] Obtain the verification residuals of the plurality of models according to the training set and the verification set;
[0114] Select the model with the smallest verification residual as the basic model;
[0115] Among them, obtaining the verification residuals of the plurality of models according to the training set and the verification set includes:
[0116] For the plurality of models, perform the following steps:
[0117] Data input step: Input the factor data in the multiple data sets of the training set into the model to obtain multiple outputs of the model;
[0118] Determine the output residuals of the model according to the multiple outputs of the model and the multiple characteristic data in the multiple data sets of the training set;
[0119] If the output residual is greater than the residual threshold, adjust the parameters of the model according to the output residual through the backpropagation algorithm, and jump to the data input step, where the parameters of the model are used to adjust the fitting of the model to the training set data;
[0120] Otherwise, input the factor data in the multiple data sets of the verification set into the model to obtain multiple outputs of the model, and determine the verification residuals of the model according to the multiple outputs of the model and the multiple characteristic data in the multiple data sets of the verification set.
[0121] In some embodiments, building the framework of the basic model includes:
[0122] Determine the total number of layers of the intermediate element functions and the total number of intermediate element functions in each layer according to the total number of data sets in the training set and the second formula, where the second formula is:
[0123]
[0124] In the formula, is the total number of intermediate element functions in each layer, is the total number of layers of the intermediate element functions, is the total number of data sets in the training group;
[0125] Build the framework of the basic model according to the third formula, the total number of layers of the intermediate element function, and the total number of intermediate element functions in each layer, where the third formula is:
[0126]
[0127] In the formula, is the output of the th intermediate element function of the is the activation function, is the th element of the weight array of the intermediate element function of the is the th output of the input function, is the flow characteristic output, is the output characteristic output, is the th element of the weight array of the flow characteristic output function, is the th element of the weight array of the output characteristic output function.
[0128] Exemplarily, in terms of determining the basic model, the embodiment of the present invention trains the basic model according to the training group, verifies the training result according to the verification group, and determines the basic model.
[0129] The basic model has a basic framework structure, as Figure 4 shown. In the figure, the input function 401 receives the input of the data set factors, the intermediate element function 402 is divided into multiple layers, and each layer has multiple intermediate element functions 402. All the intermediate element functions 402 sequentially transmit the output results, and finally two output functions 403 output the results indicating the flow characteristics and the output characteristics.
[0130] In terms of determining the framework structure, the first step is to determine the number of layers and the total number of element functions in each layer. Specifically, it should meet the constraints of the second formula:
[0131]
[0132] In the formula, J is the total number of intermediate element functions in each layer, K is the total number of layers of the intermediate element function, and S is the total number of data sets in the training group.
[0133] The entire basic model framework conforms to the expression of the third formula:
[0134]
[0135] In the formula, is the output of the intermediate element function of the th layer and the th one, is the activation function, is the th element of the weight array of the intermediate element function of the th layer and the th one, is the th output of the input function, is the flow characteristic output, is the output characteristic output, is the th element of the weight array of the flow characteristic output function, is the th element of the weight array of the output characteristic output function.
[0136] It can be seen that the final effect of this framework is related to the selection of the activation function. Therefore, in the embodiments of the present invention, on the basis of obtaining multiple activation functions, training is sequentially performed through the training set, and when the training reaches a predetermined target, the accuracy of the model is verified through the verification set, so as to achieve the effect of screening the activation function. For example, in an application scenario, using the tanh activation function can achieve good training results and can also obtain good verification results during verification.
[0137] In step 104, by repeatedly iterating between training the basic model according to the training set and verifying the accuracy of the basic model according to the verification set, the basic model is constructed into a water turbine unit operation characteristic model with an output having a deviation less than the characteristic threshold from the multiple data sets.
[0138] In some embodiments, step 104 includes:
[0139] Basic model training step: Input the factor data in the multiple data sets of the training set into the basic model to obtain multiple outputs of the model;
[0140] Determine the training residual of the basic model according to the multiple outputs of the basic model and the multiple characteristic data in the multiple data sets of the training set;
[0141] If the number of training times does not reach the training threshold, then adjust the parameters of the basic model according to the training residual through the backpropagation algorithm, and jump to the basic model training step, where the parameters of the basic model are used to adjust the fitting of the basic model to the multiple data sets of the training set;
[0142] Otherwise, input the factor data in multiple datasets of the verification group into the basic model to obtain multiple outputs of the basic model; determine the fitting deviation of the current training of the basic model according to the multiple outputs of the basic model and the multiple characteristic data in the multiple datasets of the verification group; formulate an adjustment strategy for the basic model according to the change of the current training fitting deviation and the previous training fitting deviation.
[0143] In some embodiments, the basic model includes an input function, multiple intermediate element functions, and two output functions. The multiple intermediate element functions are divided into multiple layers according to the data transmission rules. Formulating the adjustment strategy for the basic model according to the change of the current training fitting deviation and the previous training fitting deviation includes:
[0144] If the current training fitting deviation is greater than the fitting deviation threshold and the current training fitting deviation is less than the previous training fitting deviation, jump to the basic model training step;
[0145] If the current training fitting deviation is greater than the fitting deviation threshold and the current training fitting deviation is not less than the previous training fitting deviation, reduce the total number of layers of the multiple intermediate element functions and increase the total number of intermediate element functions in each layer or increase the total number of layers of the multiple intermediate element functions and reduce the total number of intermediate element functions in each layer, and then jump to the basic model training step;
[0146] If the current training fitting deviation is less than the fitting deviation threshold, use the basic model as the operating characteristic model of the water turbine unit.
[0147] Exemplarily, for the final determination of the model, it is mainly to determine the number of layers of the model and the number of intermediate element functions in each layer. When the numbers of both are too large, due to the overfitting characteristic, only a small fitting deviation is obtained at the fitting points of the training dataset, and for the part other than the fitting points, the output result deviation of the model is large, resulting in poor application effect of the model.
[0148] Therefore, in the embodiments of the present invention, first, after training several times and verifying through the verification group dataset, if the verification deviation is small, it indicates that the model construction has achieved the expected purpose, and the construction of the operating characteristic model of the unit is completed.
[0149] If the verification deviation is large and there is a downward trend compared with the previous verification, it indicates that training is still needed to expect a better fitting effect. At this time, it should return to the training step, train the basic model with the data of the training group, and then verify after training a predetermined number of times, and so on.
[0150] If the verification deviation is relatively large and there is no obvious change compared with the previous verification, or the deviation shows an increasing trend, it indicates that the training has reached or exceeded the optimal period. However, for the verification group data, a good fitting effect still cannot be obtained. At this time, the problem lies in the fact that the model structure is too complex, and the complexity of the model structure is too high relative to the data set. At this time, the number of intermediate element function layers and the number of intermediate element functions in each layer should be adjusted. Generally, to make the most correct parameters of the model unique, when reducing the number of intermediate element function layers to be adjusted, the number of intermediate element functions in each layer should be increased, or when increasing the number of intermediate element function layers to be adjusted, the number of intermediate element functions in each layer should be reduced to satisfy the following inequality:
[0151]
[0152] In the formula, is the total number of intermediate element functions in each layer, is the total number of intermediate element function layers, is the total number of data sets in the training group.
[0153] After adjusting the model structure, it is possible to return to the step of training through the training group and perform an iterative process of training and verifying the basic model.
[0154] In the implementation manner of the modeling method for the operating characteristics of the pumped-storage unit of the present invention, first, a plurality of data sets are obtained, where the data sets include a plurality of characteristic data representing the operating characteristics of the water turbine unit and a plurality of factor data representing the factors affecting the operating characteristics of the water turbine unit; then, the plurality of data sets are divided into a training group and a verification group; then, according to the training group and the verification group, a basic model is determined, where the basic model is used to express the relationship between the factors affecting the operating characteristics of the water turbine unit and the operating characteristics of the water turbine unit; finally, by repeatedly iterating between training the basic model according to the training group and verifying the accuracy of the basic model according to the verification group, the basic model is constructed into a water turbine unit operating characteristics model with an output deviation less than the characteristic threshold from the plurality of data sets.
[0155] In the implementation manner of the present invention, by repeatedly iterating the basic model between the verification group and the training group during the training and verification process, a water turbine unit operating characteristics model is constructed. Therefore, overfitting can be effectively prevented, and accurate operating characteristics can also be output for data outside the existing data, and the fitting effect is better than that of the interpolation method.
[0156] In the determination of the training group and the validation group, the present invention constructs a traffic vector and a processing vector, adopts a clustering method to divide the data set into multiple classes, and uses the target class with operating characteristics far from other classes as the validation group. The validation group obtained through the above process can better verify the ability of the basic model to express the characteristics of the water turbine unit, and a better basic model can be selected.
[0157] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0158] The following is the device embodiment of the present invention. For the details not described in detail, reference can be made to the corresponding method embodiments above.
[0159] Figure 5 is a functional block diagram of a device for modeling the operating characteristics of a pumped-storage unit provided by an embodiment of the present invention. Referring to Figure 5 , the device 5 for modeling the operating characteristics of a pumped-storage unit includes: a data set acquisition module 501, a data set grouping module 502, a basic model determination module 503, and a water turbine unit operating characteristic model construction module 504, where:
[0160] The data set acquisition module 501 is configured to acquire a plurality of data sets, where the data set includes a plurality of characteristic data representing the operating characteristics of the water turbine unit and a plurality of factor data representing the factors affecting the operating characteristics of the water turbine unit;
[0161] The data set grouping module 502 is configured to divide the plurality of data sets into a training group and a validation group;
[0162] The basic model determination module 503 is configured to determine a basic model according to the training group and the validation group, where the basic model is used to express the relationship between the factors affecting the operating characteristics of the water turbine unit and the operating characteristics of the water turbine unit;
[0163] And,
[0164] The water turbine unit operating characteristic model construction module 504 is configured to repeatedly iterate between training the basic model according to the training group and verifying the accuracy of the basic model according to the validation group, and construct the basic model into a water turbine unit operating characteristic model with an output deviation from the plurality of data sets less than a characteristic threshold.
[0165] Figure 6 is a functional block diagram of a terminal provided by an embodiment of the present invention. As Figure 6As shown, the terminal 6 of this embodiment includes: a processor 600 and a memory 601, and a computer program 602 that can run on the processor 600 is stored in the memory 601. When the processor 600 executes the computer program 602, it implements the steps in the above-mentioned various modeling methods and embodiments of the operating characteristics of pumped-storage units, such as Figure 1 the steps 101 to 104 shown
[0166] Exemplarily, the computer program 602 can be divided into one or more modules / units, and the one or more modules / units are stored in the memory 601 and executed by the processor 600 to complete the present invention.
[0167] The terminal 6 can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal 6 may include, but is not limited to, a processor 600 and a memory 601. Those skilled in the art can understand that Figure 6 merely examples of the terminal 6, which do not constitute a limitation on the terminal 6, may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the terminal 6 may further include input / output devices, network access devices, buses, etc.
[0168] The so-called processor 600 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0169] The memory 601 may be an internal storage unit of the terminal 6, such as the hard disk or memory of the terminal 6. The memory 601 may also be an external storage device of the terminal 6, such as a plug-in hard disk equipped on the terminal 6, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 601 may also include both the internal storage unit of the terminal 6 and external storage devices. The memory 601 is used to store the computer program 602 and other programs and data required by the terminal 6. The memory 601 may also be used to temporarily store data that has been output or is to be output.
[0170] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.
[0171] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0172] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0173] In the embodiments provided by the present invention, it should be understood that the disclosed device / terminal and method can be implemented in other ways. For example, the device / terminal embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.
[0174] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0175] In addition, each functional unit in various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0176] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above various method and device embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0177] The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for modeling the operating characteristics of a pumped storage unit, characterized in that: include: Acquire multiple data sets, wherein the data sets include multiple characteristic data representing the operating characteristics of the water turbine unit and multiple factor data representing the operating characteristics affecting the water turbine unit; Dividing the multiple data sets into a training group and a validation group, wherein the validation group is obtained by sampling based on clustering results of the multiple data sets; Determining a basic model according to the training group and the validation group, wherein the basic model is used to express the relationship between the factors affecting the operating characteristics of the hydraulic turbine unit and the operating characteristics of the hydraulic turbine unit; The basic model is constructed to output a hydro-turbine unit operation characteristic model whose deviation from the plurality of data sets is less than a characteristic threshold by repeatedly iterating between training the basic model according to the training set and verifying the accuracy of the basic model according to the verification set; The characteristic data includes flow rate and output, the factor data includes guide vane opening, water head and rotation speed, and the clustering process of the multiple data sets includes: constructing a flow vector according to the flow, the guide vane opening, the water head and the rotational speed, and constructing an output vector according to the output, the guide vane opening, the water head and the rotational speed; Class target selection step: randomly selecting a data set from multiple data sets to be classified as a class target, wherein the data set to be classified is a non-clustered data set among the multiple data sets; Class data set search step: search for a class data set from the multiple data sets to be classified according to the operating characteristic deviation, the class target and the first formula, wherein the first formula is: in, is the flow similarity coefficient, is the flow vector of the class target elements, is the flow vector of the data set to be classified elements, is the total number of elements of the flow vector or output vector, is the output similarity coefficient, is the output vector of the class target elements, is the output vector of the data set to be classified elements, is the running characteristic deviation between the class target and the data set to be classified, is the class threshold; If the class data set is found and the number of the data sets to be classified is not zero, the found class data sets are respectively used as class targets, and the process jumps to the class data set search step; If the class data set is not found, jump to the class target selection step.
2. The method for modeling the operating characteristics of a pumped storage unit according to claim 1, characterized in that: The step of dividing the plurality of data sets into a training group and a validation group comprises: Clustering the multiple data sets to obtain multiple classes, wherein the deviation of the operating characteristics between the multiple data sets in a class is less than a class threshold; Determine, according to the plurality of classes, a plurality of running characteristic deviation sums corresponding to the plurality of classes, wherein the running characteristic deviation sums are determined according to running characteristic deviations between a class center and class centers of other classes in the plurality of classes, and the running characteristic deviation sum between the class center and other data sets in the class is the smallest; Select the two classes with the largest running characteristic deviation as the two target classes; Determine the number of samples according to the grouping ratio and the number of data sets in the two target classes; Extracting multiple data sets as multiple target data sets from multiple classes outside the two target classes according to the sampling quantity; Using the multiple target data sets and the data sets in the two target classes as data sets of a validation group; The data sets other than the validation set in the multiple data sets are used as data sets of the training set.
3. The method for modeling the operating characteristics of a pumped storage unit according to claim 1, characterized in that: The step of determining a basic model according to the training group and the validation group includes: Building a framework of the basic model, wherein the framework includes an input function, multiple intermediate element functions and two output functions, wherein the multiple intermediate element functions are divided into multiple layers according to data transmission rules, and the two output functions correspond to flow characteristics and output characteristics respectively; Using multiple activation functions as intermediate element functions and output functions of the framework, respectively, to obtain multiple models; Obtaining validation residuals of the multiple models according to the training group and the validation group; Select the model with the smallest validation residual as the base model; Wherein, obtaining the validation residuals of the multiple models according to the training group and the validation group includes: For the multiple models, perform the following steps: Data input step: inputting factor data in the multiple data sets of the training group into the model to obtain multiple outputs of the model; Determining the output residual of the model according to the multiple outputs of the model and the multiple characteristic data in the multiple data sets of the training group; If the output residual is greater than the residual threshold, the parameters of the model are adjusted according to the output residual through the back propagation algorithm, and the process is skipped to the data input step, wherein the parameters of the model are used to adjust the fit of the model to the training set data; Otherwise, the factor data in the multiple data sets of the validation group are input into the model, multiple outputs of the model are obtained, and the validation residuals of the model are determined based on the multiple outputs of the model and the multiple characteristic data in the multiple data sets of the validation group.
4. The method for modeling the operating characteristics of a pumped storage unit according to claim 3, characterized in that: The framework for building the basic model includes: According to the total number of data sets in the training group and the second formula, the total number of layers of intermediate element functions and the total number of intermediate element functions in each layer are determined, wherein the second formula is: In the formula, is the total number of intermediate element functions in each layer, is the total number of layers of intermediate element functions, is the total number of data sets in the training set; The framework of the basic model is built according to the third formula, the total number of layers of intermediate element functions and the total number of intermediate element functions in each layer, wherein the third formula is: In the formula, For the Tier The output of the intermediate element function, is the activation function, For the Tier The weight array of the middle element function elements, The first outputs, is the flow characteristic output, Output characteristic output, The weight array of the flow characteristic output function elements, The weight array of the output characteristic output function elements.
5. The method for modeling the operating characteristics of a pumped storage unit according to any one of claims 1 to 4, characterized in that: The method of repeatedly iterating between training the basic model according to the training group and verifying the accuracy of the basic model according to the verification group, constructing the basic model to output a hydraulic turbine group operation characteristic model whose deviation from the multiple data sets is less than a characteristic threshold, comprises: Basic model training step: inputting factor data in the multiple data sets of the training group into the basic model to obtain multiple outputs of the model; Determining a training residual of the basic model according to a plurality of outputs of the basic model and a plurality of characteristic data in a plurality of data sets of the training group; If the number of training times does not reach the training threshold, adjusting the parameters of the basic model through the back propagation algorithm according to the training residual, and jumping to the basic model training step, wherein the parameters of the basic model are used to adjust the fit of the basic model to the multiple data sets of the training group; Otherwise, the factor data in the multiple data sets of the verification group are input into the basic model to obtain multiple outputs of the basic model; the fitting deviation of the basic model for this training is determined according to the multiple outputs of the basic model and the multiple characteristic data in the multiple data sets of the verification group; and an adjustment strategy for the basic model is formulated according to the changes in the fitting deviation of this training and the fitting deviation of the previous training.
6. The method for modeling the operating characteristics of a pumped storage unit according to claim 5, characterized in that: The basic model includes an input function, multiple intermediate element functions and two output functions. The multiple intermediate element functions are divided into multiple layers according to the rules of data transmission. The adjustment strategy of the basic model is formulated according to the changes of the current training fitting deviation and the previous training fitting deviation, including: If the fitting deviation of the current training is greater than the fitting deviation threshold and the fitting deviation of the current training is less than the fitting deviation of the previous training, jump to the basic model training step; If the fitting deviation of this training is greater than the fitting deviation threshold and the fitting deviation of this training is not less than the fitting deviation of the previous training, then the total number of layers of multiple intermediate element functions is reduced and the total number of intermediate element functions in each layer is increased or the total number of layers of multiple intermediate element functions is increased and the total number of intermediate element functions in each layer is reduced, and jump to the basic model training step; If the fitting deviation of this training is less than the fitting deviation threshold, the basic model is used as the operating characteristic model of the hydro-turbine unit.
7. A pumped storage unit operation characteristic modeling device, characterized in that: Used to implement the method for modeling the operating characteristics of a pumped-storage unit according to any one of claims 1 to 6, the pumped-storage unit operating characteristics modeling device comprises: A data set acquisition module, used to acquire multiple data sets, wherein the data sets include multiple characteristic data representing the operating characteristics of the hydraulic turbine unit and multiple factor data representing the factor data affecting the operating characteristics of the hydraulic turbine unit; The dataset grouping module is used to divide multiple datasets into training groups and validation groups; A basic model determination module, used to determine a basic model according to the training group and the verification group, wherein the basic model is used to express the relationship between the factors affecting the operating characteristics of the hydraulic turbine unit and the operating characteristics of the hydraulic turbine unit; as well as, A water turbine group operation characteristic model construction module is used to construct the basic model into a water turbine group operation characteristic model whose output deviation from the multiple data sets is less than a characteristic threshold by repeatedly iterating between training the basic model according to the training group and verifying the accuracy of the basic model according to the verification group.
8. A terminal comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Control method, device and equipment of variable-speed pumped storage unit and medium
CN116292057A