Training methods and related equipment for strategy prediction models of power generation processes

CN115933380BActive Publication Date: 2026-09-01INST OF AUTOMATION CHINESE ACAD OF SCI +2
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211436526.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-16
Publication Date
2026-09-01
Estimated Expiration
2042-11-16

AI Technical Summary

Technical Problem

[0004]本发明提供一种发电过程的策略预测模型的训练方法、装置、电子设备和存储介质,用以解决现有技术中发电过程中参数控制较为粗糙的缺陷,实现发电策略的动态调节,调高了发电设备运行的安全性和发电效率,具有更好的发电控制效果

Benefits of technology

[0013]本发明还提供一种非暂态计算机可读存储介质,其上存储有计算机程序,该计算机程序被处理器执行时实现如上述任一种所述发电过程的策略预测模型的训练方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115933380B_ABST
    Figure CN115933380B_ABST
Patent Text Reader

Abstract

This invention provides a training method and related equipment for a strategy prediction model of a power generation process. The method includes: initializing an initial strategy prediction model, which includes a control network, a disturbance network, a model network, and an evaluation network; acquiring historical power generation data and optimizing the weights of the model network based on the historical power generation data to obtain an optimized intermediate strategy prediction model; training the intermediate strategy prediction model based on the historical power generation data to obtain optimized evaluation, control, and disturbance networks; determining whether the trained strategy prediction model has converged, and obtaining the strategy prediction model upon confirmation of convergence. By mutually restricting the various networks and simultaneously introducing the influence of control and disturbance factors on power generation, the accuracy of power generation strategy prediction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power generation process optimization control technology, and in particular to a training method, apparatus, electronic device and storage medium for a strategy prediction model of a power generation process. Background Technology

[0002] With the continuous development and updating of technology, power generation technology is also constantly being improved and developed. However, in industry, thermal power generation is still the most commonly used power generation method. In order to generate electricity more safely and rationally, automation control technology plays a significant role in the thermal power generation process, which to a certain extent improves the safety of power generation and the rational allocation of power resources.

[0003] However, traditional automated control technologies are relatively crude in parameter control. For example, they rely on pre-set parameter adjustments and settings, ignoring the impact of changes in the environment and power generation scenarios on power generation. These settings, besides limiting power generation efficiency, can also introduce safety hazards due to environmental variability. Summary of the Invention

[0004] This invention provides a training method, apparatus, electronic device, and storage medium for a strategy prediction model of a power generation process, which addresses the shortcomings of coarse parameter control in the existing technology during power generation, enables dynamic adjustment of the power generation strategy, improves the safety and efficiency of power generation equipment operation, and has a better power generation control effect.

[0005] This invention provides a training method for a strategy prediction model of a power generation process, comprising: The initial policy prediction model is initialized, wherein the policy prediction model includes a control network, a disturbance network, a model network, and an evaluation network. Historical power generation data is acquired, and the model network is weighted and optimized based on the historical power generation data to obtain an optimized intermediate strategy prediction model. Based on historical power generation data, the intermediate strategy prediction model is trained to obtain the optimized evaluation network, the control network, and the interference network. Determine whether the trained policy prediction model has converged, and obtain the policy prediction model when convergence is confirmed.

[0006] According to the training method of a strategy prediction model for a power generation process provided by the present invention, the step of optimizing the weights of the model network based on the historical power generation data to obtain an optimized intermediate strategy prediction model includes: Based on time series, the state information corresponding to each moment is obtained from the historical power generation data, and the state information is grouped and associated to obtain several state groups, wherein the state group contains two state information and the time is adjacent. The aforementioned sets of states are used as the training set for the model network, and the model network is trained accordingly. When training is complete, an intermediate policy prediction model is obtained based on the model network with optimized weights.

[0007] According to the training method of a strategy prediction model for a power generation process provided by the present invention, the step of using the plurality of state groups as the training set of the model network and training the model network includes: Based on the timing sequence, the state information in the several groups of state groups is divided into input state and output state; The input state is input into the model network to obtain the predicted state corresponding to the input state; The state error is calculated based on the predicted state and the output state, and training is considered complete when the state error is less than a preset error value.

[0008] According to the training method of a strategy prediction model for a power generation process provided by the present invention, the step of training the intermediate strategy prediction model based on historical power generation data to obtain the optimized evaluation network, the control network, and the interference network includes: Based on time tags, the historical power generation data is used as input and input to the first evaluation network of the control network, the interference network, and the evaluation network, respectively, to obtain control output, interference output, and first evaluation output. The control output, the disturbance output, and the historical power generation data are input into the optimized model network, and the model output is obtained. The model output is input into the second evaluation network of the evaluation network, and the second evaluation output is obtained. Based on the first evaluation output and the second evaluation output, the evaluation loss value of the evaluation network is calculated, and the control loss value of the control network and the interference loss value of the interference network are obtained according to the optimal value function. Based on the evaluation loss value, the control loss value, and the interference loss value, the optimized evaluation network, the control network, and the interference network are obtained.

[0009] According to the training method of the strategy prediction model for a power generation process provided by the present invention, the step of obtaining the control loss value corresponding to the control network and the disturbance loss value corresponding to the disturbance network based on the optimal value function includes: The control network is differentiated according to the optimal value function to obtain a first derivative result, and the control loss value corresponding to the control network is calculated based on the first derivative result and the control output. The interference network is differentiated according to the optimal value function to obtain a second derivative result, and the interference loss value corresponding to the control network is calculated based on the second derivative result and the interference output.

[0010] According to the training method of a strategy prediction model for a power generation process provided by the present invention, the step of obtaining optimized evaluation networks, control networks, and disturbance networks based on the evaluation loss value, the control loss value, and the disturbance loss value includes: The evaluation loss value is compared with the loss threshold to determine whether the evaluation network, the control network and the interference network have been optimized. When it is determined that the optimization has been completed, the optimized evaluation network, the control network and the interference network are obtained. The determination of whether optimization is complete includes: If the evaluated loss value is less than or equal to the loss threshold, then the optimization is determined to be complete; If the evaluation loss value is greater than the loss threshold, then the optimization is determined to be incomplete.

[0011] According to the training method of a strategy prediction model for a power generation process provided by the present invention, the method further includes: Responding to power generation prediction commands, it receives the initial input state; Based on the trained policy prediction model, the control policy and the disturbance policy in the initial state are obtained. Based on the control strategy and the interference strategy, power generation control and regulation are performed. The present invention also provides a training apparatus for a strategy prediction model of a power generation process, comprising: An initial adjustment module is used to initialize the initial policy prediction model, wherein the policy prediction model includes a control network, a disturbance network, a model network, and an evaluation network. The first optimization module is used to acquire historical power generation data and optimize the model network based on the historical power generation data to obtain an optimized intermediate strategy prediction model. The second optimization module is used to train the intermediate strategy prediction model based on historical power generation data to obtain the optimized evaluation network, the control network, and the interference network. The convergence judgment module is used to determine whether the trained policy prediction model has converged, and to obtain the policy prediction model when convergence is determined.

[0012] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement a training method for a strategy prediction model of the power generation process as described above.

[0013] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a training method for a strategy prediction model of the power generation process as described above.

[0014] This invention provides a training method, apparatus, electronic device, and storage medium for a power generation strategy prediction model. Based on a power generation system, a corresponding power generation strategy prediction model is constructed. Several different networks are set up, and after initializing the weights of each network, the model network that will output the next state is first trained and its weights optimized. By training separately, the influence of other networks on the model network is reduced, improving the robustness and prediction accuracy of the model network. Upon completion of optimized training, the strategy prediction model is treated as a whole, and the weights of each network are optimized. Through mutual constraints between the control network, disturbance network, and evaluation network, the model's robustness and training efficiency are improved. Finally, convergence is used to ensure the model meets usage requirements. By introducing control and disturbance factors, a better overall effect of power generation control is achieved, ensuring stable equipment operation and improving the accuracy of power generation strategy prediction. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0016] Figure 1 This is a flowchart illustrating the training method for the strategy prediction model of the power generation process provided by the present invention.

[0017] Figure 2 This is a schematic diagram of the strategy prediction model provided by the present invention.

[0018] Figure 3 This is a flowchart illustrating the process of obtaining the intermediate strategy prediction model provided by the present invention.

[0019] Figure 4 This is a flowchart illustrating the process of training a model network provided by the present invention.

[0020] Figure 5 This is a flowchart illustrating the process of training an intermediate strategy prediction model provided by the present invention.

[0021] Figure 6 This is a schematic diagram of the structure of the training device for the strategy prediction model of the power generation process provided by the present invention.

[0022] Figure 7 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0024] The following is combined with Figures 1-5 The training method for the strategy prediction model of the power generation process of the present invention is described.

[0025] Figure 1 This is a flowchart illustrating the training method for the strategy prediction model of the power generation process provided by this invention. For example... Figure 1 As shown, the method includes: Step 101: Initialize the initial policy prediction model, which includes a control network, a disturbance network, a model network, and an evaluation network.

[0026] To ensure timely and accurate power generation regulation during operation, a regulation strategy needs to be pre-determined. Based on this pre-defined strategy and the actual situation and power generation status, the strategy can be adjusted promptly and accurately. Specifically, when determining the power generation regulation strategy, adaptive dynamic programming of the power generation system is used to accurately optimize the regulation strategy. For the power generation system, modeling can be performed. To predict the power generation strategy, a corresponding strategy prediction model can be constructed and initialized. Then, the model can be optimized and trained for subsequent use.

[0027] The constructed policy prediction model can be as follows: Figure 2 As shown, Figure 2 This is a schematic diagram of the strategy prediction model provided by the present invention. The strategy prediction model comprises several parts, specifically: a control network, an interference network, a model network, and several evaluation networks. Here, the number of evaluation networks is set to two. In the strategy prediction model, the control network and the interference network respectively provide reasonable control and interference quantities to the power generation strategy, making the obtained power generation strategy more consistent with actual application scenarios and usage conditions.

[0028] For each network structure in the constructed policy prediction model, the network structure can be set and adjusted according to actual needs. Specific settings can be made according to requirements, such as setting different network structures for different networks. In various embodiments, for example, the network structure of the model network can be set to 8-17-4, where 8 is the number of input layer nodes, 17 is the number of hidden layer nodes, and 4 is the number of output layer nodes; the network structure of the control network and interference network can be set to 4-8-2, where 4 is the number of input layer nodes, 8 is the number of hidden layer nodes, and 2 is the number of output layer nodes; and the network structure of the evaluation network can be set to 4-8-1, where 4 is the number of input layer nodes, 8 is the number of hidden layer nodes, and 1 is the number of output layer nodes.

[0029] Furthermore, depending on the specific needs of different networks, appropriate neural networks can be used when constructing each network, such as establishing model networks, control networks, interference networks, and / or evaluation networks based on BP neural networks.

[0030] It should be noted that for the policy prediction model, all networks need to be optimized during training. However, after the optimization training is completed, when making policy predictions, it is only necessary to obtain the control and disturbance quantities in the actual input state, and then adjust the policy based on the obtained control and disturbance quantities.

[0031] When constructing a strategy prediction model for a power generation system, system characteristics can be obtained based on the constructed model, such as whether the system satisfies specific change patterns, including but not limited to satisfying linear and nonlinear conditions. Here, we take a discrete-time stochastic linear time-invariant system as an example, where the system is: in, It is the state of the system, and The previous state, For the adjacent next state, and These are control variables and disturbance variables. , , It is a stochastic system matrix with appropriate dimensions: The above formula gives a deterministic real matrix. Control input matrix Interference input matrix and random interference input sequences They can be set to have a mean of 0 and a variance of 1. These three random disturbance input sequences are independent and uncorrelated.

[0032] In this power generation system, the simultaneous existence and action of control and disturbance are considered. In addition, a stochastic system matrix is ​​used in the system state, control and disturbance to simulate the noise that exists in the power generation process under real conditions, so that the subsequent prediction is more accurate.

[0033] Since this power generation system is a linear quadratic system, its performance indicators can be set as follows: in, It is a positive definite matrix.

[0034] The use of performance metrics is primarily for finding the Nash equilibrium solution. This minimizes the performance metrics related to control and maximizes the performance metrics related to disturbances.

[0035] In addition to performance metrics, there are also value functions. After defining and obtaining the performance metrics, their corresponding value functions can be obtained, specifically: in, This is control under optimal conditions. This is the worst-case scenario of interference. Represents the mathematical expectation. It is a unique positive definite matrix.

[0036] Therefore, in the subsequent optimization and training process, the optimization status of the network can be judged based on performance indicators or value functions to determine whether the optimization is complete.

[0037] Step 102: Obtain historical power generation data and optimize the weights of the model network based on the historical power generation data to obtain the optimized intermediate strategy prediction model.

[0038] After the strategy prediction model is built and initialized, it will be optimized and trained. Specifically, during training, relevant historical power generation data is first acquired, and then the initialized strategy prediction model is optimized and trained based on the obtained historical power generation data. Here, the weights of the model network in the strategy prediction model are first optimized, which is called optimization training.

[0039] The entire policy prediction model contains multiple different or identical networks, each playing a different role. For example, the model network is used to predict the next adjacent system state during the optimization process, the control network is used to determine the control quantity, the disturbance network is used to determine the disturbance quantity, and the evaluation model is used to determine whether the entire model has been optimized.

[0040] During training optimization, one or more networks can be optimized separately depending on the actual situation. For example, during the model training optimization process, the model networks can be optimized independently. Specifically, refer to... Figure 3 , Figure 3 This is a flowchart illustrating the process of obtaining an intermediate strategy prediction model provided by the present invention. The process includes: Step 301: Based on time series, obtain the state information corresponding to each moment in the historical power generation data, and group and associate the state information to obtain several state groups, wherein each state group contains two state information and the moments are adjacent. Step 302: Use several sets of state groups as the training set for the model network and train the model network. Step 303: When training is completed, an intermediate policy prediction model is obtained based on the model network with optimized weights.

[0041] The model network is used to deduce the next state based on the current state. Therefore, during training, it is optimized based on state data from historical power generation data. Specifically, when obtaining historical power generation data, state information is acquired from the historical power generation data. Each state corresponds to a specific time point, and these time points have a certain chronological order. The obtained state information is then grouped based on the time sequence to obtain several state groups. These state groups are then used as the training set for training the model network to optimize its weights. Finally, the intermediate policy prediction model is obtained after optimization.

[0042] For example, each data point in the historical power generation data corresponds to a unique time and exists in a certain chronological order. For instance, if the historical power generation data includes 100 data points, based on chronological order, they are: G1, G2, ..., G... n ...G 100 Where n < 100, when grouping, based on different grouping methods, the resulting groups can be: (G1, G2), (G3, G4), ... (G... n-1 G n ), ... (G) 99 G 100 ), can also be: (G1, G2), (G2, G3), ... (G n-1 G n ), ... (G) 99 G 100 (Specific restrictions are not imposed.)

[0043] After grouping the historical data, the training set obtained from the grouping is used to train the model network. The model network is pre-initialized with weights. After determining the network structure of the model network, the weights of the model network can be initialized randomly within the range of (-1, 1), and then the initialized model network is optimized and trained.

[0044] Reference Figure 4 , Figure 4 This is a flowchart illustrating the process of training a model network provided by the present invention, wherein the process includes: Step 401: Based on the timing, divide the state information in several state groups into input states and output states; Step 402: Input the input state into the model network to obtain the predicted state corresponding to the input state; Step 403: Calculate the state error based on the predicted state and the output state, and determine that training is complete when the state error is less than the preset error value.

[0045] After grouping the historical data, for each of the resulting state groups, which contains two state data, the states in each state group can be labeled based on time sequence as input state and output state respectively. Then, the input state is input into the model network, which will yield the predicted state for each input state. Finally, the state error between the predicted state and the output state is calculated to optimize the model network. When the optimization of the model network is completed, the obtained state error needs to be less than the preset error value.

[0046] For the obtained state group, the two states are set as the input state and output state of the model network according to the order or sequence of the states. For example, in the state group (G1, G2), G1 is set as the input state and G2 is set as the output state. G2 is the next state obtained based on G1 in the real situation. When optimizing the model network, the input state is input into the model network, and the model network outputs the predicted state. Then, by comparing the predicted state and the output state, it is determined whether the model network has been optimized. Specifically, the state error is obtained based on the predicted state and the output state. When the state error is small (less than the preset error value), it means that the model network has been optimized. Otherwise, it needs to be optimized and trained again.

[0047] It should be noted that there are multiple ways to determine whether optimization is complete. For example, the error between the predicted state and the output state can be calculated, and optimization can be determined when the error is less than 0.01. Alternatively, the similarity value between the predicted state and the output state can be used, and optimization can be determined when the similarity value is greater than a set threshold.

[0048] When optimizing the model network, the weights of the model network are determined, and then the model network in the policy prediction model is adjusted based on the obtained weights to obtain the intermediate policy prediction model.

[0049] Step 103: Based on historical power generation data, train the intermediate strategy prediction model to obtain the optimized evaluation network, control network, and disturbance network.

[0050] After optimizing the model network and obtaining the intermediate policy prediction model, the other networks of the overall policy prediction model will be optimized and trained. Specifically, when optimizing and training the other networks, training will also be based on historical power generation data. By training the intermediate policy prediction model, the weights of the control network, interference network, and evaluation network will be optimized and adjusted.

[0051] When training the intermediate policy prediction model, the weight optimization of the control network, interference network, and evaluation network is carried out simultaneously. At this time, the three networks that need to be optimized are treated as a whole for training. By restricting the network, the robustness of the model is improved, and the accuracy of policy prediction is also improved.

[0052] Reference Figure 5 , Figure 5 This is a flowchart illustrating the process of training an intermediate policy prediction model provided by the present invention, wherein the process includes: Step 501: Based on the time tag, historical power generation data is used as input and input into the first evaluation network of the control network, interference network and evaluation network respectively, to obtain control output, interference output and first evaluation output respectively; Step 502: Input the control output, disturbance output, and historical power generation data into the optimized model network, and output the model output. Step 503: Input the model output into the second evaluation network of the evaluation network, and obtain the second evaluation output; Step 504: Based on the first evaluation output and the second evaluation output, calculate the evaluation loss value of the evaluation network, and obtain the control loss value of the control network and the interference loss value of the interference network based on the optimal value function. Step 505: Based on the evaluation loss value, control loss value, and disturbance loss value, obtain the optimized evaluation network, control network, and disturbance network.

[0053] When training the intermediate strategy prediction model, historical power generation data is used as input to the control network, disturbance network, and first evaluation network, and the outputs are respectively the control output, disturbance output, and first evaluation output. Then, the control output and disturbance output are input into the model network together, and the historical power generation data is also input into the model network. The model network processes the data and outputs the model output. The model output is then used as input to the second evaluation network to obtain the second evaluation output. Based on the first and second evaluation outputs, the value loss of the evaluation network is obtained. At the same time, the control loss of the control network and the disturbance loss of the disturbance network are obtained according to the optimal value function. Finally, the weights of the control network, disturbance network, and evaluation network are optimized based on the obtained loss values.

[0054] In the actual training process, the output of the evaluation network is an approximate function. Using historical power generation data as input to the first evaluation network, an approximate function for each state can be obtained. After processing by the control network, disturbance network, and model network, the data is input into the second evaluation network. Since the model network itself predicts and outputs the next state based on the current state, an approximate function for the next state can be obtained for each state. At this point, for the evaluation network, its loss function is the difference between the two approximate functions. Specifically: ,in The output of the first evaluation network, This is the output of the second evaluation network.

[0055] The process of determining the loss values ​​of the control network and the interference network includes: taking the derivative of the control network based on the optimal value function to obtain a first derivative result, and calculating the control loss value corresponding to the control network based on the first derivative result and the control output; and taking the derivative of the interference network based on the optimal value function to obtain a second derivative result, and calculating the interference loss value corresponding to the control network based on the second derivative result and the interference output.

[0056] The optimal value function is defined as follows: if the expected reward of policy π is greater than that of policy π′ in all states, then policy π is said to be better than π′. In other words, π ≥ π′ if and only if vπ(s) ≥ vπ′(s) for all states s ∈ S. Naturally, the optimal policy is the policy that is better than all other policies. There may be more than one optimal policy, but they can be uniformly represented as π∗, and they all have the same state-value function, i.e., the optimal value function.

[0057] Therefore, when optimizing the control network and the interference network, optimization is performed based on the optimal value function. First, the optimal value function corresponding to the control network and the interference network is determined. Then, the derivative of the optimal value function is calculated, and the control loss value corresponding to the control network and the interference loss value corresponding to the interference network are calculated based on the derivative result and the corresponding network output.

[0058] When determining whether the control network, interference network, and model network have been optimized based on the obtained loss values, this can be done by analyzing the loss values. Specifically, this includes comparing the evaluation loss value with the loss threshold to determine whether the evaluation network, control network, and interference network have been optimized. If optimization is determined to be complete, the optimized evaluation network, control network, and interference network are obtained. Determining whether optimization is complete includes: if the evaluation loss value is less than or equal to the loss threshold, optimization is determined to be complete; if the evaluation loss value is greater than the loss threshold, optimization is determined to be incomplete.

[0059] During training, the completion of network training and optimization is determined based on the loss value of each network. Upon completion, the parameter information for each network can be determined, allowing for the prediction of power generation strategies during the overall model training process. To determine completion, the loss value of each network is compared to a corresponding threshold. If the loss value is less than or equal to the set threshold, optimization is considered complete; otherwise, further training is required.

[0060] Furthermore, when determining whether optimization is complete, the total loss value of the policy prediction model can be obtained based on the evaluation loss value, control loss value, and disturbance loss value. This total loss value is then used to determine whether optimization is complete. Specifically, when training the intermediate policy prediction model, it is optimized and trained as a whole. Therefore, to determine whether optimization is complete, the evaluation loss value, control loss value, and disturbance loss value can be summed to obtain the total loss value of the policy prediction model. This total loss value is then compared with preset values ​​to determine whether optimized control network, disturbance network, and evaluation network can be obtained. If so, these three networks are directly obtained; otherwise, step 501 is executed until the loss value meets the set conditions, such as the total loss value being less than or equal to the preset value.

[0061] When optimization is required, the weights of all networks except the model network need to be updated, and step 501 should be executed after the update. The specific update method for the weights of each network can be as follows: For the evaluation network, the weight update rule is as follows: For the control network, the weight update rule is as follows: For interference networks, the weight update rule is as follows: If the current optimized weights do not meet the set conditions, the weights of the control network, interference network, and evaluation network will be updated according to the set equal weight update rules before final determination.

[0062] Step 104: Determine whether the trained policy prediction model has converged, and obtain the policy prediction model when convergence is confirmed.

[0063] After optimizing the weights of each network in the policy prediction model, it is determined whether the optimized and trained policy prediction model has converged. If convergence is confirmed, the policy prediction model is obtained. Specifically, during the training process, the completion and termination of training are determined based on the actual training results. A common method is to determine whether the trained policy prediction model has converged. Furthermore, a test can be added during the convergence determination process to obtain a usable policy prediction model if the model converges and meets the test requirements.

[0064] To determine whether convergence has occurred, a predetermined number of training iterations can be set. Convergence is confirmed when the predetermined number of training iterations is reached, at which point the weights received by each network will be retained. Otherwise, training will continue to execute step 103.

[0065] Furthermore, after convergence is confirmed, the prediction accuracy of the strategy prediction model can be obtained, and then the usability of the model can be determined based on the prediction accuracy. During accuracy verification, historical data from the power generation process can be obtained through sampling, and then used as a test set to test the accuracy of the converged strategy prediction model. If the accuracy meets the set conditions (e.g., accuracy exceeds a set value), a strategy prediction model suitable for strategy prediction is obtained.

[0066] It should be noted that when a usable policy prediction model is obtained, in addition to determining whether it has converged and whether the accuracy meets the conditions, the accuracy can also be judged directly after the weights of each network are optimized. The accuracy can be used to determine whether training is complete. Taking the aforementioned convergence judgment training number as an example, the number of training times can be disregarded. Only when the loss meets the conditions can the accuracy be used to further judge the model.

[0067] Furthermore, the training of the strategy prediction model is for subsequent use, that is, to obtain the corresponding power generation strategy based on the current state. The power generation strategy mainly includes the control strategy and the disturbance strategy of the power generation system, so as to realize the control and regulation of the power generation process.

[0068] Specifically, the strategy prediction process includes: responding to the power generation prediction command and receiving the input initial state; obtaining the control strategy and disturbance strategy in the initial state based on the trained strategy prediction model; and performing power generation control and regulation based on the control strategy and disturbance strategy.

[0069] During the use of the model, a new initial state is given, and then the control network and interference network in the policy prediction model are processed to obtain the control policy and interference policy corresponding to the new initial state.

[0070] In the training method of the power generation strategy prediction model in the above embodiment, a corresponding power generation strategy prediction model is constructed based on the power generation system. Several different networks are set up, and after initializing the weights of each network, the model network that will output the next state is first trained and its weights optimized. By training separately, the influence of other networks on the model network is reduced, improving the robustness and prediction accuracy of the model network. Upon completion of the optimization training, the strategy prediction model is treated as a whole, and the weights of each network are optimized. Through the mutual constraints between the control network, the interference network, and the evaluation network, the robustness and training efficiency of the model are improved. Finally, convergence is used to ensure that the model meets the usage requirements. By introducing control and interference factors, a better comprehensive effect of power generation control is achieved, ensuring the stable operation of the equipment and improving the accuracy of power generation strategy prediction.

[0071] The training apparatus for the strategy prediction model of the power generation process provided by the present invention will be described below. The training apparatus for the strategy prediction model of the power generation process described below can be referred to in correspondence with the training method for the strategy prediction model of the power generation process described above.

[0072] Figure 6 This is a schematic diagram of the structure of the training device for the strategy prediction model of the power generation process provided by the present invention, as shown in the figure. Figure 7 As shown, the training device 600 for the strategy prediction model of the power generation process includes: The initial adjustment module 601 is used to initialize the initial policy prediction model, which includes a control network, a disturbance network, a model network, and an evaluation network. The first optimization module 602 is used to acquire historical power generation data and optimize the weights of the model network based on the historical power generation data to obtain the optimized intermediate strategy prediction model. The second optimization module 603 is used to train the intermediate strategy prediction model based on historical power generation data to obtain the optimized evaluation network, control network and disturbance network. The convergence judgment module 604 is used to determine whether the trained policy prediction model has converged, and obtains the policy prediction model when convergence is determined.

[0073] Based on the above embodiments, the first optimization module is further configured to: Based on time series, the state information corresponding to each moment is obtained from historical power generation data, and the state information is grouped and associated to obtain several state groups, where each state group contains two state information and the time is adjacent. Several sets of states are used as the training set for the model network, and the model network is trained. Once training is complete, an intermediate policy prediction model is obtained based on the weighted optimized model network.

[0074] Based on the above embodiments, the first optimization module is further configured to: Based on timing, the state information in several state groups is divided into input state and output state; The input state is fed into the model network to obtain the predicted state corresponding to the input state; The state error is calculated based on the predicted state and the output state, and training is considered complete when the state error is less than a preset error value.

[0075] Based on the above embodiments, the second optimization module is further configured to: Based on time tags, historical power generation data is used as input and fed into the first evaluation network of the control network, disturbance network, and evaluation network, respectively, to obtain control output, disturbance output, and first evaluation output. The control output, disturbance output, and historical power generation data are input into the optimized model network, and the output is the model output. The model output is input into the second evaluation network of the evaluation network, and the second evaluation output is obtained. Based on the first evaluation output and the second evaluation output, the evaluation loss value of the evaluation network is calculated, and the control loss value of the control network and the disturbance loss value of the disturbance network are obtained according to the optimal value function. Based on the evaluation loss value, control loss value, and disturbance loss value, the optimized evaluation network, control network, and disturbance network are obtained.

[0076] Based on the above embodiments, the second optimization module is further configured to: The control network is differentiated according to the optimal value function to obtain the first derivative result. The control loss value corresponding to the control network is calculated based on the first derivative result and the control output. The interference network is differentiated according to the optimal value function to obtain the second derivative result. The interference loss value corresponding to the control network is calculated based on the second derivative result and the interference output.

[0077] Based on the above embodiments, the second optimization module is further configured to: The evaluation loss value is compared with the loss threshold to determine whether the evaluation network, control network and interference network have been optimized. When the optimization is determined to be complete, the optimized evaluation network, control network and interference network are obtained. Determining whether optimization is complete includes: If the evaluation loss value is less than or equal to the loss threshold, then the optimization is considered complete. If the evaluation loss value is greater than the loss threshold, then the optimization is determined to be incomplete.

[0078] Based on the above embodiments, the training device for the strategy prediction model of the power generation process further includes a strategy prediction module, used for: Responding to power generation prediction commands, it receives the initial input state; Based on the trained policy prediction model, the control policy and disturbance policy in the initial state are obtained; Power generation control and regulation are carried out based on control and disturbance strategies.

[0079] Figure 7 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 7 As shown, the electronic device may include a processor 710, a communications interface 720, a memory 730, and a communication bus 740, wherein the processor 710, communications interface 720, and memory 730 communicate with each other via the communication bus 740. The processor 710 can call logic instructions in the memory 730 to execute a training method for a strategy prediction model of the power generation process. This method includes: initializing an initial strategy prediction model, which includes a control network, an interference network, a model network, and an evaluation network; acquiring historical power generation data and optimizing the weights of the model network based on the historical power generation data to obtain an optimized intermediate strategy prediction model; training the intermediate strategy prediction model based on the historical power generation data to obtain optimized evaluation, control, and interference networks; determining whether the trained strategy prediction model has converged, and obtaining the strategy prediction model upon confirmation of convergence.

[0080] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0081] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the training method for the strategy prediction model of the power generation process provided by the above methods. The method includes: initializing an initial strategy prediction model, wherein the strategy prediction model includes a control network, an interference network, a model network, and an evaluation network; acquiring historical power generation data and optimizing the weights of the model network based on the historical power generation data to obtain an optimized intermediate strategy prediction model; training the intermediate strategy prediction model based on the historical power generation data to obtain an optimized evaluation network, a control network, and an interference network; determining whether the trained strategy prediction model has converged, and obtaining the strategy prediction model when convergence is determined.

[0082] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a training method for a strategy prediction model of the power generation process provided by the methods described above. The method includes: initializing an initial strategy prediction model, wherein the strategy prediction model includes a control network, an interference network, a model network, and an evaluation network; acquiring historical power generation data and optimizing the weights of the model network based on the historical power generation data to obtain an optimized intermediate strategy prediction model; training the intermediate strategy prediction model based on the historical power generation data to obtain optimized evaluation, control, and interference networks; determining whether the trained strategy prediction model has converged, and obtaining the strategy prediction model upon determining convergence.

[0083] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0084] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0085] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A training method for a strategy prediction model of a power generation process, characterized in that, include: The initial policy prediction model is initialized, wherein the policy prediction model includes a control network, a disturbance network, a model network, and an evaluation network. Historical power generation data is acquired, and the model network is weighted and optimized based on the historical power generation data to obtain an optimized intermediate strategy prediction model. Based on historical power generation data, the intermediate strategy prediction model is trained to obtain the optimized evaluation network, the control network, and the interference network. Determine whether the trained policy prediction model has converged, and obtain the policy prediction model when convergence is confirmed. The step of training the intermediate strategy prediction model based on historical power generation data to obtain the optimized evaluation network, control network, and interference network includes: inputting the historical power generation data as input based on time labels into the first evaluation network of the control network, interference network, and evaluation network respectively, to obtain control output, interference output, and first evaluation output respectively; inputting the control output, interference output, and historical power generation data into the optimized model network to obtain model output; inputting the model output into the second evaluation network of the evaluation network to obtain second evaluation output; calculating the evaluation loss value of the evaluation network based on the first evaluation output and the second evaluation output, and obtaining the control loss value corresponding to the control network and the interference loss value corresponding to the interference network based on the optimal value function; and obtaining the optimized evaluation network, control network, and interference network based on the evaluation loss value, control loss value, and interference loss value. The step of obtaining the control loss value corresponding to the control network and the interference loss value corresponding to the interference network according to the optimal value function includes: taking the derivative of the control network according to the optimal value function to obtain a first derivative result, and calculating the control loss value corresponding to the control network according to the first derivative result and the control output; taking the derivative of the interference network according to the optimal value function to obtain a second derivative result, and calculating the interference loss value corresponding to the control network according to the second derivative result and the interference output.

2. The training method for the strategy prediction model of the power generation process according to claim 1, characterized in that, The step of optimizing the model network based on the historical power generation data to obtain the optimized intermediate strategy prediction model includes: Based on time series, the state information corresponding to each moment is obtained from the historical power generation data, and the state information is grouped and associated to obtain several state groups, wherein the state group contains two state information and the time is adjacent. The aforementioned sets of states are used as the training set for the model network, and the model network is trained accordingly. When training is complete, an intermediate policy prediction model is obtained based on the model network with optimized weights.

3. The training method for the strategy prediction model of the power generation process according to claim 2, characterized in that, The step of using the plurality of state groups as the training set of the model network and training the model network includes: Based on the timing sequence, the state information in the several groups of state groups is divided into input state and output state; The input state is input into the model network to obtain the predicted state corresponding to the input state; The state error is calculated based on the predicted state and the output state, and training is considered complete when the state error is less than a preset error value.

4. The training method for the strategy prediction model of the power generation process according to claim 1, characterized in that, The step of obtaining the optimized evaluation network, control network, and interference network based on the evaluation loss value, the control loss value, and the interference loss value includes: The evaluation loss value is compared with the loss threshold to determine whether the evaluation network, the control network and the interference network have been optimized. When it is determined that the optimization has been completed, the optimized evaluation network, the control network and the interference network are obtained. The determination of whether optimization is complete includes: If the evaluated loss value is less than or equal to the loss threshold, then the optimization is determined to be complete; If the evaluation loss value is greater than the loss threshold, then the optimization is determined to be incomplete.

5. The training method for the strategy prediction model of the power generation process according to claim 1, characterized in that, The method further includes: Responding to power generation prediction commands, it receives the initial input state; Based on the trained policy prediction model, the control policy and the disturbance policy in the initial state are obtained. Based on the control strategy and the interference strategy, power generation control and regulation are performed.

6. A training device for a strategy prediction model of a power generation process, characterized in that, An initial adjustment module is used to initialize the initial policy prediction model, wherein the policy prediction model includes a control network, a disturbance network, a model network, and an evaluation network. The first optimization module is used to acquire historical power generation data and optimize the model network based on the historical power generation data to obtain an optimized intermediate strategy prediction model. The second optimization module is used to train the intermediate strategy prediction model based on historical power generation data to obtain the optimized evaluation network, the control network, and the interference network. The convergence determination module is used to determine whether the trained policy prediction model has converged, and to obtain the policy prediction model when convergence is determined. The step of training the intermediate strategy prediction model based on historical power generation data to obtain the optimized evaluation network, control network, and interference network includes: inputting the historical power generation data as input based on time labels into the first evaluation network of the control network, interference network, and evaluation network respectively, to obtain control output, interference output, and first evaluation output respectively; inputting the control output, interference output, and historical power generation data into the optimized model network to obtain model output; inputting the model output into the second evaluation network of the evaluation network to obtain second evaluation output; calculating the evaluation loss value of the evaluation network based on the first evaluation output and the second evaluation output, and obtaining the control loss value corresponding to the control network and the interference loss value corresponding to the interference network based on the optimal value function; and obtaining the optimized evaluation network, control network, and interference network based on the evaluation loss value, control loss value, and interference loss value. The step of obtaining the control loss value corresponding to the control network and the interference loss value corresponding to the interference network according to the optimal value function includes: taking the derivative of the control network according to the optimal value function to obtain a first derivative result, and calculating the control loss value corresponding to the control network according to the first derivative result and the control output; taking the derivative of the interference network according to the optimal value function to obtain a second derivative result, and calculating the interference loss value corresponding to the control network according to the second derivative result and the interference output.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements a training method for a strategy prediction model of the power generation process as described in any one of claims 1 to 5.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements a training method for a strategy prediction model of the power generation process as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Online control method based on reinforcement learning for thickener

    CN110393954A

  • Thermal generator set combustion control optimization method and device and readable storage medium

    CN110888401A

  • Light energy power generation prediction method and device, computer equipment and storage medium

    CN112348247A