An intelligent monitoring method and system for a new energy power station
By using new value function optimization reinforcement learning algorithms and pruning algorithms in new energy power plants to build an intelligent monitoring model, the problem of insufficient model prediction accuracy in new energy power plants is solved, and efficient predictive maintenance and learning efficiency improvement is achieved.
Patent Information
- Application Number
- CN202310244169.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-09
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2043-03-09
AI Technical Summary
The existing intelligent monitoring technology has insufficient model prediction accuracy in new energy power plants, making it difficult to achieve predictive maintenance, and lacks systematic research.
A new value function optimization reinforcement learning algorithm is used to build an intelligent monitoring model, including the target network and the evaluation network, combined with the pruning algorithm optimization model, and use real-time running data for prediction.
It improves the model prediction accuracy of new energy power station equipment, enhances predictive maintenance capabilities, reduces the demand for computing resources, and improves learning efficiency and generalization.
Smart Images

Figure CN116032020B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the technical field of intelligent monitoring, and particularly relates to an intelligent monitoring method and system for a new energy power station. Background Art
[0002] In order to reduce the strong dependence of the unit on the traditional DCS (Distributed Control System) operation mode, with the help of scientific and technological means and emerging technologies such as artificial intelligence, while improving the operation safety of the unit, the workload of operation personnel is reduced, and the operation and maintenance method of predictive maintenance is achieved. The concept of intelligent monitoring was first proposed in the thermal power field. The ideal functional requirements of intelligent monitoring mainly consist of aspects such as real-time transmission and processing of production data, intelligent monitoring, intelligent trend analysis of problems such as sudden changes in equipment parameters, auxiliary early warning of problems such as fault location, intelligent screen patrol and meter reading, and automatic reporting.
[0003] At present, intelligent monitoring technology generally stays at the level of integrated visualization of sensor data and simple statistical analysis. For the core element of the model, only a small number of stations have applied prediction-related models in a small range, but the application effect is not good. The models generally have problems such as difficult to upgrade, untimely data cleaning, slow model operation speed, and insignificant improvement in prediction accuracy. It can be said that there is still a long way to go to meet the ideal functional requirements of this technology. On the other hand, intelligent monitoring mostly focuses on the thermal power field, gas turbines, and hydropower fields, and no systematic research has been carried out in new energy power stations. Summary of the Invention
[0004] The present disclosure aims to at least solve one of the technical problems in the related technologies to some extent. For this purpose, the present disclosure provides an intelligent monitoring method and system for a new energy power station, and the main purpose is to improve the model prediction accuracy during intelligent monitoring.
[0005] According to an embodiment of the first aspect of the present disclosure, an intelligent monitoring method for a new energy power station is provided, including:
[0006] Construct a training data set, where the training data set includes historical operation data of in-situ measuring points of new energy power station equipment and action label values corresponding to the historical operation data;
[0007] Construct an intelligent monitoring model, where the intelligent monitoring model adopts a new value function optimization reinforcement learning algorithm, and the new value function optimization reinforcement learning algorithm includes a target network and an evaluation network. The input of the target network includes the operation data of in-situ measuring points of new energy power station equipment and the output of the evaluation network, and the output of the target network is an action target value; the evaluation network outputs feedback parameters based on the action target value and the action label value, and the feedback parameters include reward data and adjustment data;
[0008] Train the intelligent monitoring model using the training data set to obtain a trained intelligent monitoring model, and set the feedback parameter output by the evaluation network in the trained intelligent monitoring model to zero constantly to obtain a target intelligent monitoring model;
[0009] Obtain the real-time operation data of the on-site measurement points of the new energy power station equipment;
[0010] Input the real-time operation data into the target intelligent monitoring model to output the real-time action target value, so as to realize the intelligent monitoring of the new energy power station.
[0011] In an embodiment of the present disclosure, the evaluation network outputs a feedback parameter based on the action target value and the action label value, including: if the action target value and the action label value are consistent, the reward data output by the evaluation network is a non-zero value, and the adjustment data is zero; if the action target value and the action label value are inconsistent, the evaluation network obtains the operation data action combination in the built-in database, and outputs the feedback parameter based on the operation data action combination, the action target value and the action label value.
[0012] In an embodiment of the present disclosure, outputting the feedback parameter based on the operation data action combination, the action target value and the action label value includes: searching for a target combination that matches the action target value in the operation data action combination; if the target combination does not exist, the reward data output by the evaluation network is zero, and the adjustment data is the difference between the action target value and the action label value; if the target combination exists, the reward data output by the evaluation network is a non-zero value, the adjustment data is the difference between the action target value and the action label value, and the target combination is added to the experience replay pool.
[0013] In an embodiment of the present disclosure, a pruning algorithm is adopted when training the intelligent monitoring model using the training data set.
[0014] In an embodiment of the present disclosure, the pruning algorithm is a structural sparse pruning algorithm or a temporal sparse pruning algorithm.
[0015] In an embodiment of the present disclosure, the operation data of the on-site measurement points of the new energy power station equipment includes operation system data and operation environment data. The operation system data includes the voltage, current, active power, reactive power, and the total power grid connection power of the whole plant and single unit or equipment; the operation environment data includes at least one of temperature, irradiance, wind speed, and wind direction.
[0016] According to the second aspect embodiment of the present disclosure, there is also provided a new energy power station intelligent monitoring system, including:
[0017] A modeling module for constructing a training data set, the training data set including historical operation data of in-situ measurement points of new energy power station equipment and action label values corresponding to the historical operation data, and also for constructing an intelligent monitoring model, the intelligent monitoring model adopting a new value function optimized reinforcement learning algorithm, the new value function optimized reinforcement learning algorithm including a target network and an evaluation network, the input of the target network including the operation data of in-situ measurement points of new energy power station equipment and the output of the evaluation network, and the output of the target network being an action target value; the evaluation network outputs feedback parameters based on the action target value and the action label value, and the feedback parameters include reward data and adjustment data;
[0018] A training module for training the intelligent monitoring model using the training data set to obtain a trained intelligent monitoring model, and setting the feedback parameters output by the evaluation network in the trained intelligent monitoring model to zero constantly to obtain a target intelligent monitoring model;
[0019] An acquisition module for acquiring real-time operation data of in-situ measurement points of new energy power station equipment;
[0020] An intelligent monitoring module for inputting the real-time operation data into the target intelligent monitoring model to output a real-time action target value, thereby realizing intelligent monitoring of a new energy power station.
[0021] In an embodiment of the present disclosure, the modeling module is specifically configured to: if the action target value and the action label value are consistent, the reward data output by the evaluation network is a non-zero value and the adjustment data is zero; if the action target value and the action label value are inconsistent, the evaluation network obtains an operation data action combination in the built-in database and outputs feedback parameters based on the operation data action combination, the action target value and the action label value.
[0022] In an embodiment of the present disclosure, the training module adopts a structural sparse pruning algorithm or a temporal sparse pruning algorithm when training the intelligent monitoring model using the training data set.
[0023] According to an embodiment of the third aspect of the present disclosure, there is also provided a new energy power station intelligent monitoring device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the new energy power station intelligent monitoring method proposed in the embodiment of the first aspect of the present disclosure.
[0024] In one or more embodiments of the present disclosure, a training data set is constructed. The training data set includes historical operation data of in-situ measurement points of new energy power station equipment and action label values corresponding to the historical operation data. An intelligent monitoring model is constructed. The intelligent monitoring model uses a new value function optimized reinforcement learning algorithm. The new value function optimized reinforcement learning algorithm includes a target network and an evaluation network. The input of the target network includes the operation data of in-situ measurement points of new energy power station equipment and the output of the evaluation network. The output of the target network is an action target value. The evaluation network outputs feedback parameters based on the action target value and the action label value. The feedback parameters include reward data and adjustment data. The intelligent monitoring model is trained using the training data set to obtain a trained intelligent monitoring model. The feedback parameters output by the evaluation network in the trained intelligent monitoring model are constantly set to zero to obtain a target intelligent monitoring model. The real-time operation data of in-situ measurement points of new energy power station equipment is obtained. The real-time operation data is input into the target intelligent monitoring model to output a real-time action target value, thereby realizing intelligent monitoring of the new energy power station. In this case, an intelligent monitoring model is constructed using the new value function optimized reinforcement learning algorithm. In the new value function optimized reinforcement learning algorithm, the input of the target network includes not only the operation data of in-situ measurement points of new energy power station equipment, but also the feedback parameters output by the evaluation network. The feedback parameters are obtained using the action target value and the action label value. Thus, the constructed intelligent monitoring model synthesizes the operation data, action label values, reward data, and adjustment data of in-situ measurement points of new energy power station equipment to obtain the action target value, thereby improving the accuracy of model prediction.
[0025] Additional aspects and advantages of the present disclosure will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The above and / or additional aspects and advantages of the present disclosure will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, in which:
[0027] Figure 1 The flowchart showing the process of a new energy power station intelligent monitoring method provided by an embodiment of the present disclosure;
[0028] Figure 2 The block diagram showing the new energy power station intelligent monitoring system provided by an embodiment of the present disclosure;
[0029] Figure 3 The block diagram of the new energy power station intelligent monitoring device used to implement the new energy power station intelligent monitoring method of the embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the embodiments of the present disclosure as detailed in the appended claims.
[0031] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0032] In addition, the terms "first" and "second" are used only for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include at least one of the features. In the description of the present disclosure, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined. It should also be understood that the term "and / or" used in the present disclosure refers to and includes any or all possible combinations of one or more of the associated listed items.
[0033] The embodiments of the present disclosure will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present disclosure and should not be construed as a limitation of the present disclosure.
[0034] The present disclosure provides a method and system for intelligent monitoring of a new energy power station, and the main purpose is to improve the model prediction accuracy during intelligent monitoring.
[0035] In the first embodiment, Figure 1 A flowchart showing a method for intelligent monitoring of a new energy power station provided by an embodiment of the present disclosure is shown. As Figure 1 shown, the method for intelligent monitoring of a new energy power station includes:
[0036] Step S11 : constructing a training data set, where the training data set includes historical operating data of on-site measuring points of new energy power station equipment and action label values corresponding to the historical operating data.
[0037] It is easy to understand that the historical operation data of the local measurement points of the new energy power station equipment in step S11 refers to the historically stored operation data of the local measurement points of the new energy power station equipment, wherein the operation data of the local measurement points of the new energy power station equipment includes operation system data and operation environment data.
[0038] In step S11, the operating system data includes power plant generated data such as voltage, current, active power, reactive power, and total power connected to the grid for the entire plant and individual units or equipment. The operating system data in the embodiments of the present disclosure are not limited thereto.
[0039] In step S11, the operating environment data includes at least one of air temperature, irradiance, wind speed, and wind direction. Specifically, the operating environment data refers to production-related meteorological data from the power plant's own weather station. For example, in a photovoltaic scenario, the meteorological data includes air temperature, irradiance, and wind speed; in a wind power scenario, the meteorological data includes wind speed and wind direction. The operating environment data in the embodiments of the present disclosure is not limited to this.
[0040] In step S11, each set of historical operation data has a corresponding action label value, and all sets of historical operation data and corresponding action label values are constructed to obtain a training data set.
[0041] Step S12: construct an intelligent monitoring model. The intelligent monitoring model adopts a new value function optimization reinforcement learning algorithm. The new value function optimization reinforcement learning algorithm includes a target network and an evaluation network. The input of the target network includes the operating data of the on-site measurement points of the new energy power station equipment and the output of the evaluation network. The output of the target network is the action target value; the evaluation network outputs feedback parameters based on the action target value and the action label value. The feedback parameters include reward data and adjustment data.
[0042] In step S12, the novel value function optimized reinforcement learning algorithm may refer to the optimized DQN algorithm. The DQN algorithm is the Deep Q-Network algorithm. Understandably, the DQN algorithm is a value function optimized reinforcement learning algorithm combined with deep learning and is a commonly used deep reinforcement learning algorithm at present. During training, the learning result is corrected by updating the value function therein to achieve the learning effect. Among them, reinforcement learning is one of the paradigms and methodologies of machine learning, used to describe and solve the problem that an agent maximizes the reward or achieves a specific goal through learning strategies during the interaction with the environment. The common model of reinforcement learning is the standard Markov Decision Process (MDP). The algorithms used to solve reinforcement learning problems can be divided into two categories: policy search algorithms and value function algorithms. Deep learning models can be used in reinforcement learning to form deep reinforcement learning.
[0043] In step S12, the novel value function optimized reinforcement learning algorithm includes a target network and an evaluation network.
[0044] In step S12, the input of the target network includes the operation data of the in-situ measurement points of the new energy power station equipment and the output of the evaluation network, and the output of the target network is the action target value. Specifically, the target network performs operation simulation on the operation data of the in-situ measurement points of the new energy power station equipment and outputs the learning result, which is the action target value.
[0045] In this embodiment, the target network can adopt a Convolutional Neural Network (CNN). CNN is a type of feedforward neural network (Feedforward Neural Networks) that contains convolutional calculations and has a deep structure and is one of the representative algorithms of deep learning (deeplearning). The convolutional neural network has the ability of representation learning and can perform shift-invariant classification on the input information according to its hierarchical structure. It is a commonly used embedded neural network type in reinforcement learning algorithms.
[0046] In step S12, the evaluation network outputs feedback parameters based on the action target value and the action label value. The feedback parameters include reward data and adjustment data. Specifically, the evaluation network reads the action label value in the training dataset of the current control system through the background data interface program, compares the action label value with the action target value output by the target network, and outputs different feedback parameters.
[0047] In step S12, the evaluation network outputs feedback parameters based on the action target value and the action label value, including: if the action target value and the action label value are the same, the reward data output by the evaluation network is a non-zero value and the adjustment data is zero; if the action target value and the action label value are different, the evaluation network obtains the operation data action combination in the built-in database and outputs feedback parameters based on the operation data action combination, the action target value and the action label value.
[0048] In step S12, outputting feedback parameters based on the operation data action combination, the action target value and the action label value, including: searching for a target combination that matches the action target value in the operation data action combination; if the target combination does not exist, the reward data output by the evaluation network is zero and the adjustment data is the difference between the action target value and the action label value; if the target combination exists, the reward data output by the evaluation network is a non-zero value, the adjustment data is the difference between the action target value and the action label value, and the target combination is added to the experience replay pool.
[0049] Among them, the built-in database includes multiple groups of operation data action combinations, and each group of operation data action combinations includes operation data and an action target value. All the operation data action combinations contain all possible action target values (targets) and the corresponding operation data of the UI. All the operation data action combinations can be saved in the form of a dictionary when stored in the built-in database. When the evaluation network determines that the action target value and the action label value are different, it means that there is a deviation in the model prediction at this time. In the case of reinforcement learning at this time, the reward data is set to 0. At this time, however, the evaluation network tries to match with the help of the operation data action combination in the built-in database, that is, searches for a target combination that matches the action target value in the operation data action combination; if the target combination does not exist, that is, no match is successful in the built-in database, this is a complete failure learning experience in this scenario; if the target combination exists, that is, the match is successful, then rewards are given through the reward data, and the target combination is added to the self-owned experience replay pool of the DQN algorithm to assist subsequent decision-making. In this case, by generalizing the definition of the agent's exploration success, the number of rewards obtained by the agent is increased, the ability to handle complex problems is increased, the learning process is accelerated, and the problem of difficult learning caused by sparse rewards in the reinforcement learning of complex systems is improved.
[0050] In step S12, the reward data can be, for example, the number of rewards, which can be obtained through the reward function in the evaluation network. The reward function (Rewards) is a special signal in which the goal of the agent in reinforcement learning is formally represented. The reward function is passed to the agent (such as the intelligent monitoring model) through the environment. The goal of the agent is to maximize the total reward it receives. The reward function defines the learning rate of the agent in reinforcement learning: what needs to be maximized is not the current gain, but the long-term cumulative gain.
[0051] Step S13: Use the training data set to train the intelligent monitoring model to obtain a trained intelligent monitoring model, and set the feedback parameters output by the evaluation network in the trained intelligent monitoring model to zero constantly to obtain the target intelligent monitoring model.
[0052] In step S13, during the training process, in each training task, the evaluation network sends the output feedback parameters (i.e., reward data and adjustment data) to the target network. The target network automatically calibrates the value function internally according to this feedback parameter, and then uses the new weights to learn the historical operation data of the new moment (i.e., the new training task) again, and repeats the above process until the training ends. When the training is completed, the intelligent monitoring model is sensitive to changes in device measurement points.
[0053] In step S13, a pruning algorithm is used when training the intelligent monitoring model with the training data set.
[0054] In this embodiment, the pruning algorithm is a structural sparse pruning algorithm or a temporal sparse pruning algorithm.
[0055] Specifically, considering that learning using the DQN algorithm will use a large amount of data and a large amount of data redundancy will be generated during the operation, and as the algorithm runs, there will be a deceleration problem that becomes more and more serious and grows exponentially in severity for its running speed. For the reinforcement learning algorithm applied to intelligent monitoring, the data redundancy problem of its model needs to be designed separately to make it have an effective self-pruning ability. Therefore, in the embodiments of the present disclosure, when training, the DQN optimization algorithm (i.e., the optimized DQN algorithm) also adopts two types of optimized pruning algorithms, so that it can perform pruning from two directions of structural sparsity and temporal sparsity respectively. Thus, the demand of the neural network for computing resources is effectively reduced, the operation burden is alleviated, and it has on-site testability.
[0056] For the structural sparse pruning algorithm, in this embodiment, based on the absolute value (that is, the closer a weight is to 0, the lower the importance of this weight), a small proportion of pruning is performed on the weights of the neural network for each operation, and then the remaining weights are reset to the initial values. For the pruning weights, the algorithm in this embodiment is set to fluctuate between 10% - 20%, and the algorithm adapts within this range according to the pruning effect to maintain the balance between the operation speed and operation accuracy during continuous operations.
[0057] The calculation formula for the pruning weight involved is as shown in Equation (1):
[0058]
[0059] In the formula, a represents the pruning weight, A represents the total weight, the pruning weight remains consistent as a hyperparameter during the algorithm operation, i is the pruning iteration number; is the pruning rate, that is, the percentage of the trimmed weight for each iteration.
[0060] For the temporal sparse pruning algorithm, in this embodiment, temporal sparsity is achieved through Equation (2):
[0061]
[0062] In the formula, Δy k+1 (t) is the change value of a certain convolutional neural network layer at a certain moment, k is the current neural network layer number; t is the current step, y k+1 is the output of the (k + 1)-th layer neural network; W k is the weight matrix of the network at layer k; y k is the output of the k-th layer neural network and is also the input of the (k + 1)-th layer neural network; x k is the independent variable of the k-th layer neural network, and the neural network will apply the rule to obtain y k+1 . As shown in Equation 2, the change value of the output of the convolutional layer contained in the DQN algorithm can be calculated in real time by tracking the outputs of each layer. Additionally, before actually updating the neural network through Equation (2), an anticipatory calculation of the output change value is performed, and the current data of the neural network is reduced through Equation (3), significantly reducing the subsequent computational amount and achieving the effect of temporal sparse pruning of the convolutional neural network. Equation (3) satisfies:
[0063]
[0064] In the formula, T is the threshold added in each round, i is the current neuron, and only when the output change value of each neuron in the convolutional layer exceeds this threshold will the subsequent neurons be recalculated. represents the normal neuron output formula. Y prev(t) is the reduced neuron output. In this case, it is reflected by truncating neurons with insufficient weights, and the reduction in the overall computational amount of the neural network brought about by no longer calculating achieves the savings of computing power and time.
[0065] In this embodiment, the feedback parameter output by the evaluation network in the trained intelligent monitoring model is constantly set to zero to obtain the target intelligent monitoring model.
[0066] In some other embodiments, the feedback parameter may not be set, and the trained intelligent monitoring model can be directly used as the target intelligent monitoring model for intelligent monitoring in actual scenarios.
[0067] Step S14: Obtain the real-time operation data of the in-situ measurement points of the new energy power station equipment.
[0068] Understandably, the real-time operation data of the in-situ measurement points of the new energy power station equipment refers to the operation data of the in-situ measurement points of the new energy power station equipment obtained in real time. The types of data included in the real-time operation data in step S14 are the same as those included in the historical operation data in step S11.
[0069] Step S15: Input the real-time operation data into the target intelligent monitoring model to output the real-time action target value, thereby realizing the intelligent monitoring of the new energy power station.
[0070] In step S15, monitor the real-time action target value to determine whether there is an abnormal upward or downward trend in the in-situ measurement points of the new energy power station equipment. When an abnormal upward or downward trend appears, display the result on the display screen in the centralized control room in advance and give an alarm reminder, so as to give the operators more reaction time, thereby contributing to the predictive maintenance of the new energy power station.
[0071] In step S15, it also includes monitoring the difference between the real-time action target value output in real time and the theoretical action value. When a certain threshold is reached, it indicates that the accuracy of the model has decreased, and then the intelligent monitoring model is retrained.
[0072] In some other embodiments, the intelligent monitoring method for the new energy power station in the embodiments of the present disclosure can also use the basic DQN algorithm or other improved DQN algorithms for intelligent monitoring of the new energy power station. In addition, the optimized reward sparsity method proposed in the intelligent monitoring method for the new energy power station in the embodiments of the present disclosure is applied to similar problems or scenarios; the proposed structural sparsity pruning method is applied to similar problems or scenarios; the proposed temporal sparsity pruning method is applied to similar problems or scenarios; for the proposed temporal sparsity and structural sparsity pruning method, the hyperparameters such as the pruning weights that are fixed in the design are fine-tuned during the operation of the algorithm.
[0073] In the intelligent monitoring method for a new energy power station according to an embodiment of the present disclosure, a training data set is constructed. The training data set includes historical operation data of in-situ measurement points of new energy power station equipment and action label values corresponding to the historical operation data. An intelligent monitoring model is constructed. The intelligent monitoring model adopts a novel value function optimization reinforcement learning algorithm. The novel value function optimization reinforcement learning algorithm includes a target network and an evaluation network. The input of the target network includes the operation data of in-situ measurement points of new energy power station equipment and the output of the evaluation network. The output of the target network is an action target value. The evaluation network outputs feedback parameters based on the action target value and the action label values. The feedback parameters include reward data and adjustment data. The intelligent monitoring model is trained using the training data set to obtain a trained intelligent monitoring model. The feedback parameters output by the evaluation network in the trained intelligent monitoring model are constantly set to zero to obtain a target intelligent monitoring model. The real-time operation data of in-situ measurement points of new energy power station equipment is obtained. The real-time operation data is input into the target intelligent monitoring model to output a real-time action target value, thereby realizing the intelligent monitoring of the new energy power station. In this case, an intelligent monitoring model is constructed using the novel value function optimization reinforcement learning algorithm. In the novel value function optimization reinforcement learning algorithm, the input of the target network includes not only the operation data of in-situ measurement points of new energy power station equipment, but also the feedback parameters output by the evaluation network. The feedback parameters are obtained using the action target value and the action label values. Thus, the constructed intelligent monitoring model synthesizes the operation data, action label values, reward data, and adjustment data of in-situ measurement points of new energy power station equipment to obtain the action target value, thereby improving the accuracy of model prediction. The intelligent monitoring method of the present disclosure proposes a value function optimization algorithm that has never been applied in the fields of power generation, energy, and industry. It is an intelligent monitoring method for a new energy power station based on a novel value function optimization reinforcement learning method. Based on the intelligent monitoring method of the present disclosure, it not only provides a worthy attempt for the application of reinforcement learning in power plants; it can also, through the in-depth utilization of wrong explorations, generalize the definition of successful exploration by the intelligent agent, increase the number of rewards obtained by the intelligent agent, thereby accelerating the learning process; and it can also, through the self-pruning function, self-prune the convolutional network in DQN during the long-term analysis of massive data while ensuring the learning effect, maintain the relative sparsity of the network structure, and keep the operation speed of the DQN algorithm. Another advantage brought by this pruning method is to enhance the robustness of the network, further optimize the algorithm performance, and improve the learning efficiency and generalization of reinforcement learning in the industrial field.
[0074] The following is an embodiment of the system of the present disclosure, which can be used to execute the embodiment of the method of the present disclosure. For details not disclosed in the embodiment of the system of the present disclosure, please refer to the embodiment of the method of the present disclosure.
[0075] Please refer to Figure 2 , Figure 2The block diagram of the intelligent monitoring system for a new energy power station provided by an embodiment of the present disclosure is shown. The intelligent monitoring system for the new energy power station can be implemented as all or part of the system through software, hardware, or a combination of both. The intelligent monitoring system 10 for the new energy power station includes a modeling module 11, a training module 12, an acquisition module 13, and an intelligent monitoring module 14, where:
[0076] The modeling module 11 is used to construct a training data set. The training data set includes historical operation data of in-situ measurement points of new energy power station equipment and action label values corresponding to the historical operation data. It is also used to construct an intelligent monitoring model. The intelligent monitoring model adopts a new value function optimization reinforcement learning algorithm. The new value function optimization reinforcement learning algorithm includes a target network and an evaluation network. The input of the target network includes the operation data of in-situ measurement points of new energy power station equipment and the output of the evaluation network. The output of the target network is an action target value; the evaluation network outputs feedback parameters based on the action target value and the action label value. The feedback parameters include reward data and adjustment data;
[0077] The training module 12 is used to train the intelligent monitoring model with the training data set to obtain a trained intelligent monitoring model, and set the feedback parameters output by the evaluation network in the trained intelligent monitoring model to zero constantly to obtain a target intelligent monitoring model;
[0078] The acquisition module 13 is used to acquire the real-time operation data of in-situ measurement points of new energy power station equipment;
[0079] The intelligent monitoring module 14 is used to input the real-time operation data into the target intelligent monitoring model to output a real-time action target value, so as to realize the intelligent monitoring of the new energy power station.
[0080] Optionally, the modeling module 11 is specifically used for: if the action target value and the action label value are consistent, the reward data output by the evaluation network is a non-zero value, and the adjustment data is zero; if the action target value and the action label value are inconsistent, the evaluation network obtains the operation data action combination in the built-in database, and outputs feedback parameters based on the operation data action combination, the action target value, and the action label value.
[0081] Optionally, the training module 12 adopts a structural sparse pruning algorithm or a temporal sparse pruning algorithm when training the intelligent monitoring model with the training data set.
[0082] It should be noted that when the intelligent monitoring system for new energy power stations provided in the above embodiments executes the intelligent monitoring method for new energy power stations, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the intelligent monitoring device for new energy power stations is divided into different functional modules to complete all or part of the functions described above. In addition, the intelligent monitoring system for new energy power stations provided in the above embodiments and the embodiments of the intelligent monitoring method for new energy power stations belong to the same concept. The implementation process is detailed in the method embodiments and will not be repeated here.
[0083] The serial numbers of the above embodiments of the present disclosure are only for description and do not represent the advantages or disadvantages of the embodiments.
[0084] In the intelligent monitoring system for a new energy power station according to an embodiment of the present disclosure, a modeling module is configured to construct a training data set. The training data set includes historical operation data of on-site measuring points of new energy power station equipment and action label values corresponding to the historical operation data. The modeling module is further configured to construct an intelligent monitoring model. The intelligent monitoring model uses a novel value function optimization reinforcement learning algorithm. The novel value function optimization reinforcement learning algorithm includes a target network and an evaluation network. The input of the target network includes the operation data of on-site measuring points of new energy power station equipment and the output of the evaluation network. The output of the target network is an action target value. The evaluation network outputs feedback parameters based on the action target value and the action label value. The feedback parameters include reward data and adjustment data. A training module is configured to train the intelligent monitoring model using the training data set to obtain a trained intelligent monitoring model, and set the feedback parameters output by the evaluation network in the trained intelligent monitoring model to zero constantly to obtain a target intelligent monitoring model. An acquisition module is configured to acquire the real-time operation data of on-site measuring points of new energy power station equipment. An intelligent monitoring module is configured to input the real-time operation data into the target intelligent monitoring model to output a real-time action target value, thereby realizing intelligent monitoring of the new energy power station. In this case, an intelligent monitoring model is constructed using the novel value function optimization reinforcement learning algorithm. In the novel value function optimization reinforcement learning algorithm, the input of the target network includes not only the operation data of on-site measuring points of new energy power station equipment, but also the feedback parameters output by the evaluation network. The feedback parameters are obtained using the action target value and the action label value. Thus, the constructed intelligent monitoring model synthesizes the operation data, action label values, reward data, and adjustment data of on-site measuring points of new energy power station equipment to obtain the action target value, thereby improving the accuracy of model prediction. The intelligent monitoring system of the present disclosure proposes a value function optimization algorithm that has never been applied in the fields of power generation, energy, and industry. It is an intelligent monitoring system for a new energy power station based on a novel value function optimization reinforcement learning method. The intelligent monitoring system based on the present disclosure not only provides a worthy attempt for the application of reinforcement learning in power plants; it can also generalize the definition of successful exploration by intelligent agents by deeply utilizing error exploration, increase the number of rewards obtained by intelligent agents, thereby accelerating the learning process; and it can self-prune the convolutional network in DQN during the long-term analysis of massive data while ensuring the learning effect, maintaining the relative sparsity of the network structure, and keeping the operation speed of the DQN algorithm. Another advantage brought by this pruning method is to enhance the robustness of the network, further optimize the algorithm performance, and improve the learning efficiency and generalization of reinforcement learning in the industrial field.
[0085] According to an embodiment of the present disclosure, the present disclosure further provides an intelligent monitoring device for a new energy power station, a readable storage medium, and a computer program product.
[0086] Figure 3It is a block diagram of an intelligent monitoring device for a new energy power station, which is used to implement the intelligent monitoring method for a new energy power station in an embodiment of the present disclosure. The intelligent monitoring device for a new energy power station is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The intelligent monitoring device for a new energy power station can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable electronic devices, and other similar computing devices. The components, connections and relationships of the components, and the functions of the components shown in the present disclosure are only examples and are not intended to limit the implementation of the present disclosure described and / or claimed in the present disclosure.
[0087] As Figure 3 shown, the intelligent monitoring device 20 for a new energy power station includes a computing unit 21, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 22 or a computer program loaded from a storage unit 28 into a random access memory (RAM) 23. In the RAM 23, various programs and data required for the operation of the intelligent monitoring device 20 for a new energy power station can also be stored. The computing unit 21, the ROM 22, and the RAM 23 are connected to each other through a bus 24. An input / output (I / O) interface 25 is also connected to the bus 24.
[0088] Multiple components in the intelligent monitoring device 20 for a new energy power station are connected to the I / O interface 25, including: an input unit 26, such as a keyboard, a mouse, etc.; an output unit 27, such as various types of displays, speakers, etc.; a storage unit 28, such as a disk, an optical disc, etc., and the storage unit 28 is communicatively connected to the computing unit 21; and a communication unit 29, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 29 allows the intelligent monitoring device 20 for a new energy power station to exchange information / data with other intelligent monitoring devices for new energy power stations through a computer network such as the Internet and / or various telecommunication networks.
[0089] The computing unit 21 may be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 21 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 21 executes the various methods and processes described above, such as executing the intelligent monitoring method for a new energy power station. For example, in some embodiments, the intelligent monitoring method for a new energy power station may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 28. In some embodiments, part or all of the computer program may be loaded and / or installed onto the intelligent monitoring device 20 for a new energy power station via the ROM 22 and / or the communication unit 29. When the computer program is loaded into the RAM 23 and executed by the computing unit 21, one or more steps of the intelligent monitoring method for a new energy power station described above may be executed. Alternatively, in other embodiments, the computing unit 21 may be configured to execute the intelligent monitoring method for a new energy power station in any other suitable manner (e.g., by means of firmware).
[0090] The various embodiments of the systems and techniques described above in this disclosure may be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, that may receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0091] The program code for implementing the methods of this disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0092] In the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or new energy power station intelligent monitoring device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or electronic devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage electronic devices, magnetic storage electronic devices, or any suitable combination of the foregoing.
[0093] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0094] The systems and techniques described herein may be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of the communication network include: local area network (LAN), wide area network (WAN), the Internet, and blockchain network.
[0095] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server may also be a server of a distributed system, or a server combined with a blockchain.
[0096] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this disclosure does not impose any restrictions here.
[0097] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. An intelligent monitoring method for a new energy power station, characterized in that, Including: Construct a training data set, which includes historical operation data of on-site measurement points of new energy power station equipment and action label values corresponding to the historical operation data; Construct an intelligent monitoring model. The intelligent monitoring model adopts a new value function optimization reinforcement learning algorithm. The new value function optimization reinforcement learning algorithm includes a target network and an evaluation network. The input of the target network includes the operation data of on-site measurement points of new energy power station equipment and the output of the evaluation network. The output of the target network is an action target value; the evaluation network outputs feedback parameters based on the action target value and the action label value. The feedback parameters include reward data and adjustment data; Use the training data set to train the intelligent monitoring model to obtain a trained intelligent monitoring model. Set the feedback parameters output by the evaluation network in the trained intelligent monitoring model to zero constantly to obtain a target intelligent monitoring model; Obtain the real-time operation data of on-site measurement points of new energy power station equipment; Input the real-time operation data into the target intelligent monitoring model to output real-time action target values, so as to realize the intelligent monitoring of new energy power stations.
2. The intelligent monitoring method for a new energy power station according to claim 1, wherein The evaluation network outputs feedback parameters based on the action target value and the action label value, including: If the action target value is the same as the action label value, the reward data output by the evaluation network is a non-zero value, and the adjustment data is zero; If the action target value is different from the action label value, the evaluation network obtains the operation data action combination in the built-in database and outputs feedback parameters based on the operation data action combination, the action target value and the action label value.
3. The intelligent monitoring method for a new energy power station according to claim 2, wherein Outputting the feedback parameters based on the operation data action combination, the action target value and the action label value includes: Search for a target combination that matches the action target value in the operation data action combination; If the target combination does not exist, the reward data output by the evaluation network is zero, and the adjustment data is the difference between the action target value and the action label value; If the target combination exists, the reward data output by the evaluation network is a non-zero value, the adjustment data is the difference between the action target value and the action label value, and the target combination is added to the experience replay pool.
4. The intelligent monitoring method for a new energy power station according to claim 3, characterized in that, When using the training data set to train the intelligent monitoring model, a pruning algorithm is adopted.
5. The intelligent monitoring method for a new energy power station according to claim 4, characterized in that, The pruning algorithm is a structural sparse pruning algorithm or a temporal sparse pruning algorithm.
6. The intelligent monitoring method for a new energy power station according to claim 1, wherein The operation data of on-site measurement points of new energy power station equipment includes operation system data and operation environment data. The operation system data includes the voltage, current, active power, reactive power, and total plant grid-connected power of the whole plant and single units or equipment; the operation environment data includes at least one of temperature, irradiance, wind speed, and wind direction.
7. An intelligent monitoring system for a new energy power station, characterized in that, Including: A modeling module for constructing a training data set, where the training data set includes historical operation data of on-site measurement points of new energy power station equipment and action label values corresponding to the historical operation data, and is also used to construct an intelligent monitoring model. The intelligent monitoring model uses a new value function optimization reinforcement learning algorithm, and the new value function optimization reinforcement learning algorithm includes a target network and an evaluation network. The input of the target network includes the operation data of on-site measurement points of new energy power station equipment and the output of the evaluation network, and the output of the target network is an action target value; the evaluation network outputs feedback parameters based on the action target value and the action label value, and the feedback parameters include reward data and adjustment data; A training module for training the intelligent monitoring model using the training data set to obtain a trained intelligent monitoring model, and setting the feedback parameters output by the evaluation network in the trained intelligent monitoring model to zero constantly to obtain a target intelligent monitoring model; An acquisition module for acquiring real-time operation data of on-site measurement points of new energy power station equipment; An intelligent monitoring module for inputting the real-time operation data into the target intelligent monitoring model to output a real-time action target value, so as to realize intelligent monitoring of the new energy power station.
8. The intelligent monitoring system for new energy power stations according to claim 7, characterized in that, The modeling module is specifically used for: If the action target value is the same as the action label value, the reward data output by the evaluation network is a non-zero value, and the adjustment data is zero; if the action target value is different from the action label value, the evaluation network obtains the operation data action combination in the built-in database, and outputs feedback parameters based on the operation data action combination, the action target value and the action label value.
9. The intelligent monitoring system for a new energy power station according to claim 8, wherein, When the training module trains the intelligent monitoring model using the training data set, it uses a structural sparse pruning algorithm or a temporal sparse pruning algorithm.
10. An intelligent monitoring device for a new energy power station, characterized in that, Including: At least one processor; And a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the new energy power station intelligent monitoring method according to any one of claims 1-6.
Citation Information
Patent Citations
Power grid real-time adaptive decision-making method based on deep reinforcement learning
CN114217524A
Multi-agent reinforcement learning rolling scheduling method and device, equipment and storage medium
CN115310775A