Integrated operation and maintenance resource allocation method for power information network

CN115526545BActive Publication Date: 2026-10-09STATE GRID JIANGSU ELECTRIC POWER CO LTD SUZHOU BRANCH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211368632.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-03
Publication Date
2026-10-09
Estimated Expiration
2042-11-03

AI Technical Summary

Technical Problem

但是由于电力信息网络规模巨大时,网络拓扑结构往往非常复杂,运维资源的分配存在片面不充分的问题,并且常常是出现故障后维修,较少能做到主动、预防式的运维

Benefits of technology

[0018] Compared with the prior art, the present invention has the following advantages: it optimizes the resource allocation of the smart grid, takes into account both safe allocation and efficient allocation, and can proactively prevent local power grid outages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115526545B_ABST
    Figure CN115526545B_ABST
Patent Text Reader

Abstract

The application discloses an integrated operation and maintenance resource allocation method of a power information network, which comprises the following steps: regarding the distribution of power equipment and operation and maintenance equipment in a smart grid as a state, regarding the resource allocation of the power equipment and operation and maintenance equipment to various regions as an action, obtaining the return of different resource allocation actions of the smart grid in a certain state by using a Q value function in reinforcement learning, and selecting the resource allocation action with the optimal efficiency by using a greedy strategy; introducing a risk prediction model, inputting a certain state and the corresponding resource allocation action with the optimal efficiency into the risk prediction model to obtain a predicted future state, comparing the predicted future state with the states in a risk state information set, and if the preset safety requirement is met, executing the current resource allocation action with the optimal efficiency. The integrated operation and maintenance resource allocation method provided by the application considers the allocation efficiency on the basis of safe allocation, and can actively prevent the downtime of a local power network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power operation and maintenance resource allocation technology, and in particular to an integrated operation and maintenance resource allocation method for power information networks. Background Technology

[0002] With the advancement of digitalization and intelligentization, power information networks are becoming increasingly large-scale and complex, leading to more serious security risks and uneven resource allocation. To ensure a reliable and sufficient power supply to users, timely inspection and maintenance of power equipment are essential. However, due to the massive scale and highly complex topology of power information networks, the allocation of maintenance resources is often incomplete and fragmented, and maintenance is frequently reactive, with less emphasis on proactive and preventative maintenance.

[0003] Integrated operation, maintenance, and repair (O&M) in power companies represents a new working model and a significant reform. By employing integrated O&M management methods, production resources can be utilized more rationally and effectively, the professional skills of O&M personnel can be improved, and the quality of equipment operation and maintenance can be enhanced. This approach increases safety while reducing human resource consumption, thereby improving production efficiency and quality and ensuring the safe and sustainable development of power companies. However, how to allocate resources and achieve proactive, preventative O&M in integrated O&M remains a pressing issue that needs to be addressed.

[0004] The above background information is provided only to assist in understanding the inventive concept and technical solution of this invention. It does not necessarily belong to the prior art of this patent application, nor does it necessarily provide technical teaching. In the absence of clear evidence that the above information was disclosed before the filing date of this patent application, the above background information should not be used to evaluate the novelty and inventiveness of this application. Summary of the Invention

[0005] To overcome the shortcomings of existing technologies, this invention provides an integrated operation and maintenance resource allocation method for power information networks, the specific technical solution of which is as follows:

[0006] On the one hand, an integrated operation and maintenance resource allocation method for power information networks is provided, including the following steps:

[0007] An environmental model is performed on the allocation of smart grid operation and maintenance resources to obtain a deep learning model. The distribution of power equipment and operation and maintenance equipment in the smart grid is taken as the state, and the allocation of power equipment and operation and maintenance equipment to various regions is taken as the action. In the deep learning model, the Q-value function in reinforcement learning is used to obtain the reward of different resource allocation actions of the smart grid under a certain state. Based on this, the ε-greedy strategy is used to select the most efficient resource allocation action.

[0008] A hazard prediction model is introduced, collecting historical information on past smart grid allocation accidents. This historical information includes a set of hazard state information and a set of hazard action information. The hazard prediction model is trained using this historical information. After training, a certain state and its corresponding optimal resource allocation action are input into the hazard prediction model to obtain the predicted future state. The predicted future state is then compared with the states in the hazard state information set. If the preset safety requirements are met, the optimal resource allocation action is executed; otherwise, other resource allocation actions are executed.

[0009] Furthermore, the state of region i at time j and the corresponding optimal resource allocation action are input into the hazard prediction model to obtain the predicted next hidden state. The set of hazard state information corresponding to region i is subjected to feature extraction and flattening to obtain the corresponding hidden state set Z; the next predicted hidden state is then... The mean squared error loss is calculated with the set of hidden states Z. If the mean squared error loss exceeds a preset threshold, the most efficient resource allocation action is executed. Otherwise, the resource allocation action that maximizes the Q value is executed instead.

[0010] Furthermore, if other resource allocation actions are to be performed, they also need to undergo the safety check of the hazard prediction model. The state and the corresponding selected other resource allocation actions are input into the hazard prediction model to obtain the corresponding predicted future state. The predicted future state is then compared with the state in the hazard state information set. If the preset safety requirements are met, the corresponding resource allocation action is executed; otherwise, a different resource allocation action needs to be executed. Before the resource allocation action is executed, the safety requirements of the hazard prediction model must be met at least once.

[0011] Furthermore, the training process of the hazard prediction model includes: extracting features and flattening the hazard state information to obtain the corresponding hidden state; inputting the hidden state and the corresponding hazard action information into the prediction network of the hazard prediction model to obtain the predicted next hidden state; extracting features and flattening the actual next hidden state to obtain the actual next hidden state; calculating the mean squared error loss between the predicted next hidden state and the actual next hidden state; and performing gradient descent processing on the variance loss to train and update the network parameters of the hazard prediction model.

[0012] Furthermore, in the deep learning model, the agent needs to iteratively calculate the Q value each time it interacts with the environment to adapt to the resource allocation of the smart grid.

[0013] Furthermore, the set of hazard status information corresponding to region i is denoted as A three-layer convolutional neural network is used to process the danger state information corresponding to region i at time j. Feature extraction is performed, followed by flattening, and then compression is achieved using a network consisting of two fully connected layers to obtain the corresponding hidden states.

[0014] Furthermore, from the set of danger status information corresponding to region i Extract the state at time j+1. State The hidden state is obtained after feature extraction, flattening, and compression.

[0015] Furthermore, historical information on past smart grid distribution accidents is regularly supplemented and updated, and the hazard prediction model is supplemented and updated accordingly.

[0016] In another aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the aforementioned integrated operation and maintenance resource allocation method.

[0017] On the other hand, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the aforementioned integrated operation and maintenance resource allocation method.

[0018] Compared with the prior art, the present invention has the following advantages: it optimizes the resource allocation of the smart grid, takes into account both safe allocation and efficient allocation, and can proactively prevent local power grid outages. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the model network of the hazard prediction model in the integrated operation and maintenance resource allocation method for power information networks provided in this embodiment of the invention;

[0020] Figure 2 This is a schematic diagram of the overall architecture of the integrated operation and maintenance resource allocation method for power information networks provided in this embodiment of the invention;

[0021] Figure 3 This is a flowchart of the model training process in the integrated operation and maintenance resource allocation method for power information networks provided in this embodiment of the invention.

[0022] Figure 4 This is the model decision flowchart corresponding to the integrated operation and maintenance resource allocation method for power information networks provided in this embodiment of the invention. Detailed Implementation

[0023] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0024] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion.

[0025] Reinforcement learning, a method that seeks to maximize cumulative reward through interaction with the environment, is an important part of machine learning. In recent years, the application of reinforcement learning to practical applications has achieved remarkable results in fields such as automatic control. The Q-function in reinforcement learning can be used to evaluate the quality of actions performed in a given state. This invention applies this concept to smart grid resource allocation, using the distribution of power equipment and maintenance equipment as the state space, and smart grid resources as the multi-dimensional action space. The Q-function is used to evaluate the quality of resource allocation under the current power equipment and maintenance equipment conditions, enabling the smart grid to spontaneously perform resource scheduling and allocation without human intervention.

[0026] To achieve self-organized resource allocation, the Q-function from classical reinforcement learning is used as a key indicator to evaluate the quality of current resource allocation. Simultaneously, past data on insecure smart grid allocations is used as samples to train a pre-allocation prediction model network to avoid insecure resource allocation by the smart grid. When the smart grid needs to allocate resources, the allocation by the prediction model network is used to check whether the allocation is safe and reasonable, while the Q-learning method of reinforcement learning is continuously used to find the optimal allocation scheme. Through real-time evaluation and allocation, the entire smart grid dynamically and self-organizes resource allocation, thereby maximizing the allocation of smart grid resources.

[0027] In one embodiment of the present invention, an integrated operation and maintenance resource allocation method for power information networks is provided. This method performs environmental modeling on the allocation of smart grid operation and maintenance resources to obtain a deep learning model, using the distribution of power equipment and operation and maintenance equipment in the smart grid as the state. (See also...) Figure 1 and Figure 2Each region trains its own local model, and the overall results are then aggregated at the central node, where the global server outputs the final action selection. In this deep learning model, the Q-value function from reinforcement learning is used to obtain the rewards for different resource allocation actions of the smart grid under a certain state. Based on this, an ε-greedy policy is used to select the most efficient resource allocation action. In the ε-greedy policy, a larger ε value indicates a higher probability of using a random action and a lower probability of using the action with the highest current Q-value.

[0028] A hazard prediction model is introduced, and historical information on past smart grid allocation accidents is collected. The historical information includes a set of hazard state information and a set of hazard action information. The hazard prediction model is trained using the historical information, and the historical information on past smart grid allocation accidents is periodically supplemented and updated. The hazard prediction model is then further trained and updated accordingly.

[0029] After training, a certain state and its corresponding optimal resource allocation action are input into the hazard prediction model to obtain the predicted future state. The predicted future state is then compared with the states in the hazard state information set. If the preset safety requirements are met, the currently optimal resource allocation action is executed; otherwise, another resource allocation action is executed. If another resource allocation action is executed, it also needs to undergo a safety check by the hazard prediction model. The state and the corresponding selected other resource allocation actions are input into the hazard prediction model to obtain the corresponding predicted future state. The predicted future state is then compared with the states in the hazard state information set. If the preset safety requirements are met, the corresponding resource allocation action is executed; otherwise, a different resource allocation action is executed. Before each resource allocation action is executed, the safety requirements of the hazard prediction model must be met at least once. In a preferred embodiment, in region i, the state at time j and its corresponding optimal resource allocation action are input into the hazard prediction model to obtain the corresponding predicted next hidden state. The dangerous state information set is subjected to feature extraction and flattening to obtain the corresponding hidden state set Z; the next predicted hidden state is then... The mean squared error loss is calculated with the set of hidden states Z. If the mean squared error loss exceeds a preset threshold, the most efficient resource allocation action is executed. Otherwise, the resource allocation action that maximizes the Q value is executed instead.

[0030] Specifically, the integrated operation and maintenance resource allocation method includes the following steps:

[0031] Environment modeling

[0032] In a smart grid, the distribution of power equipment and maintenance equipment is represented by state X, and the allocation of resources to various regions is represented by action H. Simultaneously, a model for predicting future hazards is needed, i.e., a hazard prediction model. It is also necessary to collect historical information on past smart grid distribution accidents, with the set of hazardous state information for each region denoted as . The corresponding dangerous action information is recorded as follows

[0033] Step a mainly involves environmental modeling of power equipment and smart grids, which allows reinforcement learning methods to solve resource allocation problems.

[0034] b. Feature extraction of smart grid accident information

[0035] In reinforcement learning, specifically regarding prediction, training in the latent space is required to accurately predict future events. (See also...) Figure 1 Using the three-layer convolutional neural network in the model corresponding to region i Information on the status of dangerous accidents at time j in region i Feature extraction was performed to obtain Its feature extraction formula is: Will After flattening, it consists of two fully connected layers. Function networks compress it to a hidden state z of size 32. i Its formula is:

[0036] Step b mainly involves extracting state features from the environment to facilitate training in the latent space and improve prediction accuracy.

[0037] c Predicting future states

[0038] In acquiring the hidden state Then, associate it with the corresponding action. They are fed together as input into a prediction network consisting of three fully connected layers. Get the future state Future State This means predicting the future state of j+1 based on region i and time j.

[0039] d. Training the hazard prediction model

[0040] After steps b and c, the future state of the accident at time i+1 is obtained. From the collection of hazardous status information Extract the state of region i at time j+1. State The hidden state is obtained after the same process as in step b. To train the hazard prediction model more accurately Predicting the future state and the future state at the actual next moment Calculate the mean squared loss to obtain See Figure 3 And perform gradient descent on l to train the prediction model. The network parameters are adjusted until the training termination condition is met, so that a predictive model sensitive to future dangerous states can be obtained.

[0041] Steps c and d involve collecting hazardous data on past accidents caused by irrational resource allocation in the smart grid, and training a hazardous prediction model sensitive to hazardous resource allocation actions. This allows for the spontaneous prediction of future state distributions when allocating resources in a smart grid.

[0042] e performs a Q-value evaluation, using the Q-value to assess whether the current resource allocation is reasonable.

[0043] In reinforcement learning, the Q-value can be used to determine the long-term reward after performing an action in the current state. Applying this idea to the smart grid, where the distribution state of power devices in region i at time j is... Perform resource allocation actions as use The Q-value indicates the quality of resource allocation. Using the Q-value, we can clearly understand which resource allocation will yield the greatest return in the current state, providing a foundation for finding the most efficient resource allocation method.

[0044] f uses Q-value bootstrapping to find the optimal resource allocation scheme.

[0045] During the continuous resource allocation in a smart grid, the Q-value is updated using the Q-value formula to obtain a new evaluation function. After the Q-value is updated, an ε-greedy strategy is selected based on the new Q-value, choosing the most efficient resource allocation action according to the greedy strategy.

[0046] Steps e and f use the Q value to judge the quality of resource allocation and select the action that maximizes the Q value as the resource allocation.

[0047] Steps e and f utilize the Q-value in reinforcement learning to evaluate the quality of the action performed in the current state. Indicates the current state Execute resource allocation actions The degree of good or bad. Each time the agent interacts with the environment, iteratively calculates the Q-value to better adapt to the resource allocation of the smart grid. The update formula is: In the formula, α is the preset learning rate, typically a real number between 0 and 1. The larger α is, the greater the proportion of results obtained by adopting new attempts and the smaller the proportion of results obtained by retaining old ones. γ is the discount rate, typically a real number between 0 and 1. The larger γ is, the greater the weight of the prediction effect in the allocation decision. For smart grids in predictive state The rewards for different resource allocation actions.

[0048] By employing an ε-greedy strategy in steps e and f, the resource allocation action that maximizes allocation efficiency can be selected while simultaneously considering exploration. Furthermore, the Q-value obtained through continuous interaction and iteration with the environment can adapt well to the dynamic resource changes of the smart grid, enabling it to select the best resource allocation action and maximize resource utilization efficiency.

[0049] g utilizes a risk prediction model to avoid risky resource allocation actions.

[0050] While using Q-values ​​for ε-greedy policy selection can, to some extent, lead to more rational resource allocation in smart grids, risky resource allocation still exists, potentially causing localized network outages. To address this issue, a pre-trained risk prediction model is used. This method predicts the future state of resource allocation actions selected by the ε-greedy strategy under the current power equipment state distribution. The predicted future state is then compared with the historical accident information state distribution. If the future state is dangerous, the current resource allocation action is not executed; instead, the action with the highest Q value is executed. If the future state is safe, the current resource allocation action is executed.

[0051] Step g mainly utilizes the hazard prediction model trained in steps c and d. It predicts whether the current resource allocation action is reasonable, and if not, it will automatically choose other safe resource allocation actions.

[0052] Although the ε-greedy strategy and Q-value enable smart grids to self-organize resource allocation and improve the overall resource utilization efficiency of the power system, unsafe resource allocation still exists in smart grids. Therefore, a hazard prediction model is introduced to avoid unsafe resource allocation actions. When the distribution status of power equipment and maintenance equipment in region i at time j is extracted and entered into the latent space, its state is: At this point, the ε-greedy strategy is used to select resource allocation actions. By training the prediction model Predict the state z at the next time step. j+1 The formula is: At the same time, the collection of hazard status information for region i The hidden state set after feature extraction is obtained by performing feature extraction in step b. Predicting the future state The set of hidden states corresponding to the set of dangerous state information Similarity comparison is performed, with the comparison function being the future state. The mean squared error loss between each state sample in Z: See Figure 4 ,if If the value is less than the threshold δ, then the current resource allocation action will be executed. This is dangerous; the smart grid will abandon this resource allocation action and instead execute the resource allocation action that maximizes the Q value. If L i If the value is greater than the threshold δ, it means that the current resource allocation action will be executed. It is safe, and the smart grid will allow the current resource allocation actions to proceed. After each node makes a decision, the data is transmitted to the central node's global server. The global server further checks for safety via the network and outputs the final decision set. The overall architecture can be found in [reference needed]. Figure 2 .

[0053] This invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the aforementioned integrated operation and maintenance resource allocation method. The concept of this device embodiment is the same as the working process of the integrated operation and maintenance resource allocation method in the above embodiments. Therefore, the entire content of the above-described integrated operation and maintenance resource allocation method embodiments is incorporated into this device embodiment by means of full reference, and will not be repeated here.

[0054] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned integrated operation and maintenance resource allocation method. The concept of this storage medium embodiment is the same as the working process of the integrated operation and maintenance resource allocation method in the above embodiments. Therefore, the entire content of the above-described integrated operation and maintenance resource allocation method embodiment is incorporated into this storage medium embodiment by means of full reference, and will not be repeated here.

[0055] The integrated operation and maintenance resource allocation method provided by this invention utilizes a hazard prediction model for pre-allocation, reducing the number of hazardous resource allocation schemes. Simultaneously, it leverages Q-learning in reinforcement learning to find the optimal resource allocation for the smart grid. This invention seeks the optimal resource allocation method while ensuring safe resource allocation. It uses historical hazardous allocation information as data to train the hazard prediction model, pre-allocating resources through the hazard prediction model before actual allocation. If a hazardous allocation is detected, alternative allocations are implemented; otherwise, the current allocation scheme continues.

[0056] The above description is merely a preferred embodiment of the present invention and does not limit its patent scope. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, whether directly or indirectly applied to other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A method for integrated operation and maintenance resource allocation in a power information network, characterized in that, Includes the following steps: An environmental modeling process is used to model the allocation of operation and maintenance resources in a smart grid, resulting in a deep learning model. The distribution of power equipment and operation and maintenance equipment within the smart grid is considered as the state, and the allocation of these resources to various regions is considered as the action. In this deep learning model, the Q-value function from reinforcement learning is used to obtain the rewards for different resource allocation actions under a given state. Based on this, further... A strategy is selected to obtain the most efficient resource allocation action; among which, in In the strategy, The larger the value of , the greater the probability of using a random action as the most efficient resource allocation action, and the smaller the probability of using the action with the largest current Q function value as the most efficient resource allocation action. A hazard prediction model is introduced, collecting historical information on past smart grid allocation accidents. This historical information includes a set of hazard state information and a set of hazard action information. The hazard prediction model is trained using this historical information. The training process includes: extracting features and flattening the hazard state information to obtain the corresponding hidden states; inputting the hidden states and corresponding hazard action information into the prediction network of the hazard prediction model to obtain the predicted next hidden state; extracting features and flattening the actual next hidden state to obtain the actual next hidden state; calculating the mean squared error loss between the predicted and actual next hidden states; and applying gradient descent to the variance loss to train and update the network parameters of the hazard prediction model. After training, a certain state and its corresponding optimal resource allocation action are input into the hazard prediction model to obtain the predicted future state. The predicted future state is compared with the states in the hazard state information set. If the preset safety requirements are met, the optimal resource allocation action is executed; otherwise, other resource allocation actions are executed. If other resource allocation actions are performed, they also need to undergo the safety check of the hazard prediction model. The state and the corresponding selected other resource allocation actions are input into the hazard prediction model to obtain the corresponding predicted future state. The predicted future state is then compared with the state in the hazard state information set. If the preset safety requirements are met, the corresponding resource allocation action is executed; otherwise, a different resource allocation action needs to be executed. Before the resource allocation action is executed, the safety requirements of the hazard prediction model must be met at least once.

2. The integrated operation and maintenance resource allocation method according to claim 1, characterized in that, The state of region i at time j and the corresponding optimal resource allocation action are input into the hazard prediction model to obtain the predicted next hidden state. The set of hazard state information corresponding to region i is subjected to feature extraction and flattening to obtain the corresponding hidden state set Z; the next predicted hidden state is then... The mean squared error loss is calculated with the set of hidden states Z. If the mean squared error loss exceeds a preset threshold, the most efficient resource allocation action is executed. Otherwise, the resource allocation action that maximizes the Q value is executed instead.

3. The integrated operation and maintenance resource allocation method according to claim 1, characterized in that, In the deep learning model, the agent needs to iteratively calculate the Q value every time it interacts with the environment in order to adapt to the resource allocation of the smart grid.

4. The integrated operation and maintenance resource allocation method according to claim 2, characterized in that, The set of hazard status information corresponding to region i is denoted as j= A three-layer convolutional neural network is used to process the danger state information corresponding to region i at time j. Feature extraction is performed, followed by flattening, and then compression is achieved using a network consisting of two fully connected layers to obtain the corresponding hidden states. .

5. The integrated operation and maintenance resource allocation method according to claim 4, characterized in that, From the set of hazard status information corresponding to region i Extract the state at time j+1. , for state The hidden state is obtained after feature extraction, flattening, and compression. .

6. The integrated operation and maintenance resource allocation method according to claim 1, characterized in that, Historical information on past smart grid distribution accidents is regularly supplemented and updated, and the hazard prediction model is supplemented and updated accordingly.

7. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the integrated operation and maintenance resource allocation method according to any one of claims 1 to 6.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the integrated operation and maintenance resource allocation method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Power communication network utility maximization resource allocation policy generation method based on Q-learning

    CN108092804A

  • Power information equipment safety monitoring method

    CN112418398A