Method for generating contingency plan for icing disposal of power transmission and distribution lines based on Q-Learning algorithm

The construction of an ice-covered disposal intelligent model through the Q-Learning algorithm solved the problem that the ice-covered disposal plan is inconvenient for simulation drills and dynamic changes, and realized intelligent decision-making support for ice-covered disposal, which improved the scientificity and efficiency of ice-covered disposal.

CN114548540BActive Publication Date: 2025-07-18NORTH CHINA ELECTRIC POWER UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210146606.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-17
Publication Date
2025-07-18
Estimated Expiration
2042-02-17

AI Technical Summary

Technical Problem

In the prior art, the ice-covered disposal plan is not convenient for simulation drills and dynamic changes, resulting in inefficient ice-covered disposal work.

Method used

The Q-Learning algorithm is used to build an agent learning model of the ice-covering treatment action plan. By establishing an environmental model of ice-covering growth and development, the agent is trained and the optimal processing strategy is generated, and output it to word documents and animations, providing decision support for ice-covering treatment.

Benefits of technology

The intelligent generation of ice-covering disposal plans has been achieved, the scientificity and efficiency of ice-covering disposal has been improved, simulation drills and dynamic adjustments have been supported, and economic and safety losses caused by ice disasters have been reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114548540B_ABST
    Figure CN114548540B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for generating a contingency plan for icing disposal of power transmission and distribution lines based on the Q-Learning algorithm, belonging to the technical field of power line deicing optimization. The steps are as follows: establish an environmental model for ice accretion growth, development, and disposal; based on the Q-Learning algorithm, construct an intelligent agent learning model for exploring the contingency plan for icing disposal and complete the training; select a specific environmental model as the environment corresponding to the contingency plan; apply the trained intelligent agent model to the contingency plan environment to calculate the optimal treatment strategy; output each change in the above environment and each measure taken to a word document of the contingency plan; construct an animated demonstration of the contingency plan for icing disposal to assist in training the disposal personnel. By constructing data related to the line state into an environment, the present invention uses reinforcement learning to complete decision optimization, constructs a learning model using the Q-Learning algorithm to complete the learning of the intelligent agent and the evaluation function; the intelligent agent interacts with it to output the corresponding contingency plan, and outputs the contingency plan in word format.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of power line deicing optimization, and particularly relates to a method for generating an icing disposal plan for transmission and distribution lines based on the Q-Learning algorithm. Background Art

[0002] With the development of computer technology and communication technology, various automation devices have been continuously installed in the power grid system, and they have collected a large amount of useful data, such as ice thickness, tower inclination, wire tension, environmental temperature, humidity, wind speed and direction, etc. At present, the work of predicting the weather and ice accretion development in the next few days based on these data has become a research hotspot. Directly giving corresponding ice accretion disposal plans based on these predictions will greatly improve the scientificity and efficiency of the disposal work. However, the plans for ice accretion disposal are mostly generated based on empirical rules, which are not convenient for simulation drills and dynamic changes. Therefore, there is an urgent need for a method with a simple model, strong description ability, easy to implement simulation and dynamic adjustment in the ice accretion disposal work, so as to utilize the relevant data and automatically generate ice accretion disposal plans under different scenarios. Summary of the Invention

[0003] The purpose of the present invention is to provide a method for generating an icing disposal plan for transmission and distribution lines based on the Q-Learning algorithm, aiming to solve the technical problem that the existing icing disposal plan is not convenient for simulation drills and dynamic changes.

[0004] To solve the above technical problems, the technical solution adopted by the present invention is:

[0005] A method for generating an icing disposal plan for transmission and distribution lines based on the Q-Learning algorithm, comprising the following steps:

[0006] Step 1: Establish an environmental model for ice accretion growth, development and disposal based on the OpenAI gym library;

[0007] Step 2: Based on the Q-Learning algorithm, construct an intelligent agent learning model for exploring the icing disposal action plan and complete the training of the model;

[0008] Step 3: Based on the environmental model in Step 1, select a specific environmental model as the environment corresponding to the plan;

[0009] Step 4: Apply the trained intelligent agent model in Step 2 to the plan environment in Step 3 to calculate the optimal treatment strategy;

[0010] Step 5: Based on python-docx, output each step change of the environment and each step measure taken in Step 4 to a word document of the disposal plan;

[0011] Step 6: Build a demonstration animation for the ice-covered disposal plan based on the gym library to assist in training disposal personnel.

[0012] Further, the said Step 1 includes the following steps:

[0013] Step 11: Build a static information model of the line

[0014] (1) Build the type and connection model of the tower and the line;

[0015] (2) Build the tolerance tension model of the line;

[0016] (3) Build the micro-environment meteorological model and the conductor surface temperature model of the line;

[0017] Step 12: Build a discretization model of the dynamic operation data of the line

[0018] (1) Build the historical ice-covered thickness and fault information models such as pole collapse and wire breakage of the line;

[0019] (2) Build the ice-covered monitoring data input model of the line;

[0020] (3) Build the micro-meteorological information input model of the line;

[0021] Step 13: Build a single-step interaction model between the disposal behavior and the environmental change;

[0022] Step 14: Build a human and financial upper limit model for ice-covered disposal to restrict the rewards for ice-covered disposal behaviors;

[0023] Step 15: Build a reward method for the environment to the disposal behavior in order to train the optimal disposal strategy.

[0024] Further, the said Step 2 includes the following steps:

[0025] Step 21: Initialize the set of ice-covered disposal behaviors;

[0026] Step 22: Build the initial Q-table of the environment;

[0027] Step 23: Adopt the time difference algorithm, try to make a random disposal behavior, and update the Q-table in a loop.

[0028] Further, the said Step 3 includes the following steps:

[0029] Step 31: Consider some extreme situations such as continuous rain and snow weather, combinations of temperature and humidity, and generate their combinations;

[0030] Step 32: Combine the above situations with the static information model of the line to generate different emergency disposal scenarios, that is, some determined environmental models.

[0031] Further, step 4 includes the following steps:

[0032] Step 41: Select the actions in each state from the Q-table using the greedy strategy;

[0033] Step 42: Record the complete state-action sequence to obtain the optimal strategy.

[0034] Further, step 5 includes the following steps:

[0035] Step 51: Install the python-docx plugin;

[0036] Step 52: Output the sequence recorded in step 42 to a word document according to the specified format to obtain the emergency response plan document.

[0037] Further, step 6 includes the following steps:

[0038] Step 61: The user specifies the possible environmental conditions that may occur;

[0039] Step 62: The system generates the corresponding deterministic environment according to the environmental changes specified by the user;

[0040] Step 63: Select actions from the Q-table generated by training using the greedy strategy;

[0041] Step 64: Execute the selected actions in the environment, and call the render in the environment to draw the state for each execution to form a frame-by-frame demonstration animation.

[0042] The beneficial effects of adopting the above technical solutions are as follows: Compared with the prior art, the present invention constructs the data related to the line state into an environment, uses reinforcement learning to complete decision optimization, uses the Q-Learning algorithm to construct an intelligent agent learning model for exploring the ice-covering disposal action plan. On the basis of the learning of the environment and the intelligent agent, given the scenario combination, the learning of the intelligent agent and the evaluation function is completed to generate the value estimation Q-table; for each scenario, let the intelligent agent interact with it, and the output corresponding environment and action sequence is the disposal plan, and the plan is output in word format. The present invention regards the ice period disposal behavior as a processing process of an intelligent agent, models the ice-covered line and its state, and uses the method of reinforcement learning to preview the development trend of ice covering, providing help for the ice-covering disposal decision-making. This method has important practical significance and theoretical value. Description of the Drawings

[0043] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments.

[0044] Figure 1It is a flowchart of a method for generating an icing treatment plan for a power transmission and distribution line based on a Q-Learning algorithm provided by an embodiment of the present invention;

[0045] Figure 2 It is an ice treatment process based on reinforcement learning in an embodiment of the present invention;

[0046] Figure 3 This is the process of constructing the ice treatment environment model in the embodiment of the present invention. DETAILED DESCRIPTION

[0047] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0048] With the continuous development of deep learning technology, its application fields are becoming more and more extensive. Reinforcement learning has good environmental exploration capabilities and strategy optimization capabilities. Reinforcement learning has achieved great success in the fields of robot path planning and game theory. In recent years, reinforcement learning has also begun to be applied to power-related fields. Some literatures have proposed a regional power grid reactive voltage optimization control method based on reinforcement learning theory, a power grid emergency control strategy based on deep reinforcement learning, an interconnected power grid CPS self-correction control method based on reinforcement learning, a wind storage cooperation decision-making based on reinforcement learning methods, and a multi-yield decay equipment preventive maintenance strategy based on reinforcement learning. These studies are all beneficial attempts to apply reinforcement learning in power systems, and have achieved good results.

[0049] In view of the successful application of reinforcement learning algorithms in the fields of robots, game AI and power system control, it is a new attempt to regard ice-disposal behavior as an intelligent processing process, model ice-covered lines and states, and preview the development trend of ice-covering. The present invention uses reinforcement learning in ice-covering disposal, which helps to optimize ice-covering disposal plans and reduce economic losses and safety benefits caused by ice disasters. This method has important practical significance and theoretical value.

[0050] The present invention provides a method for generating an ice treatment plan for power transmission and distribution lines based on a Q-Learning algorithm. The specific process is as follows: Figure 1 As shown, the following steps are included:

[0051] Step 1: Establish an environmental model for ice growth, development, and disposal based on Open AI's gym library; specifically, the following steps are included:

[0052] Step 11: Build the static information model of the line;

[0053] Step 111: Build the model of the types of poles and towers and the line and their connections;

[0054] (1) Use a relational database table to store pole and tower information, including pole and tower numbers, locations, types, and terrain information;

[0055] (2) Use a relational database table to store the information of overhead conductors, including line type, length, service life, and load conditions of the line;

[0056] (3) Use a relational database table to store the connection relationship between poles and towers and the line.

[0057] Step 112: Build the tolerance tension model of the line;

[0058] (1) Use a relational database table to store the tension resistance information of each overhead conductor;

[0059] (2) Use a relational database table to store the tension resistance information of each pole and tower.

[0060] Step 12: Build the discretization model of the dynamic operation data of the line;

[0061] Step 121: Build the model of historical ice coating thickness and fault information such as pole collapse and wire breakage of the line;

[0062] (1) Import information such as ice coating thickness, pole collapse, and wire breakage from historical faults;

[0063] (2) Mine the relationship between ice coating thickness and faults, and build an environmental model on this basis.

[0064] Step 122: Build the input model of ice coating monitoring data of the line;

[0065] Design a table for storing the change of dynamic ice coating thickness;

[0066] Store the ice coating thickness monitoring data in the table at a certain frequency (such as collecting a set of data every 15 minutes).

[0067] Step 123: Build the input model of micro-meteorological information of the line;

[0068] (1) Design the table structure for storing meteorological information and temperature information, including environmental temperature, humidity, wind speed, wind direction, and conductor surface temperature data;

[0069] (2) Store the meteorological and conductor temperature data in the table at a certain frequency (such as collecting a set of data every 15 minutes).

[0070] Step 13: Construct a single-step interaction model between the disposal behavior and environmental changes, as follows Figure 2 shown:

[0071] Step 131: Use the Python language to design a class for icing disposal to implement the construction of the environmental model; including states (discrete situations of the environment) and behaviors.

[0072] Step 132: Implement the step function in the class to change the environmental state according to the disposal behavior and calculate the reward value obtained by this behavior.

[0073] Step 14: Construct a model for the upper limits of human and financial resources for icing disposal to restrict the rewards for icing disposal behaviors.

[0074] (1) When calculating the behavior reward algorithm, use the upper limits of human and financial resources for icing disposal as reference factors.

[0075] (2) When the continuous behaviors of the agent lead to the upper limits of human and material resources, give a large penalty value (negative reward) to the behavior.

[0076] Step 15: Construct a reward method for the environment for disposal behaviors in order to train the optimal disposal strategy.

[0077] (1) Construct a reward function at the end of the icing period. Reward value = number of faults * (-100) + (1 - probability of fault) * 10 + remaining percentage * (10).

[0078] (2) According to the reward function, calculate the reward value based on the line state, number of faults, remaining human and financial resources at the end.

[0079] In summary, the construction process of the environmental model for icing disposal is as follows Figure 3 shown. Design the state model based on the inherent state and changing state, design the behavior model based on the line set and action set, and design the reward model according to the overall goal.

[0080] Step 2: Based on the Q-Learning algorithm, construct an intelligent agent learning model for exploring the icing disposal action plan and complete the training of the model; specifically including the following steps:

[0081] Step 21: Initialize the set of icing disposal behaviors.

[0082] Step 211: Construct a set of objects to be disposed of, with the entire line as the basic object of treatment. The size of the disposal set is the number of lines, and the set of lines is L.

[0083] Step 212: Construct a basic set of disposal actions including preventive maintenance, transformer de-icing, manual de-icing, no treatment, etc. This set is S.

[0084] Step 213: Combine the above two sets to form the action set A = L × S.

[0085] Step 22: Construct the initial Q-table for the environment; the construction process of the initial Q-table is as follows:

[0086]

[0087] Step 221: Construct the state set S of the environment according to the number of lines in the environment, the discrete combinations of meteorological data, and the discrete values of icing states;

[0088] Step 222: Initialize the estimated value corresponding to each state to 0 to obtain the initial Q-table.

[0089] Step 23: Adopt the temporal difference algorithm, try to make a random disposal behavior, and update the Q-table cyclically.

[0090] Step 231: Set an outer infinite loop, and each loop is a scene, that is, all time steps from the start of icing to the end of icing;

[0091] Step 232: Calculate and update the Q value for each state;

[0092] Step 233: Judge whether to terminate the outer loop according to the set increment threshold.

[0093] Step 3: Based on the environment model in Step 1, select a specific environment model as the environment corresponding to the pre-plan; specifically, it includes the following steps:

[0094] Step 31: Consider some extreme situations such as continuous rain and snow weather, combinations of temperature and humidity, and generate their combinations.

[0095] Step 311: Construct the discrete weather set W = {rain, snow, sunny, cloudy}, and the combined relationship with the duration (days) of the weather;

[0096] Step 312: Construct the discrete combinations of temperature and humidity, and their combined relationships with the duration;

[0097] Step 313: Construct the discrete set of icing thickness I = {zero, light, moderate, severe};

[0098] Step 314: Combine the above multi-step sets to generate the combined weather, that is, the icing situation;

[0099] Step 315: Select several typical situations as the typical situations for emergency disposal.

[0100] Step 32: Combine the above situations with the static information model of the line to generate different emergency disposal scenarios, that is, a set of determined environment models.

[0101] Step 321: Generate the states in the environment according to the environment model;

[0102] Step 322: Implement the states in the model using the Python language and implement the storage of the states in a file manner.

[0103] Step 4: Apply the trained agent in Step 2 to the pre-plan environment in Step 3 to calculate the optimal handling strategy; specifically including the following steps:

[0104] Step 41: Select the actions in each state from the Q-table using the greedy strategy;

[0105] Step 411: Search the learned Q-table to find the action a with the maximum value of the value of the action in the current state.

[0106] Step 421: Execute the action a in the environment and record the new state obtained after executing a.

[0107] Step 42: Record the complete state-action sequence to obtain the optimal strategy.

[0108] Step 5: Based on python-docx, output each step change of the environment and each step measure taken in Step 4 to the word document of the disposal pre-plan. Specifically including the following steps:

[0109] Step 51: Install the python-docx plugin;

[0110] Step 52: Output to the word document according to the sequence recorded in Step 42 in the specified format to obtain the emergency disposal pre-plan document.

[0111] Step 521: Select a combination relationship of the environment, that is, an emergency situation;

[0112] Step 522: In the emergency pre-plan generation program, output the state and action sequence relationship obtained in 42 to the word document.

[0113] Step 6: Build a demonstration animation for the ice-covered disposal pre-plan based on the gym library to provide assistance for training the disposal personnel. Specifically including the following steps:

[0114] Step 61: The user specifies the possible environmental situations;

[0115] Step 611: The user sets the values of various environments in the demonstration program;

[0116] Step 612: The user specifies the occurrence frequency and time length (days) of the set environmental parameters during the ice-covered period.

[0117] Step 62: The system generates a corresponding deterministic environment according to the environmental changes specified by the user.

[0118] Step 621: Generate environment-related parameters according to the environmental combination specified by the user.

[0119] Step 622: Design a new reward function according to the human and financial resources specified by the user and go online.

[0120] Step 623: Generate an environment class according to the above steps.

[0121] Step 63: Select an action from the Q-table generated by training using the greedy strategy.

[0122] Step 631: Start cycling from the first day of the icing period until the end of the icing period.

[0123] Step 632: Select the corresponding environmental state according to the time step (the time length of each scene) set by the user.

[0124] Step 633: According to the environmental state and the trained Q-table, select the behavior with the maximum value using the greedy strategy, and execute the step function to enter the next state.

[0125] Step 64: Execute the selected action in the environment, and call the render in the environment to draw the state for each execution, forming a frame-by-frame demonstration animation.

[0126] Step 641: Execute following the loop steps of Step 631.

[0127] Step 642: Use the graphics output class of python to complete the graphics drawing of each frame in the animation, and delay for a period of time so as to show the state change in the form of an animation.

[0128] In summary, a method for generating a contingency plan for icing disposal of power transmission and distribution lines based on the Q-Learning algorithm provided by the present invention can not only complete the modeling of the line state during icing and the modeling of the intelligent agent for exploring the optimization scheme, but also automatically generate a contingency plan according to the line state and the icing situation. Therefore, the present invention can provide guidance for the design of contingency plan software and is of great significance for improving the effectiveness of icing disposal work.

[0129] Many specific details are set forth in the above description to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed above.

Claims

1. A method for generating an ice covering disposal plan for power transmission and distribution lines based on the Q-Learning algorithm, characterized in that, It includes the following steps: Step 1: Establish an environmental model for ice accretion growth, development, and disposal based on the OpenAI gym library; it includes the following steps: Step 11: Build a static information model of the line (1) Build a model of the types and connections of poles and lines; (2) Build a model of the line's tolerance tension; (3) Build a microenvironment meteorological model and a conductor surface temperature model of the line; Step 12: Build a discretization model of the line's dynamic operation data (1) Build a model of historical ice accretion thickness and fault information such as pole collapse and wire breakage of the line; (2) Build an input model for the line's ice accretion monitoring data; (3) Build an input model for the line's micro-meteorological information; Step 13: Build a single-step interaction model of disposal behavior and environmental changes; Step 14: Build a model of the upper limits of human and financial resources for ice accretion disposal to constrain the rewards for ice accretion disposal behavior; Step 15: Build a reward method for the environment's disposal behavior to train the optimal disposal strategy; (1) Build a reward function at the end of the ice accretion period, reward value = number of faults * (-100) + (1 - probability of fault) * 10 + remaining percentage * (10); (2) Calculate the reward value according to the reward function and the line state, number of faults, remaining human and financial resources at the end; Step 2: Based on the Q-Learning algorithm, build an intelligent agent learning model for exploring the ice accretion disposal action plan and complete the training of the model; it includes the following steps: Step 21: Initialize the set of ice accretion disposal behaviors; Step 22: Build the initial Q-table of the environment; Step 23: Use the temporal difference algorithm to make random disposal behaviors and update the Q-table in a loop; Step 3: Based on the environmental model in Step 1, select a specific environmental model as the environment corresponding to the pre-plan; Step 4: Apply the trained intelligent agent model in Step 2 to the pre-plan environment in Step 3 to calculate the optimal processing strategy; Step 5: Based on python-docx, output each step change of the environment and each step measure taken in Step 4 to a word document of the disposal pre-plan; Step 6: Build a demonstration animation of the ice accretion disposal pre-plan based on the gym library to assist in training disposal personnel.

2. The method for generating an ice coating disposal plan for power transmission and distribution lines based on the Q-Learning algorithm according to claim 1, wherein The said Step 3 includes the following steps: Step 31: The specific environmental model is an extreme weather condition, including combinations of continuous rain and snow weather, temperature, and humidity; Step 32: Combine the above combination with the static information model of the line to generate different emergency disposal scenarios.

3. The method for generating an ice coating disposal plan for power transmission and distribution lines based on the Q-Learning algorithm according to claim 1, wherein The said Step 4 includes the following steps: Step 41: Select actions in each state from the Q-table using the greedy strategy; Step 42: Record the complete state-action sequence to obtain the optimal strategy.

4. The method for generating an ice covering disposal plan for power transmission and distribution lines based on the Q-Learning algorithm according to claim 3, characterized in that, The said Step 5 includes the following steps: Step 51: Install the python-docx plugin; Step 52: Output according to the sequence recorded in Step 42 to a word document in the specified format to obtain the emergency disposal pre-plan document.

5. The method for generating an ice covering treatment plan for power transmission and distribution lines based on the Q-Learning algorithm according to claim 4, wherein, The said Step 6 includes the following steps: Step 61: The user specifies the possible environmental conditions; Step 62: The system generates a corresponding deterministic environment according to the environmental changes specified by the user; Step 63: Select an action from the Q-table generated during training using the greedy strategy; Step 64: Execute the selected action in the environment. For each execution, call render in the environment to draw the state, forming a frame-by-frame demonstration animation.

Citation Information

Patent Citations

  • Coal mine accident simulating method and system based on multi-intelligent agent

    CN102508995A

  • Maximum power point tracking method and device of wind energy conversion system

    CN111222718A

  • Deep reinforcement learning emergency control strategy extraction method for power system

    CN114004282A