DQN-based icing disposal decision optimization method and system
Through the DQN-based decision optimization method for ice-covering treatment, the problems of low decision efficiency and poor adaptability of traditional ice-covering treatment methods in extreme weather are solved, and the intelligence and systematization of ice-covering decisions are realized, and the accuracy and real-timeness of decisions are improved.
Patent Information
- Application Number
- CN202411964951.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-13
AI Technical Summary
Traditional ice-covering treatment methods have problems such as strong subjectivity, low decision-making efficiency, and difficulty in adapting to complex environments, and cannot meet the decision-making requirements in extreme weather.
The DQN-based ice-covering treatment decision optimization method is adopted, and the optimal ice-covering treatment decision is made based on the real-time environmental status by establishing multi-objective deicing interrupt optimization problems, using deep reinforcement learning network learning decisions, and making the best ice-covering treatment decisions based on the real-time environmental status.
The systematization and intelligence of ice-covered decisions have been achieved, the accuracy and real-time nature of decisions have been improved, the risk of damage to transmission lines by extreme weather has been reduced, and the errors and limitations of human decision-making have been avoided.
Smart Images

Figure CN119989870A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power transmission system maintenance, and in particular to a DQN-based icing disposal decision optimization method and system. Background Art
[0002] Icing is extremely harmful to the power grid. Traditional icing decision-making methods are usually based on experience and rules, which are highly subjective, inefficient, and difficult to adapt to complex environments. They cannot meet the decision-making requirements of overhead transmission lines in extreme weather conditions that take into account both human and equipment limitations. Therefore, a more intelligent and accurate icing decision-making method is needed to improve efficiency and quality.
[0003] This patent first establishes a multi-objective de-icing interruption optimization problem, taking the number of manpower, de-icing equipment and risk loss as measurement criteria. By modeling the environmental status, behavior and rewards in the de-icing process, an environment for intelligent agent training is provided. The DQN algorithm is then used to train the deep reinforcement learning network to learn decisions, so that it can make the best de-icing decision according to the real-time environmental status, improve the efficiency and quality of de-icing, and finally realize that the transmission line can continue to work at the lowest cost, ensuring the stability and safety of power supply. Summary of the invention
[0004] In view of the above-mentioned problems, the present invention is proposed.
[0005] Therefore, the technical problem solved by the present invention is: the traditional methods of dealing with ice accumulation mainly include thermal ice melting, mechanical ice breaking and manual deicing, etc., and these methods have their own advantages and disadvantages. Although thermal ice melting can effectively eliminate ice and snow accumulation, it consumes a lot of energy and has high requirements for equipment; mechanical ice breaking requires a lot of manual intervention, which is inefficient and has high risks; although manual deicing is flexible, it is difficult to cover all lines due to limited human resources. In response to these problems, in recent years, intelligent deicing systems that combine big data analysis and artificial intelligence technology have gradually attracted the attention of researchers.
[0006] Existing intelligent deicing decision-making methods mostly use traditional rule-based models or some simple machine learning algorithms to make deicing predictions and decisions. However, these methods often fail to fully consider complex meteorological factors and the physical characteristics of transmission lines, resulting in low accuracy and practicality of decisions. Specifically, existing technologies usually ignore the precise modeling of risk assessment and reward functions, and fail to effectively quantify the costs and effects of various deicing methods. Therefore, existing technologies are difficult to cope with complex icing environments and decision-making needs under multiple constraints.
[0007] In order to solve the above technical problems, the present invention provides the following technical solutions: a DQN-based ice disposal decision optimization method, comprising: selecting a first task mode for a first object and modeling the first task problem.
[0008] A first comprehensive mathematical model and a first comprehensive function are established to obtain an environment suitable for the first task.
[0009] Use the first algorithm to calculate and learn the optimal decision.
[0010] As a preferred solution of the DQN-based icing disposal decision optimization method described in the present invention, wherein: the first task mode is a different working mode when performing a task on the first object.
[0011] As a preferred solution of the DQN-based icing disposal decision optimization method described in the present invention, wherein: the modeling of the first task problem includes first confirming the problem to be solved for the first task, modeling and solving the problem, and obtaining a decision solution for the first task problem.
[0012] As a preferred solution of the DQN-based icing disposal decision optimization method described in the present invention, wherein: the first comprehensive mathematical model includes a calculation model of the first task and a calculation model of the first object.
[0013] As a preferred solution of the DQN-based icing disposal decision optimization method described in the present invention, the first comprehensive function includes establishing a reward function based on the first task mode and setting rewards and penalties for the first task.
[0014] As a preferred solution of the DQN-based icing disposal decision optimization method described in the present invention, wherein: the first object includes but is not limited to a transmission line, and the first task includes but is not limited to de-icing the first object.
[0015] The first task method includes but is not limited to thermal ice melting, mechanical ice breaking and other methods.
[0016] The first comprehensive mathematical model includes an icing growth model based on meteorological conditions and a risk assessment model based on meteorological, icing and transmission line characteristics.
[0017] The first comprehensive function is a reward function based on risk, labor, and equipment limitations.
[0018] The first algorithm includes, but is not limited to, the DQN algorithm.
[0019] As a preferred solution of the DQN-based icing disposal decision optimization method described in the present invention, wherein: the use of the first algorithm to calculate and learn the optimal decision includes the intelligent agent using the DQN algorithm for training and iterative learning.
[0020] A DQN-based icing treatment decision optimization system, characterized by: including:
[0021] The modeling module selects a first task mode for the first object and models the first task problem.
[0022] The calculation module establishes a first comprehensive mathematical model and a first comprehensive function to obtain an environment suitable for the first task and uses the first algorithm to calculate and learn the optimal decision.
[0023] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0024] A computer-readable storage medium stores a computer program, which implements the steps of the method described above when executed by a processor.
[0025] The beneficial effects of the present invention are as follows: the systematization and intelligence of icing decision-making is achieved, the interference of human factors in decision-making is reduced, and the accuracy and real-time performance of decision-making are improved. The risk of damage to power transmission lines caused by extreme weather is effectively reduced. Automatic optimization of de-icing decisions reduces the risks and waste of resources caused by excessive or delayed de-icing. Through the DQN algorithm, the decision-making strategy can be automatically adjusted to gradually improve the effect of the decision, so that the system can still maintain a high decision-making accuracy rate when facing changeable weather and resource constraints, avoiding the errors and limitations of human decision-making, and greatly improving the system's adaptability and generalization capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. Among them:
[0027] Figure 1 An overall flow chart of an ice disposal decision optimization method and system based on DQN is provided for the first embodiment of the present invention.
[0028] Figure 2 A Markov decision process diagram of a DQN-based icing disposal decision optimization method and system provided in the first embodiment of the present invention.
[0029] Figure 3 A flowchart of a DQN-based icing disposal decision optimization method and a system risk assessment model provided in the first embodiment of the present invention.
[0030] Figure 4 A reward function diagram of a DQN-based ice disposal decision optimization method and system provided in the first embodiment of the present invention.
[0031] Figure 5 A DQN-based ice disposal decision optimization method and system agent and action interaction diagram is provided for the first embodiment of the present invention. DETAILED DESCRIPTION
[0032] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in the art without creative work should fall within the scope of protection of the present invention.
[0033] Example 1, reference Figure 1 to Figure 5 , which is an embodiment of the present invention, provides an icing disposal decision optimization method based on DQN, including:
[0034] S1: Select a first task method for a first object and model the first task problem.
[0035] In the present invention, the first object is a power transmission line, the first task is deicing, and the first task method is thermal deicing, mechanical deicing, and other deicing methods. Direct current deicing of thermal deicing and manual operation of mechanical deicing are used as decision actions of the present invention.
[0036] In the present invention, the first task includes but is not limited to de-icing of a power transmission line, and may also be power transmission line fault elimination and power transmission line load optimization.
[0037] In an optional embodiment of the present invention, the first task is to troubleshoot the transmission line. In some cases, the transmission line may have local failures or unstable operation due to icing or other factors. The task is to quickly detect and repair the failure to restore power supply.
[0038] The modeling of the first task problem includes:
[0039] Task problem: This task involves selecting the most appropriate repair method (such as adjusting load, switching backup circuits, repairing fault points) according to the fault type (such as line break, ground fault, etc.) and location when a transmission line fails. By establishing a fault detection and response model, ensure that the power outage time and losses are minimized.
[0040] Modeling method: Using real-time monitoring data of transmission lines (such as current, voltage, temperature, fault location, etc.) and historical fault data, an intelligent decision-making system based on the DQN algorithm is established. By training the intelligent agent, the optimal repair decision made under different fault conditions is learned to improve the repair efficiency and minimize economic losses.
[0041] In an optional embodiment of the present invention, the first task is to optimize the load of the transmission line. When the transmission line is affected by icing, it may be necessary to reasonably distribute the load according to the current line condition to reduce the risk of line overload.
[0042] The modeling of the first task problem includes:
[0043] Task problem: The load optimization task of the transmission line requires reasonable adjustment of the load of each line during the icing period to ensure that the power supply capacity of the system is maximized without failure. A decision model is established by evaluating the real-time load, meteorological conditions and health status of the line.
[0044] Modeling method: Use the DQN algorithm to build a load optimization model based on real-time data of the line (such as current, voltage, temperature, load, etc.). Optimize the load distribution strategy by training the intelligent agent to ensure that the power grid can operate smoothly and reduce losses under the influence of icing.
[0045] In this invention, the first task is to solve the problem of n transmission lines under different conditions at different locations. Through the DQN algorithm, the current and future risks and costs are evaluated to make a decision action set (natural ice melting, artificial deicing, DC ice melting) to maximize the reward. Among them, the factors affecting the formation of ice cover are: meteorological and transmission line characteristics, and the decision constraints are: economic cost, equipment quantity limit, and labor quantity limit.
[0046] The problem is modeled as a Markov decision process (MDP), including the environment, agent, action, reward and state, such as Figure 2 shown.
[0047] It should be noted that there is currently no open environment that can be used for ice melting decision-making. Therefore, this topic proposes an environment design idea for intelligent agent learning.
[0048] The solution process includes monitoring the weather and transmission line status, establishing an ice melting decision-making environment model based on gym, determining the state space set, action space set, and reward function of the environment, determining the DQN algorithm to construct a transmission line decision plan in a high-dimensional continuous space, exploring the intelligent agent learning model, conducting experiments through simulated data, completing the model training, and verifying the effectiveness of the model.
[0049] It should be noted that the first object includes but is not limited to a power transmission line, and may also be a wind turbine generator set and a railway signal system.
[0050] In an optional embodiment of the present invention, the first object is a wind turbine generator set, and the first task is to de-ice the wind turbine generator set. Wind turbine generator sets are prone to ice accumulation in extremely low temperature environments, which affects the normal operation of the unit. In this case, the task is to de-ice the blades of the wind turbine generator set to ensure its normal operation.
[0051] The modeling of the first task problem includes:
[0052] Task problem: For the de-icing task of wind turbines, multiple factors need to be considered, such as wind speed, temperature, humidity, and the structure of wind turbine blades. When modeling, it is necessary to conduct a comprehensive analysis of the current state of the unit (such as the ice thickness of the blades, meteorological conditions, and the operating status of the wind turbine), and evaluate the effects of different de-icing methods (such as hot air de-icing, mechanical de-icing, and electric heating de-icing) on the unit's restoration of normal operation.
[0053] Modeling method: Establish a multi-objective optimization model based on the health status of the unit (such as whether the blades are frozen, the thickness of the ice, etc.), meteorological conditions (such as temperature, wind speed, humidity, etc.) and unit performance data (such as power generation, speed, etc.). Use the DQN algorithm to train the agent to learn to take the best deicing measures under different meteorological and working conditions, maximize the stable operation time of the generator set, and minimize equipment damage and manual intervention costs.
[0054] In an optional embodiment of the present invention, the first object is a railway signal system, and the first task is to de-ice the railway signal system. In cold weather conditions, sensors, signal lights and other equipment of the railway signal system may not work properly due to ice, affecting the safety of the railway. Therefore, the task is to de-ice the railway signal system.
[0055] The modeling of the first task problem includes:
[0056] Task problem: This task requires selecting different deicing methods (such as hot air deicing, electric deicing, etc.) according to the freezing conditions of different parts of the railway signal system (such as signal lights, sensors, communication equipment, etc.). The thickness of the ice layer, the operating temperature range of the equipment, and the possible types of failures need to be considered when modeling.
[0057] Modeling method: A model based on the health status of railway signal system equipment, meteorological conditions and equipment layout is established. Through DQN algorithm training, the agent learns to choose different deicing strategies under different environmental conditions to avoid equipment failure and reduce human intervention. The reward function considers factors such as equipment recovery function, reduced failure risk and reduced deicing cost.
[0058] S2: Establish a first comprehensive mathematical model and a first comprehensive function to obtain an environment suitable for the first task.
[0059] In the present invention, the first comprehensive mathematical model includes an icing growth model based on meteorological conditions and a risk prediction model based on meteorological, icing and transmission line characteristics, and the first comprehensive function is a reward function based on risk, labor and equipment limitations.
[0060] It should be noted that the ice growth model can be established based on the material and surface characteristics of the first object, namely the transmission line, in addition to meteorological conditions. The risk assessment model can also be established based on the load of the power grid, the health status of the transmission line, the aging degree of the equipment, and the ability of the operator, in addition to meteorological, icing and transmission line characteristics.
[0061] In an optional embodiment of the present invention, an ice growth model based on the material and surface characteristics (such as material, shape, etc.) of the transmission line is established. Taking into account the different effects of transmission lines of different materials on ice accumulation, the model will further adjust the speed and distribution pattern of ice growth by analyzing the effects of thermal conductivity, hygroscopicity and shape of different surface materials on ice accumulation. Transmission lines of different materials and shapes will cause ice to form at different speeds at different locations, and may have different ice thickness distributions.
[0062] In an optional embodiment of the present invention, a systematic risk assessment is performed based on a multi-factor environment. In addition to considering a single icing factor, a comprehensive risk assessment is also conducted in combination with factors such as the load of the power grid, the health of the transmission lines, the aging of the equipment, and the capabilities of the operators. The model conducts a comprehensive monitoring of the transmission system and combines it with real-time data streams to assess the various risks that the transmission lines in a specific area may encounter in the future. This comprehensive risk model emphasizes the interactive impact between different factors within the system, helping to achieve a more comprehensive fault prevention and handling solution.
[0063] In the present invention, temperature, humidity, wind speed and rainfall sensors are installed on the transmission lines to obtain meteorological data in real time. Historical data can also be used for training and then transplanted into reality. The present invention uses the public data sets of the European Centre for Medium-Range Weather Forecasts (ECMWF): ERA5 hourly data on pressure levels from 1940 to present and ERA5 hourly data on single levels from 1940 to present for training.
[0064] The temperature unit of the dataset was converted from Kelvin to Celsius. The wind direction was converted into the angle relative to the north direction based on u10 (the component of the wind in the east-west direction) and v10 (the component of the wind in the north-south direction), and was converted from radians to degrees. Data from early December to the end of March were selected as environmental simulation parameters because there would be no possibility of ice cover in other time periods and they did not need to be used as training data.
[0065] The ice growth model is used to calculate the ice thickness. However, in the absence of publicly available ice data, the present invention selects a mechanism model to calculate the ice thickness. The environment simulates the rime coverage process. Rime has a large density and is more destructive. Therefore, the ice growth model is divided into the conditions of rainfall and no rainfall.
[0066] When the rainfall is less than 0.0000005, the Imai model is used, and the formula is as follows:
[0067]
[0068] R eq =R-R0
[0069] Among them, R is the radius of the wire after ice covering, v is the wind speed, T is the temperature, t is the ice covering time, R eq is the radius of the transmission line, and R0 is the original radius of the transmission line.
[0070] When the rainfall is greater than 0.0000005, the Jones model is used, and the formula is as follows:
[0071]
[0072] According to the empirical formula proposed by Best, the liquid water content w is calculated by the rainfall P, R eq is the radius of the transmission line, ρ0 is the density of liquid water, generally 1g / cm 3 , ρ i The ice density of rime is generally 0.9g / cm 3
[0073] By judging the rainfall conditions, the thickness of ice covering the transmission lines can be calculated.
[0074] Collect the transmission line model and related parameters. Transmission line breakage and tower collapse are generally caused by extreme weather. In addition to considering the above-mentioned ice thickness, the characteristics of the transmission line itself must also be considered. The risk estimation model needs to use the collected transmission line model, outer diameter, resistance at 0°C, weight, elastic coefficient, expansion coefficient, maximum breaking force and safety factor.
[0075] Calculate the self-weight load:
[0076]
[0077] Calculation of ice loads:
[0078] W=10 -6 ×ρ i πg n R eq (R eq +D)L t
[0079] Calculate wind loads:
[0080]
[0081] To calculate the rain load, first obtain the raindrop diameter range based on the rainfall, as shown in Table 1.
[0082] Table 1 Raindrop diameter range
[0083]
[0084] The raindrop diameters are divided into 0-1mm, 1-3mm and 3-6mm, and the final speed and number of raindrops are calculated based on the raindrop diameters.
[0085] Count the number of raindrops:
[0086]
[0087] Calculate the final velocity of raindrops. When the raindrop diameter is ≤1, the final velocity of raindrops is calculated using the modified Sha Yuqing formula:
[0088] x=[28.32+6.524lg(0.1D r )-(lg0.1D r ) 2 ] 0.5 -3.665
[0089] v = 0.496 × 10 x
[0090] When 1<raindrop diameter≤3, the final speed of raindrop falling adopts Newton's formula:
[0091] v=(17.20-0.844D r )(0.1D r ) 0.5
[0092] When the raindrop diameter is greater than 3, the following formula is used:
[0093]
[0094] Calculate the force exerted on a transmission line by raindrops of different diameters
[0095]
[0096] Calculate the force exerted by all raindrops on the transmission line and convert F d Add them together to get F r .
[0097] Calculate the comprehensive load H and specific load γ:
[0098]
[0099] γ=4H / πD 2 L t
[0100] Calculation of the maximum stress σ per unit length of the transmission line cross-sectional area:
[0101]
[0102] Calculate the risk probability:
[0103]
[0104] The risk reward function, the probability multiplied by the loss of tower disconnection, can be set according to different situations.
[0105] Among them, L t is the line length, km. D is the transmission line radius, mm. w1 is the transmission line weight, kg / km. g n is the acceleration due to gravity, and its value is 9.8m / s 2 ρ w is the air density, g / cm 3 V1 is the horizontal wind speed, V2 is the wind speed perpendicular to the line direction, m / s. r is the diameter of the raindrop, mm. v is the final velocity of the raindrop, m / s.
[0106] The flow chart of the risk assessment model is as follows Figure 3 shown.
[0107] Furthermore, in the present invention, the first comprehensive function is a reward function based on risk, labor, and equipment limitations, and may also be a reward function based on system sustainability and optimization.
[0108] In an optional embodiment of the present invention, the first comprehensive function is a reward function based on system sustainability and optimization, in which the agent not only pays attention to the current risks and costs, but also considers the effects of several future decisions. Its main goal is to reduce dependence on resources and enhance the self-recovery ability of the system. For example, in some cases, the agent can choose not to completely de-ice, but to delay high-intensity intervention on equipment by controlling the load of the transmission line and monitoring the risks, thereby avoiding excessive consumption of resources.
[0109] This reward function pays more attention to long-term system stability and reliability, allowing the system to maintain a low failure rate and low operating costs in different ice-covering cycles. The learning process of the agent is not only for the current task, but also to enable the system to show higher efficiency and better fault prevention capabilities in future tasks. In this process, the reward value includes not only the benefits brought by the current decision, but also the impact of the agent's decision on the long-term health of the system.
[0110] In the present invention, the possible rewards and punishments in the environment are as follows:
[0111] The rewards and penalties for manual deicing include: if manual deicing is still chosen when there is already a shortage of manpower, reward-n; labor costs, calculated based on the ice thickness, number of people, and labor capacity, the required time, reward-m; and the loss caused by power outage, reward-deicing time × t.
[0112] The rewards and penalties for DC de-icing include: reward-n for choosing DC de-icing when the equipment is insufficient; electricity cost, calculated based on the de-icing current and de-icing time, assuming that all ice can be melted in one hour; reward-q for the loss caused by power outage; reward-t for the loss caused by power outage.
[0113] The risk loss caused by the state update after executing the action is reward-loss. For every 24 steps completed without tower collapse or disconnection, reward+x. The episode caused by tower collapse or disconnection ends, reward-y.
[0114] The reward function of the present invention is as follows: Figure 4 shown.
[0115] S3: Use the first algorithm to calculate and learn the optimal decision.
[0116] In the present invention, the first algorithm is the DQN algorithm.
[0117] It should be noted that the first algorithm includes but is not limited to the DQN algorithm, and may also be the PPO algorithm.
[0118] In an optional embodiment of the present invention, the first algorithm uses the PPO algorithm, in which the agent's strategy is approximated by a neural network (usually a fully connected network or a convolutional network). The strategy network accepts the current state as input and outputs a probability distribution to guide the agent on how to choose an action. Unlike the Q value calculation of DQN, PPO directly optimizes the strategy and uses a "probabilistic strategy", that is, for each state, the value output by the strategy network represents the probability of selecting each action in that state.
[0119] PPO uses a method called Clipped Objective Function to update the policy. The core idea of this method is to avoid excessive changes in the policy during the update process, thereby ensuring the stability of the training process. Specifically, PPO optimizes the policy by maximizing a clipped advantage function. The advantage function is used to measure the quality of the current policy relative to the baseline policy, which represents the advantage of choosing a certain action in a certain state.
[0120] In the PPO algorithm, the optimization objective function includes the "probability ratio" of the current strategy (i.e., the ratio of the new and old strategies), and by clipping this ratio (i.e., limiting it to a reasonable range), the strategy is prevented from updating too quickly, which leads to unstable learning. Usually, by setting a hyperparameter ε, the amplitude of the strategy update is limited so that the difference between the new and old strategies is not too large, which can improve the stability of training.
[0121] Similar to the experience replay mechanism of DQN, the PPO algorithm also relies on collecting experience from the environment, but PPO focuses more on experience replay based on time series. After the agent takes an action at each time step, it stores the state, action, reward, and next state. By sampling these data, PPO performs strategy optimization.
[0122] In PPO, the sampling process generates a batch of trajectory data through multiple interactions with the environment. Each trajectory includes a state, action, and reward sequence. After certain preprocessing, these data are used to estimate the state value and action advantage, and to calculate the objective function. The advantage of PPO is that it optimizes and updates through multiple epochs, so that the strategy can be effectively learned and improved within a certain time frame.
[0123] Another key part of the PPO algorithm is to use a value function to estimate the value of each state. The role of the value function is to predict the long-term cumulative reward that the agent can obtain in a certain state while following the current policy. In order to estimate the value of a state, PPO usually uses a value network, which is separate from the policy network. The value network takes the state as input and outputs the estimated value of the state.
[0124] During training, PPO updates the policy by calculating the advantage function (the difference between the actual reward and the estimated value). The advantage function reflects the improvement of a state-action pair relative to the average performance, which is crucial for policy updates.
[0125] During the training process, PPO uses multi-step updates and pruning strategies to further improve training efficiency. Specifically, PPO does not perform a single data update each time, but updates the strategy through multiple data accumulations. This helps stabilize the update and reduce the amount of computation required for each update.
[0126] During the pruning process, PPO uses a technique called "clipping" on the objective function to prevent the difference between the new and old strategies from being too large. When the strategy update is too large, it may lead to an unstable learning process. By comparing the probability ratio of the new and old strategies and setting a threshold, PPO can control the magnitude of the update, thereby avoiding strategy collapse caused by excessive updates.
[0127] The training process of PPO is an iterative process. Each time through interaction with the environment, data is accumulated, and the strategy and value function are updated until the strategy of the agent tends to be stable. During the training process, the agent continuously explores new strategies and behaviors, and maximizes the cumulative reward by taking the best action for each state.
[0128] When the agent's strategy reaches a convergence state after multiple training cycles, it means that the agent has learned how to make the best decision in a complex environment. At this point, the training process of the PPO algorithm ends, and the agent can choose the most appropriate ice treatment method based on the environmental state.
[0129] In the present invention, a neural network is constructed, and the neural network structures and parameters of TargetNetwork and OnlineNetwork are set to approximate the Q value function. Currently commonly used neural network structures include convolutional neural networks (CNN) and fully connected neural networks. An experience replay buffer is set. A loss function is defined, and the mean square error (MSE) loss is proposed.
[0130] The Online Network receives the current state as input and outputs the Q value Q(s,a) of each action. The DQN agent will choose the next action based on these Q values. The agent selects the optimal action with the largest Q value or randomly explores the action with an ε probability, executes the action, interacts with the environment, and observes the reward and the next state.
[0131] Agents interact with actions such as Figure 5 shown.
[0132] Experience replay, implement the experience replay mechanism, store the interaction experience (S, A, R, S') in the buffer. Randomly sample a batch of experience data from the experience cache, including state s, action a, reward r and next state s'. For all state sets s' of the batch, calculate the target Q value according to TargetNetwork, combine the cumulative discount return factor γ, and use the sum of the current reward r and γQ(s',a) as the target. The difference between the target and Q(s,a) is used as the TD error to update the estimate of the value function.
[0133] Calculate the loss function loss:
[0134]
[0135] Update the parameters of the Online Network through the back-propagation algorithm and optimize the loss function so that the current estimated Q value gradually approaches the target Q value. Regularly update the parameters of the Target Network and copy the parameters of the Online Network to the Target Network to stabilize the estimate of the target Q value. After multiple iterations and learning, adjust the action selection strategy according to the feedback and reward mechanism of the environment, and gradually find the best ice melting disposal strategy. Continuously interact with the environment, replay experience, and update network parameters until the set number of training times or convergence conditions are reached.
[0136] The computer device may be a server. The computer device includes a processor, a memory, an input / output interface (I / O for short) and a communication interface. The processor, the memory and the input / output interface are connected via a system bus, and the communication interface is connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data cluster data of the power monitoring system. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a DQN-based ice disposal decision optimization method is implemented.
[0137] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.
[0138] Embodiment 2 is an embodiment of the present invention, which provides an icing disposal decision optimization method and system based on DQN. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through simulation comparative experiments.
[0139] In this experiment, the goal is to verify an icing disposal decision optimization method based on the deep Q-learning (DQN) algorithm. For comparison, this embodiment also uses a traditional rule-based decision method as the prior art. The prior art mainly uses simple conditional rules and traditional physical models to conduct icing risk assessment and disposal decisions, which usually ignores the impact of environmental changes on decision-making and is inefficient when dealing with multiple variable interactions.
[0140] First, by analyzing meteorological conditions (temperature, humidity, wind speed, etc.), transmission line characteristics (outer diameter, elastic modulus, weight, etc.) and related cost factors (labor, power equipment, equipment maintenance, etc.), select the appropriate deicing method. In the prior art, the deicing method is usually determined by rules, such as selecting DC deicing when the temperature is lower than a certain value, and selecting mechanical deicing when the temperature is higher than a certain value; the present invention uses the DQN algorithm and combines the actual environmental data to automatically select the optimal deicing method (such as thermal deicing, mechanical deicing, DC deicing, etc.).
[0141] The present invention uses meteorological data (such as rainfall, wind speed, humidity, etc.) and transmission line characteristics (such as weight, outer diameter, breaking force, etc.) to establish an ice growth model. The risk assessment model calculates the load of the transmission line under different environmental conditions, and then predicts the risk of line breakage or tower collapse that the line may encounter. Compared with the static model of the prior art, the present invention provides a more accurate risk prediction through a model based on dynamic meteorological data.
[0142] In order to optimize the de-icing decision, a reward function is designed that comprehensively considers factors such as risk, labor, and equipment limitations. The reward function comprehensively considers the economic losses, equipment consumption, and labor costs caused by the de-icing operation. In the prior art, the reward function is usually too simple and fails to fully consider the weight differences of multiple factors. The reward function of the present invention is more sophisticated and can weigh various decision factors in actual operations.
[0143] Based on the environmental simulator, the present invention constructs a reinforcement learning environment suitable for transmission line icing decision-making. The environment can obtain meteorological data in real time and simulate the effects of different deicing methods to support the DQN agent to learn the optimal decision. In the prior art, environmental simulation often lacks the combination with actual meteorological data, so the effectiveness of the decision is limited.
[0144] The experimental results are shown in Table 2.
[0145] Table 2 Experimental results
[0146]
[0147]
[0148] By comparing the data in the table, it can be clearly seen that the DQN optimization decision-making method proposed in the present invention has significant advantages over the prior art in many aspects.
[0149] Under the same meteorological conditions (such as wind speed of 5m / s and temperature of -5℃), the strategy of the present invention can reduce the ice thickness (12mm for method 1, compared with 15mm for the prior art) and significantly reduce the risk of line disconnection (from 20% to 10% for the prior art). This optimization shows that the decision-making method based on DQN can more accurately predict the growth of ice and adjust the deicing method in time, thereby reducing the probability of accidents.
[0150] Through the optimized reward function, the present invention can not only reduce equipment consumption (for example, 3.0 kWh in method 1, compared with 5.5 kWh in the prior art), but also significantly reduce labor costs (150 yuan in method 1, compared with 200 yuan in the prior art). This shows that through intelligent decision-making, operation and maintenance costs can be reduced while ensuring safety.
[0151] Under different environments (such as wind speed 8m / s, temperature -7℃), the three decision-making methods of the present invention (thermal ice melting, DC ice melting, etc.) are more adaptable than the fixed rule methods of the prior art. For example, using the DC ice melting strategy (method 2), the equipment consumption is 6.0kWh, the labor cost is 180 yuan, and the risk of disconnection is reduced to 15%. Compared with the prior art, various indicators are significantly optimized, especially under extreme climate conditions, the performance is more stable and efficient.
[0152] Embodiment 3 is an embodiment of the present invention, including an icing treatment decision optimization system based on DQN, specifically:
[0153] A modeling module, selecting a first task mode for a first object and modeling the first task problem;
[0154] The computing module establishes a first comprehensive mathematical model and a first comprehensive function to obtain an environment suitable for the first task; and uses the first algorithm to calculate and learn the optimal decision.
[0155] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A DQN-based icing treatment decision optimization method, characterized in that: include: Selecting a first task method for a first object and modeling a first task problem; Establishing a first comprehensive mathematical model and a first comprehensive function to obtain an environment suitable for a first task; Use the first algorithm to calculate and learn the optimal decision.
2. The DQN-based icing treatment decision optimization method according to claim 1, characterized in that: The first task mode is a different working mode when performing a task on the first object.
3. The DQN-based icing treatment decision optimization method according to claim 2, characterized in that: The modeling of the first task problem includes firstly confirming the problem to be solved for the first task, modeling and solving the problem, and obtaining a decision solution for the first task problem.
4. The DQN-based icing treatment decision optimization method according to claim 3, characterized in that: The first integrated mathematical model includes a computational model of the first task and a computational model of the first object.
5. The DQN-based icing treatment decision optimization method according to claim 4, characterized in that: The first comprehensive function includes establishing a reward function based on the first task mode and setting rewards and penalties for the first task.
6. The DQN-based icing treatment decision optimization method according to claim 5, characterized in that: The first object includes but is not limited to a power transmission line, and the first task includes but is not limited to de-icing the first object; The first task method includes but is not limited to thermal ice melting, mechanical ice breaking and other methods; The first comprehensive mathematical model includes an ice growth model based on meteorological conditions and a risk estimation model based on meteorological, icing and transmission line characteristics; The first comprehensive function is a reward function based on risk, labor, and equipment limitations; The first algorithm includes, but is not limited to, the DQN algorithm.
7. The DQN-based icing treatment decision optimization method according to claim 6, characterized in that: The using of the first algorithm to calculate and learn the optimal decision includes the intelligent agent using the DQN algorithm for training and iterative learning.
8. A DQN-based icing treatment decision optimization system using the method according to any one of claims 1 to 7, characterized in that: A modeling module, selecting a first task mode for a first object and modeling the first task problem; A calculation module, establishing a first comprehensive mathematical model and a first comprehensive function, and obtaining an environment suitable for a first task; Use the first algorithm to calculate and learn the optimal decision.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Maintenance optimization method and system for heating tile in blister equipment
CN121073440A
Self-adaptive cooperative control method for operation of multi-type road deicing equipment
CN121578655A