Intelligent method based on electric power communication and power grid fusion

Through the combination of LSTM, SAC and TD3 algorithms, the intelligent scheduling of power communication and power grid resources is realized, the real-time and coordination of resource scheduling in the existing system is solved, and the prediction accuracy and scheduling efficiency of the power communication network are improved.

CN120433162APending Publication Date: 2025-08-05SICHUAN SIJI TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510379894.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

In existing power communication systems, resource scheduling and allocation between the power grid and the communication network lacks real-time and flexibility, making it difficult to adapt to changes in dynamic environments, resulting in inefficiency and poor coordination.

Method used

The LSTM model is used for time series prediction, and the resource matching and scheduling optimization are combined with SAC and TD3 algorithms to realize intelligent decision-making and adaptive resource allocation.

Benefits of technology

It improves the prediction accuracy and flexibility and stability of the power communication network, and can quickly respond to emergencies in a dynamic environment and improve overall network performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120433162A_ABST
    Figure CN120433162A_ABST
Patent Text Reader

Abstract

The invention claims to protect an electric power communication and power grid fusion method based on deep reinforcement learning, and aims to solve the problems of insufficient intelligent level, low coordination efficiency and the like of an existing electric power communication network. The method comprises the following steps: firstly, performing time sequence prediction on a power load and a communication demand by using an LSTM neural network to provide a basis for subsequent resource matching and scheduling; then, carrying out preliminary matching on power grid resources by using an SAC algorithm, and automatically adjusting a power grid and communication resource allocation scheme; and then the efficiency and the stability of a resource scheduling strategy are further improved through a TD3 algorithm, the problem of unstable resource allocation caused by dynamic environment change is solved, intelligent decision making is performed according to different requirements of a power grid, the coordination between different tasks is ensured, and the overall network energy efficiency is optimized. Through the deep reinforcement learning algorithm, for complex tasks in the power communication network and the power grid, the problems of resource scheduling, operation and maintenance optimization and the like of the power communication network can be effectively solved, intelligent decision support is provided, and the stability, the reliability and the operation and maintenance efficiency of the power grid are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of power communication and power grid, and particularly relates to an intelligent method based on the integration of power communication and power grid. Background Art

[0002] The integration of power communication networks and power grid systems plays a vital role in modernizing power management and improving grid efficiency. Power grids and communication networks operate in a highly dynamic and complex environment, with frequent fluctuations in power loads and evolving communication demands. Traditional resource allocation methods often lack the flexibility and adaptability to handle these fluctuations in real time, limiting the system's ability to respond to dynamic demand and environmental conditions. In this context, deep reinforcement learning has emerged as a promising approach to enhance the decision-making processes for resource allocation, scheduling, and optimization within these systems. Therefore, it is essential to consider the design of intelligent management systems that can predict future demand, dynamically adjust allocation plans, and effectively optimize resource utilization.

[0003] In existing power communication systems, resource scheduling and allocation between the power grid and communication networks are typically handled through pre-set models or manual intervention, which are unresponsive to the rapid and unpredictable changes inherent in these systems. This static approach leads to inefficiencies and heavy manual workloads, such as power outages, communication delays, and suboptimal energy utilization. Furthermore, faced with environmental uncertainty and the need for real-time decision-making, traditional systems struggle to ensure stability, resulting in a lack of coordination between different network components.

[0004] The SAC algorithm is a model-free reinforcement learning algorithm that focuses on maximizing the balance between exploration and exploitation through entropy regularization. This helps improve sample efficiency, which is crucial in power grid scenarios, where acquiring training data is both expensive and time-consuming. Unlike traditional Q-learning, SAC is designed to efficiently handle continuous action spaces. This is crucial in the context of resource allocation in power grids and communication networks. In addition, SAC introduces an entropy term in the objective function, which encourages exploration and prevents the model from prematurely converging to suboptimal solutions. This feature makes SAC more robust when solving complex high-dimensional problems compared to standard Q-learning or DQN.

[0005] After searching, the application publication number CN118589489A, a power grid load forecasting optimization system and method based on deep learning, belongs to the field of electric power grid technology. The present invention solves the problem of inaccurate prediction of existing methods. By setting the time node and the one-time collection duration, the data collection terminal regularly collects the power grid load and transmits it regularly according to the transmission interval time point to ensure the timeliness of data transmission; by establishing a deep learning network model, it is trained and tested multiple times and outputs multiple sets of prediction results; based on the multiple sets of prediction results, corresponding strategy optimization schemes are established to form a database; and search keywords are established so that the client can search according to the current prediction results or the current power grid load attributes to obtain the optimized scheduling strategy for the current power grid load, thereby avoiding the problems of low accuracy, slow data transmission and unstable prediction effect in the existing power grid system, and improving the accuracy of the prediction method of the system.

[0006] Patent CN118589489A focuses on optimizing grid load forecasting. It uses a deep learning model for multiple training and testing to improve forecasting accuracy, and combines it with a strategic optimization solution to provide optimized grid load scheduling. However, this patent has the following shortcomings: 1. While this patent primarily addresses grid load forecasting and improves forecasting accuracy and data transmission efficiency, it does not address resource coordination within the power communication network. 2. While this patent utilizes deep learning for load forecasting, the selection of optimization solutions may still rely on rule-based or traditional optimization methods, lacking adaptive optimization capabilities. The lack of reinforcement learning methods may result in poor system adaptability in complex dynamic environments (such as sudden load fluctuations and grid anomalies), making it incapable of autonomously optimizing resource allocation. 3. While the patent utilizes a timed data collection and transmission approach, which improves the timeliness of data transmission, this approach may not be able to provide a real-time response to sudden grid events (such as load surges or power failures).

[0007] This invention not only considers grid load forecasting but also integrates the optimized scheduling of power communication resources, overcoming the shortcomings of existing patents on multiple levels. 1. This invention not only performs time series prediction (LSTM) of grid load but also considers the needs of the power communication network to ensure the transmission efficiency and reliability of power dispatch instructions. Through resource matching and intelligent scheduling, it improves the coordination between power communication and the power grid, optimizing overall network performance. 2. It introduces deep reinforcement learning (SAC + TD3) to adaptively optimize resource scheduling. SAC automatically adjusts the allocation of grid and communication resources based on load forecast results, improving resource matching efficiency. TD3 further optimizes resource scheduling and enhances policy stability, enabling the system to quickly adapt in dynamic environments and avoiding the rigidity of traditional scheduling methods. 3. Using reinforcement learning, the system can achieve real-time resource adjustments in response to sudden load changes and grid anomalies, rather than relying on fixed-time data collection and scheduling strategies. Integrating data from the power communication network enables more accurate and efficient load scheduling, ensuring timely decision-making. Summary of the Invention

[0008] The present invention aims to solve the above problems of the prior art. It proposes an intelligent method based on the integration of power communication and power grid. The technical solution of the present invention is as follows:

[0009] An intelligent method based on the integration of power communication and power grid, comprising the following steps:

[0010] Step 1: Use the LSTM model to predict future charge fluctuations and changes in communication network demand;

[0011] Step 2: Use the Soft Actor-Critic (SAC) algorithm to perform a preliminary match between grid resources and power communication. By observing the system status, the grid and communication resource allocation scheme is automatically adjusted.

[0012] Step 3: After completing the preliminary matching of grid resources through step 2, the TwinDelayed Deep Deterministic Policy Gradient (TD3) algorithm is used to further optimize resource scheduling efficiency. Based on the different needs of power communication and power grid, intelligent decision-making is implemented to ensure that the system can maximize energy utilization efficiency.

[0013] Furthermore, the step 1: using the LSTM model to predict future charge fluctuations and changes in communication network demand, specifically includes:

[0014] Step 1.1: Collect historical data of the power system, including power load changes, communication network load, resource scheduling historical data, and system response data;

[0015] Step 1.2: Preprocess and extract features from the collected data;

[0016] Step 1.3: Design, train, and evaluate the LSTM model.

[0017] Step 1.4: Use the LSTM model to predict power load and communication demand.

[0018] Furthermore, the step 2: using the SAC algorithm to preliminarily match the grid resources in the power communication, and automatically adjusting the grid and communication resource allocation plan by observing the system status, specifically includes:

[0019] Step 2.1: Collect grid data and environmental information;

[0020] Step 2.2: Define a value network to evaluate state value and action value, specifically including the current load of the power system, communication requirements, energy supply status, equipment health status, and adjust the load distribution of power routes;

[0021] Step 2.3: Define the maximum entropy, specifically including the maximization of system energy utilization, the satisfaction of power communication requirements, and the prevention of system failures;

[0022] Step 2.4: By maximizing the expected return and the entropy of the strategy, a certain flexibility is maintained in the power resource scheduling process.

[0023] Furthermore, step 3: after completing the preliminary matching of grid resources through step 2, the TD3 algorithm is used to further optimize the resource scheduling efficiency and implement intelligent decision-making based on the different needs of power communication and grid, which specifically includes the following steps:

[0024] Step 3.1: Collect power system historical data, power load change data, resource dispatch historical data, and power system response data;

[0025] Step 3.2: Design the environment of the power communication network, including state space variables, action space variables, and reward functions;

[0026] Step 3.3: Train the TD3 algorithm model;

[0027] Step 3.4: Through repeated training, TD3 gradually optimizes the grid resource scheduling strategy to ensure load balancing and maximize communication needs.

[0028] Furthermore, the step 1.3 of designing and evaluating the LSTM model specifically includes:

[0029] The specific components of the LSTM model are as follows:

[0030] (1) Forget gate: determines how much past information is forgotten;

[0031]

[0032] Among them, f t is the output of the forget gate, σ is the Sigmoid activation function, h t-1 is the hidden state of the previous moment, x t is the current input, including power load and communication demand; b f is the bias term of the forget gate.

[0033] (2) Input gate: determines how much new information is added to the cell state;

[0034] i t =σ(W i ·[h t-1 ,x t ]+b i )

[0035] where i t is the output of the input gate, W i is the weight matrix, b i is bias;

[0036] (3) Candidate cell state: Generate new candidate cell state

[0037]

[0038] in is the candidate cell state, tanh is the hyperbolic tangent activation function; b C A bias term representing the candidate cell state.

[0039] (4) Update cell state: Combine the forget gate, input gate and candidate cell state to update the cell state;

[0040]

[0041] Among them C t is the cell state at the current moment, used for long-term memory storage, C t-1 is the cell state at the previous moment, i t is the output of the input gate.

[0042] (5) Output gate: determines the output of the cell state;

[0043] o t =σ(W o ·[h t-1 ,x t ]+b o )

[0044] h t =o t tanh(C t )

[0045] Among them t is the output of the output gate, the hidden state h t Used as output for prediction. b o Represents the bias term of the output gate.

[0046] The goal of the LSTM model is to optimize the network parameters by minimizing the prediction error, and MSE is used for evaluation:

[0047]

[0048] where y i is the true value, is the predicted value of the LSTM model, and N is the number of samples.

[0049] Furthermore, the value network in step 2.2 can be specifically expressed as:

[0050] (1) The SAC strategy evaluation formula is based on the Q-value function and the V-value function, where the Q-value function Q π (s t ,a t ) takes into account the entropy of rewards and policies and expresses it as:

[0051]

[0052] where s t is the current state, a t For the action taken, r t is the immediate reward, γ is the discount factor, V π (s t+1 ) is in the next state s t+1 The value function under

[0053] (2) The V-value function is closely related to the Q-value function. The V-value function is a function that is used to calculate the value of a given state s. t Under this condition, the expected return of taking strategy π is expressed as

[0054]

[0055] Where α is the weight of controlling entropy, which is used to balance task completion and exploration; π(a t ∣s t ) Strategy in state s t Next select action a t probability.

[0056] Furthermore, the specific function for maximizing entropy in step 2.3 can be expressed as:

[0057]

[0058] where τ = (s0, a0, r0, s1, a1, r1, ...) is the trajectory generated by the policy π, s0, a0, and r0 represent the initial state, the action taken, and the reward at the initial time, respectively, and γ t Represents the discount factor at the current moment, r(s t ,a t ) represents the reward for taking an action in the current state, α is the weight of controlling entropy, H(π) is the entropy of the strategy, r t For immediate rewards.

[0059] Furthermore, the entropy update strategy in step 2.4 is specifically expressed as:

[0060] The policy update goal of SAC is to maximize entropy based on the variational derivation of the policy gradient, so as to gradually maximize the weighted sum of the expected return and entropy of the policy. The parameters of the policy are updated by gradient ascent:

[0061]

[0062] in For the experience replay pool.

[0063] Furthermore, the meanings of the variables in step 3.2 are as follows:

[0064] (1) State space variables are used to describe the current environmental state so that TD3 can make reasonable decisions based on this information. t Defined as

[0065] s t =[s1,s2,s3,s4,s5]

[0066] Where s1-s5 represent the power load status, communication demand status, energy supply status, equipment health status and power network topology respectively;

[0067] (2) In the power system, the action space a t Includes all operations that can affect the configuration of grid resources. Specifically, t Expressed as:

[0068] a t =[a1,a2,a3,a4]

[0069] Where a1-a4 represent power load distribution, backup power switch, equipment switch and energy storage control respectively;

[0070] (3) The reward function of power communication and grid resource scheduling can directly reflect the optimization goal of the system. Its comprehensive reward function can be defined as

[0071] r t =λ1f1+λ2f2+λ3f3+λ4f4

[0072] Among them, λ1-λ4 are the weight coefficients of each reward, and f1-f4 represent the load balancing reward, communication demand reward, device monitoring reward and energy efficiency reward respectively.

[0073] Furthermore, the specific steps of model training in step 3.3 are as follows:

[0074] (1) Randomly initialize the Critic network and the Actor network, the target network (Q 1' , Q 2' ) and the target policy network (π');

[0075] (2) In each round of training, sample s from the environment t 、a t 、r t 、s t+1 ;

[0076] (3) Update the Critic network, using the current state s t and action a t , calculate the target Q value network:

[0077]

[0078] Among them, Q1 and Q2 are two critic networks, γ is the discount factor, and the min operation is used to avoid excessive and ensure the stability of the Q value;

[0079] (4) After updating the Critic network twice, update the Actor network again, according to the current state s t Generate a specific scheduling action to maximize the Q value of the Critic network;

[0080]

[0081] where π * (s t ) represents the optimal scheduling strategy under the current state;

[0082] (5) After each update of the Critic network, the target network (Q 1' , Q 2') and the target policy network (π') are soft updated as follows:

[0083] Q1′←τQ1+(1-τ)Q1′

[0084] Q2′←τQ2+(1-τ)Q2′

[0085] π′←τπ+(1-τ)π′

[0086] Where τ is the soft update step size, which ensures that the changes of the target network are slow and smooth;

[0087] (6) Sampling and updating the Critic network and Actor network until the model converges.

[0088] The advantages and beneficial effects of the present invention are as follows:

[0089] This paper introduces an innovative approach to integrating power communication networks with the power grid using LSTM, SAC, and TD3 algorithms. This approach aims to address the limitations of existing systems, particularly the lack of intelligence and inefficient coordination between power grid operations and communication networks. Traditional systems rely on fixed schedules and manual intervention, making them difficult to adapt to dynamic changes in power load and communication demands. The proposed approach leverages machine learning to enable intelligent, real-time decision-making, optimize resource allocation, and ensure the stability and efficiency of the integrated system.

[0090] Step 1: Use an LSTM neural network for time series forecasting, predicting grid load changes and communication demand, providing accurate data support for subsequent resource matching and scheduling. Through historical data training, the system learns power load patterns and communication demand fluctuations, improving forecast accuracy.

[0091] Beneficial Effects: 1. Improved forecasting accuracy. Compared to traditional statistical methods (such as ARIMA), LSTM can better capture nonlinear time series characteristics. 2. Simultaneous forecasting of grid load and communication demand provides a data foundation for the coordinated optimization of power and communication resources, ensuring forward-looking scheduling plans. 3. Enhanced adaptability to dynamic changes. LSTM can learn long-term dependencies, improving the ability to predict sudden load fluctuations.

[0092] Step 2: Use the SAC reinforcement learning algorithm to intelligently match power grid resources and communication resources and automatically adjust the allocation plan. SAC introduces the maximum entropy strategy in reinforcement learning, so that decisions can not only optimize resource utilization but also maintain a certain degree of exploratory nature to avoid local optimality. Through intelligent agents, the optimal mapping relationship between power grid load, communication demand, and resource allocation can be autonomously learned.

[0093] Beneficial Effects: 1. Avoiding the fixed rules of traditional scheduling methods, SAC can adaptively optimize resource allocation based on historical interaction data, improving flexibility. 2. Maximum entropy optimization enhances policy exploration capabilities, balancing utilization and exploration during resource allocation to find the optimal scheduling strategy. 3. Improved resource utilization, rationally matching power and communication resources, and avoiding overload and resource waste.

[0094] Step 3: Based on the initial SAC scheduling, the TD3 algorithm is introduced to further optimize the resource scheduling strategy. TD3 uses a dual Q-value network and a delayed policy update mechanism to reduce the deviation of Q-value estimation and improve learning stability. It is suitable for dynamic environmental changes, such as sudden load fluctuations and power grid anomalies, making the scheduling strategy more stable and reliable.

[0095] Beneficial Effects: 1. Enhanced scheduling policy stability. TD3 reduces Q-value overestimation, reduces decision jitter, and accelerates convergence. 2. Dynamically adapts to environmental changes, responds to emergencies, and improves the reliability of grid-communication coordinated scheduling. 3. Optimizes task coordination, ensuring efficient collaboration between different tasks (load scheduling, communication resource allocation), and improving overall network energy efficiency.

[0096] Cleverness: This invention cleverly combines LSTM prediction + SAC resource matching + TD3 scheduling optimization, giving it significant advantages in prediction accuracy, resource scheduling adaptability, stability and energy efficiency optimization. Compared with traditional methods, it can better improve the intelligence level and scheduling efficiency of the power communication network. BRIEF DESCRIPTION OF THE DRAWINGS

[0097] Figure 1 It is a schematic diagram of an intelligent method based on the integration of power communication and power grid according to a preferred embodiment of the present invention;

[0098] Figure 2 This is the basic structure of the LSTM algorithm of the present invention;

[0099] Figure 3 This is a training flow chart based on the LSTM model of the present invention;

[0100] Figure 4 This is a basic block diagram of the deep reinforcement learning model based on the SAC algorithm of the present invention;

[0101] Figure 5 This is the basic block diagram of the deep reinforcement learning model based on the TD3 algorithm of the present invention. DETAILED DESCRIPTION

[0102] The following will describe the technical solutions in the embodiments of the present invention in detail with reference to the accompanying drawings. The described embodiments are only a part of the embodiments of the present invention.

[0103] The technical solution of the present invention to solve the above technical problems is:

[0104] An intelligent method based on the integration of power communication and power grid, comprising the following steps:

[0105] Step 1: Use the LSTM model to predict future charge fluctuations and changes in communication network demand, plan resource scheduling strategies in advance, and ensure the system's responsiveness to load fluctuations or emergencies.

[0106] Step 2: Use the SAC algorithm to preliminarily match grid resources in power communication. By observing the system status, the grid and communication resource allocation plan is automatically adjusted;

[0107] Step 3: After completing the initial grid resource matching in Step 2, the TD3 algorithm is used to further optimize resource scheduling efficiency. Based on the different needs of power communications and the grid, intelligent decision-making is implemented to ensure that the system can maximize energy efficiency.

[0108] The intelligent method based on the integration of power communication and power grid described in step 1 includes the following steps:

[0109] Step 1.1: Collect historical data of the power system, including power load changes, communication network load, resource scheduling historical data, and system response data;

[0110] Step 1.2: Preprocess and extract features from the collected data to make it suitable for LSTM model training;

[0111] Step 1.3: Design, train, and evaluate the LSTM model.

[0112] Step 1.4: The LSTM model predicts power load and communication demand.

[0113] Furthermore, the step 1.3 designs and evaluates the LSTM model as follows:

[0114] The specific components of the LSTM model are as follows:

[0115] (1) Forget gate: determines how much past information is forgotten.

[0116] f t =σ(W f ·[h t-1 ,x t ]+b f )

[0117] Among them, f t is the output of the forget gate, σ is the Sigmoid activation function, h t-1 is the hidden state of the previous moment, x tis the current input, such as power load, communication demand, etc. b f is the bias term of the forget gate.

[0118] (2) Input gate: determines how much new information is added to the cell state.

[0119] i t =σ(W i ·[h t-1 ,x t ]+b i )

[0120] where i t is the output of the input gate, W i is the weight matrix, b i is the bias.

[0121] (3) Candidate cell state: Generate new candidate cell state

[0122]

[0123] in is the candidate cell state, tanh is the hyperbolic tangent activation function, b C A bias term representing the candidate cell state.

[0124] (4) Update cell state: Combine the forget gate, input gate and candidate cell state to update the cell state.

[0125]

[0126] C t is the cell state at the current moment, used for long-term memory storage, C t-1 is the cell state at the previous moment, i t is the output of the input gate.

[0127] (5) Output gate: determines the output of the cell state.

[0128] o t =σ(W o ·[h t-1 ,x t ]+b o )

[0129] h t =o t tanh(C t )

[0130] Among them t is the output of the output gate, the hidden state h t Used as output for prediction. b o Represents the bias term of the output gate.

[0131] The goal of the LSTM model is to optimize the network parameters by minimizing the prediction error, and MSE is used for evaluation:

[0132]

[0133] where y i is the true value, is the predicted value of the LSTM model, and N is the number of samples.

[0134] The step 2 comprises the following steps:

[0135] Step 2.1: Collect grid data and environmental information;

[0136] Step 2.2: Define a value network to evaluate state value and action value, including the current load of the power system, communication needs, energy supply status, equipment health status, and load distribution of power routes.

[0137] Step 2.3: Define the maximum entropy, specifically including maximizing the system energy utilization, satisfying the power communication requirements, and preventing system failures.

[0138] Step 2.4: By maximizing the expected return and the entropy of the strategy, a certain flexibility is maintained in the power resource scheduling process;

[0139] Moreover, the value network in step 2.2 can be specifically expressed as:

[0140] (1) The SAC strategy evaluation formula is based on the Q-value function and the V-value function, where the Q-value function Q π (s t ,a t ) takes into account the entropy of rewards and policies and expresses it as:

[0141]

[0142] where s t is the current state, a t For the action taken, r t is the immediate reward, γ is the discount factor, V π (s t+1 ) is in the next state s t+1 The value function below.

[0143] (2) The V-value function is closely related to the Q-value function. The V-value function is a function that is used to calculate the value of a given state s. t Under this condition, the expected return of taking strategy π is expressed as

[0144]

[0145] Where α is the weight of controlling entropy, which is used to balance task completion and exploration; π(a t ∣s t ) Strategy in state s t Next select action a t probability.

[0146] Moreover, the specific function for maximizing entropy in step 2.3 can be expressed as:

[0147]

[0148] where τ = (s0, a0, r0, s1, a1, r1, ...) is the trajectory generated by the policy π, s0, a0, and r0 represent the initial state, the action taken, and the reward at the initial time, respectively, and γ t Represents the discount factor at the current moment, r(s t ,a t ) represents the reward for taking an action in the current state, α is the weight of controlling entropy, H(π) is the entropy of the strategy, r t For immediate rewards.

[0149] Moreover, the entropy update strategy in step 2.4 is specifically expressed as:

[0150] The policy update goal of SAC is to maximize entropy based on the variational derivation of the policy gradient, thereby gradually maximizing the weighted sum of the expected return and entropy of the policy. The parameters of the policy are updated by gradient ascent:

[0151]

[0152] in For the experience replay pool.

[0153] Furthermore, the step 3 includes the following steps:

[0154] Step 3.1: Collect power system historical data, power load change data, resource scheduling historical data, and power system response data;

[0155] Step 3.2: Design the environment of the power communication network, including state space variables, action space variables, and reward functions;

[0156] Step 3.3: Train the TD3 algorithm model;

[0157] Step 3.4: Through repeated training, TD3 gradually optimizes the grid resource scheduling strategy to ensure load balancing and maximize communication needs.

[0158] Furthermore, the meanings of the variables in step 3.2 are as follows:

[0159] (1) State space variables are used to describe the current environmental state so that TD3 can make reasonable decisions based on this information. t Defined as

[0160] s t =[s1,s2,s3,s4,s5]

[0161] Among them, s1-s5 represent the power load status, communication demand status, energy supply status, equipment health status and power network topology respectively.

[0162] (2) In the power system, the action space a t Includes all operations that can affect the configuration of grid resources. Specifically, t Expressed as:

[0163] a t =[a1,a2,a3,a4]

[0164] Where a1-a4 represent power load distribution, backup power switch, equipment switch, and energy storage control, respectively. Specific actions can be continuous or discrete. For example, power load distribution can be continuous, indicating the load distribution ratio of each regional power grid.

[0165] (3) The reward function of power communication and grid resource scheduling can directly reflect the optimization goal of the system. Its comprehensive reward function can be defined as

[0166] r t =λ1f1+λ2f2+λ3f3+λ4f4

[0167] Among them, λ1-λ4 are the weight coefficients of each reward, and f1-f4 represent the load balancing reward, communication demand reward, device monitoring reward and energy efficiency reward respectively.

[0168] Moreover, the specific steps of model training in step 3.3 are as follows:

[0169] (1) Randomly initialize the Critic network and the Actor network, the target network (Q 1' , Q 2' ) and the target policy network (π');

[0170] (2) In each round of training, sample s from the environment t 、a t 、r t 、s t+1 ;

[0171] (3) Update the Critic network, using the current state s t and action at , calculate the target Q value network:

[0172]

[0173] Among them, Q1 and Q2 are two critic networks, γ is the discount factor, and the min operation is used to avoid excessive and ensure the stability of the Q value.

[0174] (4) After updating the Critic network twice, update the Actor network again, according to the current state s t Generate a specific scheduling action to maximize the Q value of the Critic network.

[0175]

[0176] where π * (s t ) represents the optimal scheduling strategy under the current state.

[0177] (5) After each update of the Critic network, the target network (Q 1' , Q 2' ) and the target policy network (π') are soft updated as follows:

[0178] Q1′←τQ1+(1-τ)Q1′

[0179] Q2′←τQ2+(1-τ)Q2′

[0180] π′←τπ+(1-τ)π′

[0181] Where τ is the soft update step size, which ensures that the changes of the target network are slow and smooth.

[0182] (6) Sampling and updating the Critic network and Actor network until the model converges.

[0183] In this embodiment: Figure 1 This is a schematic diagram of the intelligent method based on the integration of power communication and power grid of the present invention. Figure 2 This is the basic structure of the LSTM algorithm of the present invention; Figure 3 This is a training flow chart based on the LSTM model of the present invention; Figure 4 This is a basic block diagram of the deep reinforcement learning model based on the SAC algorithm of the present invention; Figure 5 This is the basic block diagram of the deep reinforcement learning model based on the TD3 algorithm of the present invention.

[0184] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions.

[0185] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0186] The above embodiments should be understood as merely illustrating the present invention and not as limiting the scope of protection of the present invention. After reading the contents of the present invention, technicians may make various changes or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention.

Claims

1. An intelligent method based on the integration of power communication and power grid, characterized in that: The following steps are involved: Step 1: Use the LSTM model to predict future charge fluctuations and changes in communication network demand; Step 2: Use the soft actor-critic (SAC) algorithm to perform a preliminary match between grid resources and power communication, and automatically adjust the grid and communication resource allocation plan by observing the system status; Step 3: After completing the preliminary matching of grid resources through step 2, the delayed deep deterministic policy gradient TD3 algorithm is used to further optimize the resource scheduling efficiency. According to the different needs of power communication and power grid, intelligent decision-making is implemented to ensure that the system can maximize energy utilization efficiency.

2. The intelligent method based on the integration of power communication and power grid according to claim 1, characterized in that: Step 1: Using the LSTM model to predict future charge fluctuations and changes in communication network demand, specifically includes: Step 1.1: Collect historical data of the power system, including power load changes, communication network load, resource scheduling historical data, and system response data; Step 1.2: Preprocess and extract features from the collected data; Step 1.3: Design, train, and evaluate the LSTM model. Step 1.4: Use the LSTM model to predict power load and communication demand.

3. The intelligent method based on the integration of power communication and power grid according to claim 1, characterized in that: Step 2: Using the SAC algorithm to preliminarily match the grid resources in the power communication, and automatically adjusting the grid and communication resource allocation plan by observing the system status, specifically including: Step 2.1: Collect grid data and environmental information; Step 2.2: Define a value network to evaluate state value and action value, specifically including the current load of the power system, communication requirements, energy supply status, equipment health status, and adjust the load distribution of power routes; Step 2.3: Define the maximum entropy, specifically including the maximization of system energy utilization, the satisfaction of power communication requirements, and the prevention of system failures; Step 2.4: By maximizing the expected return and the entropy of the strategy, a certain flexibility is maintained in the power resource scheduling process.

4. The intelligent method based on the integration of power communication and power grid according to claim 1, characterized in that: Step 3: After completing the preliminary matching of grid resources in step 2, the TD3 algorithm is used to further optimize the resource scheduling efficiency and implement intelligent decision-making based on the different needs of power communication and grid. Specifically, the following steps are included: Step 3.1: Collect power system historical data, power load change data, resource dispatch historical data, and power system response data; Step 3.2: Design the environment of the power communication network, including state space variables, action space variables, and reward functions; Step 3.3: Train the TD3 algorithm model; Step 3.4: Through repeated training, TD3 gradually optimizes the grid resource scheduling strategy to ensure load balancing and maximize communication needs.

5. The intelligent method based on the integration of power communication and power grid according to claim 2, characterized in that: Step 1.3 of designing and evaluating the LSTM model specifically includes: The specific components of the LSTM model are as follows: (1) Forget gate: determines how much past information is forgotten; f t =σ(W f ·[h t-1 ,x t ]+b f ) Among them, f t is the output of the forget gate, σ is the Sigmoid activation function, h t-1 is the hidden state of the previous moment, x t is the current input, including power load and communication demand; b f is the bias term of the forget gate. (2) Input gate: determines how much new information is added to the cell state; i t =σ(W i ·[h t-1 ,x t ]+b i ) where i t is the output of the input gate, W i is the weight matrix, b i is bias; (3) Candidate cell state: Generate new candidate cell state in is the candidate cell state, tanh is the hyperbolic tangent activation function, b C A bias term representing the candidate cell state. (4) Update cell state: Combine the forget gate, input gate and candidate cell state to update the cell state; C t is the cell state at the current moment, used for long-term memory storage, C t-1 is the cell state at the previous moment, i t is the output of the input gate. (5) Output gate: determines the output of the cell state; the t =σ(W o ·[h t-1 ,x t ]+b o ) h t =o t ·tanh(C t ) Among them t is the output of the output gate, the hidden state h t Used as output for prediction. b o Represents the bias term of the output gate. The goal of the LSTM model is to optimize the network parameters by minimizing the prediction error, and MSE is used for evaluation: where y i is the true value, is the predicted value of the LSTM model, and N is the number of samples.

6. The intelligent method based on the integration of power communication and power grid according to claim 3, characterized in that: The value network in step 2.2 can be specifically expressed as: (1) The SAC strategy evaluation formula is based on the Q-value function and the V-value function, where the Q-value function Q π (s t ,a t ) takes into account the entropy of rewards and policies and expresses it as: where s t is the current state, a t is the current action, r t is the immediate reward, γ is the discount factor, V π (s t+1 ) is in the next state s t+1 The value function under (2) The V-value function is closely related to the Q-value function. The V-value function is a function that is used to calculate the value of a given state s. t Under this condition, the expected return of taking strategy π is expressed as Where α is the weight of controlling entropy, which is used to balance task completion and exploration; π(a t ∣s t ) Strategy in state s t Next select action a t probability.

7. The intelligent method based on the integration of power communication and power grid according to claim 3, characterized in that: The specific function of maximizing entropy in step 2.3 can be expressed as: where τ = (s0, a0, r0, s1, a1, r1, ...) is the trajectory generated by the policy π, s0, a0, and r0 represent the initial state, the action taken, and the reward at the initial time, respectively, and γ t Represents the discount factor at the current moment, r(s t ,a t ) represents the reward for taking an action in the current state, α is the weight of controlling entropy, H(π) is the entropy of the strategy, r t For immediate rewards.

8. The intelligent method based on the integration of power communication and power grid according to claim 3, characterized in that: The entropy update strategy in step 2.4 is specifically expressed as: The policy update goal of SAC is to maximize entropy based on the variational derivation of the policy gradient, so as to gradually maximize the weighted sum of the expected return and entropy of the policy. The parameters of the policy are updated by gradient ascent: in For the experience replay pool.

9. The intelligent method based on the integration of power communication and power grid according to claim 4, characterized in that: The meanings of the variables in step 3.2 are as follows: (1) State space variables are used to describe the current environmental state so that TD3 can make reasonable decisions based on this information. t Defined as s t =[s1,s2,s3,s4,s5] Where s1-s5 represent the power load status, communication demand status, energy supply status, equipment health status and power network topology respectively; (2) In the power system, the action space a t Includes all operations that can affect the configuration of grid resources. Specifically, t Expressed as: <h2 style=";text-align:left;direction:ltr">a<h2 style=";text-align:left;direction:ltr"> t <h2 style=";text-align:left;direction:ltr"> (a1,a2,a3,a4) Where a1-a4 represent power load distribution, backup power switch, equipment switch and energy storage control respectively; (3) The reward function of power communication and grid resource scheduling can directly reflect the optimization goal of the system. Its comprehensive reward function can be defined as r t =λ1f1+λ2f2+λ3f3+λ4f4 Among them, λ1-λ4 are the weight coefficients of each reward, and f1-f4 represent the load balancing reward, communication demand reward, device monitoring reward and energy efficiency reward respectively.

10. The intelligent method based on the integration of power communication and power grid according to claim 9, characterized in that: The specific steps of model training in step 3.3 are as follows: (1) Randomly initialize the Critic network and the Actor network, the target network (Q 1' , Q 2' ) and the target policy network (π'); (2) In each round of training, sample s from the environment t 、a t 、r t 、s t+1 ; (3) Update the Critic network, using the current state s t and action a t , calculate the target Q value network: Among them, Q1 and Q2 are two critic networks, γ is the discount factor, and the min operation is used to avoid excessive and ensure the stability of the Q value; (4) After updating the Critic network twice, update the Actor network again, according to the current state s t Generate a specific scheduling action to maximize the Q value of the Critic network; where π * (s t ) represents the optimal scheduling strategy under the current state; (5) After each update of the Critic network, the target network (Q 1' , Q 2' ) and the target policy network (π') are soft updated as follows: Q1′←τQ1+(1-τ)Q1′ Q2′←τQ2+(1-τ)Q2′ π′←τπ+(1-τ)π′ Where τ is the soft update step size, which ensures that the changes of the target network are slow and smooth; (6) Sampling and updating the Critic network and Actor network until the model converges.

Citation Information

Patent Citations

  • Power grid load prediction optimization system and method based on deep learning

    CN118589489A