Base station energy storage scheduling method and device based on time and graph embedding reinforcement learning

By combining time coding and graph convolutional network reinforcement learning algorithms, the energy storage scheduling of 5G base stations is optimized, which solves the problem of insufficient utilization of graph structure features in existing technologies, realizes more efficient and safer power grid scheduling, and reduces operating costs.

CN120414530AActive Publication Date: 2025-08-01ZHEJIANG UNIV

Patent Information

Application Number
CN202510901151.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-08-01
Estimated Expiration
2045-07-01

AI Technical Summary

Technical Problem

Existing 5G base station energy storage scheduling algorithms fail to fully utilize graph structure feature information, resulting in insufficient decision-making capabilities and difficulty in achieving real-time and secure power grid scheduling in a high-reliability power grid environment. Furthermore, traditional methods involve large computational loads and are susceptible to environmental interference.

Method used

We employ a reinforcement learning approach based on time and graph embedding, combining Time2Vec temporal coding and GCN graph convolutional network to design a deep deterministic policy gradient algorithm, T2V-GCN-DDPG. By optimizing decisions through a safety constraint layer, we improve the agent's state information extraction and decision-making capabilities.

Benefits of technology

It improves the scheduling and decision-making performance of 5G base station energy storage resources, reduces power grid operating costs, and enhances the economy and security of power grid operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120414530A_ABST
    Figure CN120414530A_ABST
Patent Text Reader

Abstract

The invention discloses a base station energy storage scheduling method and device based on time and graph embedding reinforcement learning, and belongs to the field of power distribution network scheduling. Obtaining a graph topological structure according to the power grid system containing 5G base station energy storage; b pieces of power grid training data are randomly selected to train the reinforcement learning agent decision network model, and a trained reinforcement learning agent decision network model is obtained; analyzing and deciding the state observation information of the power grid system at the current moment according to the trained reinforcement learning agent decision network model to obtain an original action vector at the current moment, and performing security constraint to obtain a security action vector; and finally, the power grid system is adjusted through the safety action vector, and scheduling of the power grid system is completed. The scheduling action can be executed according to the real-time state information of the power distribution network, the schedulable capacity of the energy storage of the base station is fully exerted on the premise of ensuring the safe standby capacity of the energy storage of the base station, new energy consumption is promoted, and the operation cost of the power grid is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of distribution network scheduling, and in particular relates to a base station energy storage scheduling method and device based on time and graph embedded reinforcement learning. Background Art

[0002] In recent years, 5G base stations have been built and promoted on a large scale. By December 2024, China had built and commissioned over 4.1 million 5G base stations. During construction, 5G base stations were equipped with energy storage batteries as uninterruptible power supplies. This ensures power supply reliability and communication service quality while providing a degree of power flexibility. This allows for the potential to participate in grid scheduling, leverage peak load shaving, and accommodate renewable energy generation. However, in the current high-reliability power grid environment, the energy storage batteries in 5G base stations remain idle for extended periods, remaining in a floating charge state. This results in significant power wastage and failure to fully utilize their energy storage advantages. It is urgent to prioritize the characteristics of 5G base stations and control their participation in grid scheduling to maximize their energy storage scheduling potential. Furthermore, because the power load of 5G base stations is closely linked to the communication load, the large number of 5G base stations introduces additional uncertainty into the scheduling process, placing higher demands on the algorithm's environmental adaptability and ability to withstand the effects of uncertainty.

[0003] At present, domestic and foreign scholars have conducted extensive research on the control strategy of 5G base station energy storage participating in distribution networks containing new energy. Traditional 5G base station energy storage scheduling algorithms mainly use mathematical optimization methods. The advantage is that the algorithm results have a relatively complete theoretical explanation. However, they rely on accurate prediction information and are more susceptible to environmental interference, which may lead to deviations from the optimal solution. In addition, the model solution process takes a long time and has a large amount of calculation, making it difficult to use this algorithm in real-time adjustment processes.

[0004] With the rise of artificial intelligence algorithms, reinforcement learning algorithms have achieved remarkable results in solving sequential decision-making problems. Reinforcement learning's ability to learn scheduling rules based on environmental data rivals expert experience, enabling it to provide effective adjustment plans in real time and offering strong resilience to uncertainty. However, existing reinforcement learning scheduling algorithms are relatively straightforward in extracting features from the real-time status information of 5G base station energy storage. They lack a targeted structural design tailored to the characteristics of 5G base station distribution, underutilize the structural features of the base station distribution graph, and fail to fully explore these relationships. This fails to fully leverage the advantages of reinforcement learning, and leaves much to be desired in terms of decision-making capabilities. Therefore, there is an urgent need for an algorithm that can fully exploit the features of the base station graph, deliver better decision-making results, and provide more effective real-time scheduling strategies for the safe and economical operation of distribution networks containing 5G base station energy storage resources. Summary of the Invention

[0005] The purpose of the present invention is to address the deficiencies of the existing technology and provide a base station energy storage scheduling method and device based on time and graph embedded reinforcement learning.

[0006] The object of the present invention is achieved by the following technical solutions: A base station energy storage scheduling method based on time and graph embedding reinforcement learning, comprising the following steps: (1) Obtain the graph topology structure according to the power grid system including 5G base station energy storage ; (2) Obtain the power grid training data of the power grid system at multiple moments and put them into the experience replay pool; (3) Randomly select B pieces of power grid training data from the experience replay pool to train the reinforcement learning agent decision network model, and obtain the trained reinforcement learning agent decision network model; (4) Analyze and make decisions on the state observation information of the power grid system at the current moment according to the trained reinforcement learning agent decision network model, obtain the original action vector at the current moment and perform safety constraints to obtain the safe action vector; finally, adjust the power grid system through the safe action vector to complete the scheduling of the power grid system.

[0007] Further, the step (1) is specifically: According to the power grid system including 5G base station energy storage, determine the connection relationship between nodes and branches, and obtain the access location of the superior power grid, the number of nodes , the number of new energy units , the access location and capacity parameters, and the number of 5G base stations and the access location as the graph topology structure of the power grid system .

[0008] Further, the power grid training data includes the state observation information at time , the original action vector , the reward value , and the state observation information at time ; The state observation information at time includes the state observation information at time , the state observation information at time , the set of power grid node state information at time , the set of active power of new energy units , and the set of active power of the superior power grid ; The set of power grid node state information at time includes the active load, reactive load, voltage amplitude, phase angle of all nodes at time , and the state of charge of the 5G base station energy storage; The set of active power of new energy units at time ; The set of power grid node state information at time includes the active load, reactive load, voltage amplitude, phase angle of all nodes at time and the state of charge of the 5G base station energy storage; The set of active power of new energy units at time ; The set of active power of new energy units at time ​ The active power injected by new energy units of all nodes including the moment ; the active power set of the superior power grid at the moment The active power injected by the superior network of all nodes including the moment ; The moment Reward value By normalizing the operating cost of the entire power grid system with 5G base station energy storage And then performing a subtraction operation to calculate.

[0009] Furthermore, the original action vector at the moment Is obtained through the following sub-steps: (a.1) First, use the Time2Vec time encoding embedding model to process the moment To obtain a time encoding vector ; (a.2) Subsequently, use the GCN graph convolution model to process the power grid node status information set To obtain a node feature matrix ; (a.3) Concatenate the time encoding vector , the node feature matrix , the output power set of all new energy units at the moment And the interaction power of the superior power grid at the moment And the moment To form a column vector and input it into a fully connected neural network model to obtain the original action vector ; The original action vector Includes the predicted output power of New energy units, the predicted output power of the superior network, and At the moment The predicted output power of 5G base stations.

[0010] Furthermore, the state observation information at the moment Is obtained through the following sub-steps: (b.1) Input the original action vector Into the power grid system with 5G base station energy storage, and adjust the output power of the superior power grid, New energy units, and 5G base stations in the power grid system; After adjustment, calculate through the active power equation relationship and reactive power equation relationship in the AC power flow to obtain the moment ​​​The active load, reactive load, voltage amplitude, phase angle, active power injected by the energy unit, and active power injected by the superior network of all nodes; (b.2) Calculate the state of charge of the energy storage of 5G base stations at all nodes at time through the energy transfer relationship of the battery power; (b.3) Combine the active load, reactive load, voltage amplitude, phase angle, active power injected by the new energy unit, and active power injected by the superior network of all nodes at time and the state of charge of the energy storage of 5G base stations at all nodes at time to obtain the state observation information at time ; .

[0011] Furthermore, the reinforcement learning agent decision network model consists of an actor network and a critic network; the actor network includes a main policy network and a target policy network; the critic network includes an active action value network and a target action value network; the neural network parameters of the main policy network are , the neural network parameters of the target policy network are , the neural network parameters of the active action value network are , and the neural network parameters of the target action value network are .

[0012] Furthermore, step (3) is specifically: (3.1) Randomly select B grid training data from the experience replay pool. Any grid training data is represented by a 4-tuple , where represents the state observation information at time , represents the original action vector at time , represents the reward value at time , represents the state observation information at time ; (3.2) Construct a loss function and a policy gradient objective function through B grid training data; The calculation formula of the loss function is: , where represents the evaluation value of the active action value network, represents the target value of the target action value network; The target value of the target action value network For , where represents the discount factor in the reinforcement learning algorithm; represents the output action predicted by the target policy network for the state observation information ; The policy gradient objective function has the following calculation formula: ; Subsequently, the gradient of the policy gradient objective function is calculated : , where represents the gradient of the active action value network with respect to the main policy network, represents the gradient of the main policy network with respect to its own neural network parameters ; (3.3) By minimizing the loss function and with respect to the gradient , through the gradient ascent method, the policy gradient objective function is maximized, and the neural network parameters of the active action value network and the neural network parameters of the main policy network are optimized; (3.4) Use to softly update the neural network parameters of the target policy network, and use to softly update the neural network parameters of the target action value network, where represents the neural network parameters of the target policy network after soft update, represents the neural network parameters of the target action value network after soft update; (3.5) Repeat steps (3.1) - (3.4) until the training process of the reinforcement learning agent decision network model converges, and the trained reinforcement learning agent decision network model is obtained.

[0013] Furthermore, the original action vector at the current moment is subjected to safety constraints to obtain a safe action vector, specifically: By minimizing the squared Euclidean distance between the predicted output power of 5G base stations in the original action vector at the current moment and the predicted output power of the 5G base stations after constraint at the current moment, and using the predicted output power of the 5G base stations after constraint at the current moment to replace the corresponding content in the original action vector at the current moment, the corresponding safe action vector is obtained; The predicted output power of the th 5G base station at the current moment needs to satisfy the following constraint conditions: c.1) When the current moment The energy storage power of the th 5G base station is greater than the upper limit of the energy storage power of the th 5G base station at the next moment. When this occurs, the following constraint is satisfied: and where represents the upper limit of the discharging power of the th 5G base station; c.2) When the energy storage power of the th 5G base station at the current moment is less than the lower limit of the energy storage power of the th 5G base station at the next moment, the following constraint is satisfied: and where represents the upper limit of the charging power of the th 5G base station; c.3) Otherwise, the following constraint is satisfied:

[0014] The present invention further includes a base station energy storage scheduling device based on time and graph embedding reinforcement learning, including a memory and one or more processors. Executable code is stored in the memory. When the one or more processors execute the executable code, it is used for the above-mentioned base station energy storage scheduling method based on time and graph embedding reinforcement learning.

[0015] The present invention further includes a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the above-mentioned base station energy storage scheduling method based on time and graph embedding reinforcement learning.

[0016] The beneficial effects of the present invention are as follows: 1) In view of the time characteristics and graph-like distribution structure characteristics presented when 5G base stations participate in the distribution network scheduling, the present invention designs a deep deterministic policy gradient algorithm T2V-GCN-DDPG that integrates time embedding coding and graph convolutional neural network, improving the ability of the reinforcement learning agent to extract real-time state information of the environment, thereby improving the performance of the agent's real-time decision-making; 2) To ensure the safety of decision-making actions, the present invention proposes a safety constraint layer based on quadratic programming, which fine-tunes the original actions output by the model. Through the smallest adjustment amplitude, the actions are made to enter the safe domain, ensuring the safety of the backup power of the base station energy storage; 3) The structural improvement proposed by the present invention plays a key role in enhancing the decision-making ability of the intelligent agent. The final decision result shows that the proposed algorithm can reasonably mobilize the energy storage resources of 5G base stations, reduce the grid operation cost, and improve the overall operation economy. Description of the Drawings

[0017] Figure 1 is the topology diagram of the power grid system with 5G base station energy storage; Figure 2 is the flowchart of a base station energy storage scheduling method based on time and graph embedding reinforcement learning; Figure 3 is the comparison diagram of the cumulative rewards obtained before and after adding time coding and graph convolution embedding; Figure 4 is the comparison diagram of the cumulative rewards obtained before and after imposing security constraints; Figure 5 is the structural diagram of a base station energy storage scheduling device based on time and graph embedding reinforcement learning in Embodiment 2. Detailed Embodiments

[0018] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention will be further described in detail below in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0019] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention will be further described in detail below in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0020] Embodiment 1: In this embodiment, the power grid system targeted is a power grid system with 5G base station energy storage. The topology diagram of the power grid system with 5G base station energy storage is as Figure 1As shown in the figure, the power grid system with 5G base station energy storage is divided into different regions according to functional attributes, including residential areas, campus areas, commercial areas, and industrial areas. There are significant differences in the user power load and the communication load rate of 5G base stations among different regions. Therefore, the present invention sets different power load curves and base station communication load fluctuation curves for four different regions (Region 1, Region 2, Region 3, and Region 4) respectively. This power grid system with 5G base station energy storage is equipped with a 1.5MW photovoltaic power generation unit (PV) at Node 6 and a 3.5MW wind power generation unit (WT) at Node 12. The 5G base station energy storage is converted into an approximate distance according to the branch impedance between different nodes, and the number of 5G base stations to be accessed under each node is determined with reference to the principle of the communication coverage radius of 5G base stations. The scheduling duration of the power grid system is 24h, and the scheduling step interval is 1h. To simulate the prediction deviation in the real scenario, the prediction information of the power grid environment is the actual value superimposed with the prediction deviation, and the prediction deviation follows a Gaussian distribution with a mean of 0 and a standard deviation of 10% of the actual value.

[0021] In this embodiment, the power grid environment runs on MATLAB2022b, and the real-time power flow state information is calculated using MATPOWER 7.1. The intelligent agent algorithm model runs on Python 3.8 and is constructed using Pytorch 2.0.1. The communication between the power grid environment and the intelligent agent is established through matlab engine for python 9.13 for connection and interaction.

[0022] As Figure 2 shown, the present invention provides a base station energy storage scheduling method based on time and graph embedding reinforcement learning, including the following steps: (1) Obtain the graph topology structure according to the power grid system with 5G base station energy storage.

[0023] The specific content of step (1) is as follows: According to the power grid system with 5G base station energy storage, determine the connection relationship between node branches, and obtain the access location of the superior power grid, the number of nodes , the number of new energy generating units , the access location and capacity parameters, and the number of 5G base stations and the access location as the graph topology structure of the power grid system .

[0024] (2) Obtain the power grid training data of the power grid system at multiple moments and put them into the experience replay pool.

[0025] The power grid training data at moment in the experience replay pool includes the state observation information at moment , the original action vector , the reward value , and at moment and at moment State observation information .

[0026] The said moment State observation information is obtained from the power grid system with 5G base station energy storage. The said moment State observation information includes the power grid node state information set at moment , moment , the active power set of new energy units , and the active power set of the superior power grid , that is . The said moment The power grid node state information set at includes the active load, reactive load, voltage amplitude, phase angle of all nodes at moment and the state of charge of 5G base station energy storage at moment . The said moment The power grid node state information set at is , where , represents the characteristic information vector of the th node at moment , , represents the active load of the th node at moment , represents the reactive load of the th node at moment , represents the voltage amplitude of the th node at moment , represents the phase angle of the th node at moment , represents the state of charge of 5G base station energy storage of the th node at moment ; the size of the power grid node state information set is . The said moment The active power set of new energy units at includes the active power injected by energy units of all nodes at moment , the said moment The active power set of new energy units at is , where represents the th The active power injected by new energy units at a node. At the moment The set of active power of the superior power grid Includes the active power injected by the superior network of all nodes at the moment At the moment The set of active power of the superior power grid Is .

[0027] At the moment The original action vector Is obtained through the following sub-steps:

[0028] In this embodiment, an improved T2V-GCN-DDPG model is proposed. The T2V-GCN-DDPG model includes a Time2Vec time encoding embedding model, a GCN graph convolutional model, and a fully connected neural network model DDPG.

[0029] (a.1) First, use the Time2Vec time encoding embedding model to process the moment To obtain a time encoding vector , specifically: Input the moment Into the Time2Vec time encoding embedding model. First, perform a linear term transformation on the moment , and use Calculate the first element Of the time encoding vector, and use it as the non-periodic feature of the time encoding vector; And Represent the model weight parameters when calculating the Th element Of the time encoding vector. Subsequently, use a periodic change method Calculate the other elements of the time encoding vector as the periodic features of the time encoding vector; finally, output the k-dimensional time encoding vector Calculated by the Time2Vec time encoding embedding model. The role of Time2Vec encoding is to encode the moment Into a k-dimensional time encoding vector That can represent other implicit information: , enabling subsequent models to simultaneously mine the non-periodic and periodic features in time information, facilitating algorithm decision-making.

[0030] (a.2) Subsequently, use the GCN graph convolutional model to process the set of power grid node state information To obtain a node feature matrix , specifically including the following sub-steps: The GCN graph convolutional model described in (a.2.1) consists of L graph convolutional networks; according to the graph topology obtain an adjacency matrix of size The adjacency matrix is used to represent the interconnection relationship between nodes in the power grid system with 5G base station energy storage.

[0031] (a.2.2)When the power grid node state information set is input into the GCN graph convolutional model, directly use the power grid node state information set as the node feature matrix extracted by the first layer of the graph convolutional network; use to extract the implicit features in the graph information, and through layer-by-layer convolution, obtain the node feature matrix extracted by the th layer of the graph convolutional network, where , represents the identity matrix of size ; the matrix is , represents the element in the th row and the th column of the matrix , , represents the element in the th row and the th column of the matrix , , . represents the weight matrix of the th layer of the graph convolutional network; represents the node feature matrix extracted by the th layer of the graph convolutional network; represents the non-linear activation function. In this embodiment, the ReLU non-linear activation function is used. Through layer-by-layer convolution operations, the GCN graph convolutional model can aggregate information from the neighbor nodes of each node, which means that as the number of network layers increases, the nodes can capture more global information through their embeddings.

[0032] (a.3)Combine the time encoding vector , the node feature matrix , the output power set of all new energy units at time and the interaction power of the superior power grid at time ​After splicing into a column vector and inputting it into the fully connected neural network model DDPG, the original action vector is obtained. ; The original action vector contains the predicted output power of new energy units, the predicted output power of the upper-level network, and the predicted output power of 5G base stations at a specific time, specifically: First, splice the time encoding vector , the node feature matrix , all elements in the output power set of new energy units at time , and the interaction power of the upper-level power grid at time into a single-dimensional column vector ; Then input this column vector into the fully connected neural network model. After forward propagation, the original action vector is output; The original action vector is , where represents the predicted output power of the th new energy unit at time , represents the predicted output power of the upper-level network at time , represents the predicted output power of the th 5G base station at time . To adapt to the characteristics of different regional loads, the energy storage of 5G base stations adopts a zonal scheduling method, and 5G base stations in the same group share the same scheduling strategy. The output original action vector contains the output power of each controllable unit in the entire power grid system with 5G base station energy storage, which can be used as a coordinated action instruction for the distribution network to control each unit to act together and improve the overall operation effect of the power grid.

[0033] The state observation information at time is obtained through the following sub-steps: (b.1) Input the original action vector into the power grid system with 5G base station energy storage, and adjust the output power of the upper-level power grid, new energy units, and 5G base stations in the power grid system; After adjustment, calculate the time The active load, reactive load, voltage amplitude, phase angle, active power injected by the energy unit, and active power injected by the superior network of all nodes.

[0034] Specifically: Input the original action vector into the power grid system with 5G base station energy storage, and adjust the output power of the superior power grid, new energy units, and 5G base stations in the power grid system; After adjustment, calculate the active load, reactive load, voltage amplitude, phase angle, active power injected by the new energy unit, and active power injected by the superior network of all nodes at time through the active power flow equation relationship and reactive power flow equation relationship in the AC power flow, where represents the active power injected by the superior power grid at the th node at time ; represents the active power injected by the new energy unit at the th node at time ; represents the active load at the th node at time ; represents the active power injected by the 5G base station at the th node at time ; represents the reactive power injected by the superior power grid at the th node at time ; represents the reactive power injected by the new energy unit at the th node at time ; represents the reactive load at the th node at time ; represents the voltage amplitude at the th node at time ; represents the voltage amplitude at the th node at time ; ; represents the real part of the element in the th row and th column of the node admittance matrix of the power grid system, represents the imaginary part of the element in the th row and th column of the node admittance matrix of the power grid system; represents the th node at time The phase angle difference between the th node and the th node, indicating the phase angle of the th node at time ; indicating the phase angle of the th node at time ;

[0035] (b.2) Calculate the state of charge of the energy storage of the 5G base stations of all nodes at time through the energy transfer relationship of the battery power.

[0036] Specifically: Calculate the state of charge of the energy storage of the 5G base stations of each node according to the energy transfer relationship of the battery power , where , represents the charging efficiency, represents the discharging efficiency, represents the charging power of the th 5G base station at time ; at time the discharging power of the th 5G base station; represents the rated capacity of the battery of the th 5G base station.

[0037] (b.3) Combine the active load, reactive load, voltage amplitude, phase angle, active power injected by the new energy unit, and active power injected by the superior network of all nodes at time and time , as well as the state of charge of the energy storage of the 5G base stations of all nodes at time to obtain the state observation information at time .

[0038] The state observation information at the said time includes the set of grid node state information at time , the set of active power of the new energy units , and the set of active power of the superior power grid , that is . The set of grid node state information at the said time includes the active load, reactive load, voltage amplitude, phase angle, and state of charge of the energy storage of the 5G base stations of all nodes at time at time . The said time Set of grid node status information is , where represents the moment the th feature information vector of the node , represents the active power load of the th node at the moment represents the reactive power load of the th node at the moment represents the voltage amplitude of the th node at the moment represents the phase angle of the th node at the moment represents the state of charge of the energy storage of the 5G base station of the th node at the moment ; Set of grid node status information has a size of . The active power set of the new energy units at the moment contains the active power injected by the energy units of all nodes at the moment , and the active power set of the new energy units at the moment is , where represents the active power injected by the new energy unit at the th node at the moment . The active power set of the superior power grid at the moment contains the active power injected by the superior network of all nodes at the moment , and the active power set of the superior power grid at the moment is . The reward value at the moment is calculated by normalizing the operating cost of the entire power grid system with 5G base station energy storage and then performing a subtraction operation, specifically:

[0039] To enable the reward fed back by the distribution network environment with 5G base station energy storage to play a role in guiding the training of the reinforcement learning agent, is used as the reward function, where ​​​​​​​​Represents the overall operating cost of the power grid, including the power purchase cost from the superior power grid , the cost of wind and light curtailment of new energy units , and the maintenance cost of 5G base stations participating in scheduling ; According to the active power of the superior power grid and the power purchase cost per unit time , the power purchase cost of the superior power grid is calculated ; According to the cost of wind and light curtailment per unit time and the new energy unit at time the maximum power and the difference between the actual power , the cost of wind and light curtailment of new energy units is calculated ; According to the cost coefficient of 5G base station energy storage participating in scheduling , the charging power of 5G base stations at time , the discharging power , the maintenance cost of all 5G base stations participating in scheduling is calculated ; In addition, the dimensional is transformed into a normalized value between [0, 1] through a normalization function , and then a negative operation is performed on the normalized value to transform the original cost minimization problem into a cumulative reward maximization problem, which is convenient for algorithm training and convergence.

[0040] (3) Randomly select B power grid training data from the experience replay pool to train the reinforcement learning agent decision network model, and obtain the trained reinforcement learning agent decision network model.

[0041] The reinforcement learning agent decision network model consists of an actor network and a critic network; the actor network includes a main policy network and a target policy network; the critic network includes an active action value network and a target action value network; the neural network parameters of the main policy network are , the neural network parameters of the target policy network are , the neural network parameters of the active action value network are , the neural network parameters of the target action value network are .

[0042] The specific steps of step (3) include the following sub-steps: (3.1) Randomly select B power grid training data from the experience replay pool, and any power grid training data is represented by a 4-tuple , where represents the time State observation information Indicates the moment The original action vector Indicates the moment The reward value Indicates the moment State observation information

[0043] (3.2) Construct the loss function and the policy gradient objective function through B grid training data And the policy gradient objective function .

[0044] The loss function The calculation formula is: , where Represents the evaluation value of the active action value network Represents the target value of the target action value network

[0045] The target value of the target action value network Is , where Represents the discount factor in the reinforcement learning algorithm, which is a hyperparameter of the algorithm; Represents the output action predicted by the target policy network for the state observation information

[0046] The policy gradient objective function The calculation formula is: .

[0047] Subsequently, calculate the gradient of the policy gradient objective function : , where Represents the gradient of the active action value network with respect to the main policy network Represents the gradient of the main policy network with respect to its own neural network parameters , and then use the chain rule to obtain the gradient of the active action value network With respect to

[0048] (3.3) Optimize the neural network parameters of the active action value network and the main policy network by minimizing the loss function And for the gradient Maximize the policy gradient objective function through gradient ascent ,

[0049] (3.4) Use Soft update the neural network parameters of the target policy network, use​​​​ Soft-update the neural network parameters of the target action-value network, where represents the neural network parameters after soft-update of the target policy network, and represents the neural network parameters after soft-update of the target action-value network.

[0050] (3.5) Repeat steps (3.1) - (3.4) until the training process of the reinforcement learning agent's decision network model converges, obtaining the trained reinforcement learning agent's decision network model.

[0051] (4) Analyze and make decisions on the state observation information of the power grid system at the current moment according to the trained reinforcement learning agent's decision network model, obtaining the original action vector at the current moment and perform safety constraints to obtain the safe action vector ; finally, adjust the power grid system through the safe action vector to complete the scheduling of the power grid system.

[0052] The original action vector at the current moment is subjected to safety constraints to obtain the safe action vector, specifically: The current moment of the original action vector is , where represents the predicted output power of the th new energy unit at the current moment, represents the predicted output power of the superior network at the current moment , represents the predicted output power of the th 5G base station at the current moment.

[0053] By minimizing the squared Euclidean distance between the predicted output power of the 5G base stations in the original action vector at the current moment and the constrained predicted output power of the 5G base stations at the current moment , and using the constrained predicted output power of the 5G base stations at the current moment to replace the corresponding content in the original action vector at the current moment, the safe action vector corresponding to the current moment is obtained; where represents the predicted output power of the th 5G base station at the current moment, represents the constrained predicted output power of the

[0054] ​The security action vector corresponding to the current moment is .

[0055] The constrained predicted output power of the th 5G base station at the current moment needs to satisfy the following constraints: c.1) When the energy storage power of the th 5G base station at the current moment is greater than the upper limit of the energy storage power of the th 5G base station at the next moment then it means that the th 5G base station at the current moment can only discharge, and the predicted output power of the th 5G base station at the current moment must be less than 0, that is, the output power needs to reduce the energy storage power to below the upper limit of the energy storage power of the next moment The constrained predicted output power of the th 5G base station at the current moment needs to satisfy the following constraints: And where represents the upper limit of the discharge power of the th 5G base station, represents the discharge efficiency.

[0056] ​​​​​​​​​​​​​​​​​​​​​The upper limit of the charging power of a 5G base station, represents the charging efficiency.

[0057] c.3) Conversely, when at the current moment the energy storage power of the nth 5G base station is not less than the lower limit of the energy storage power of the nth 5G base station at the next moment and not greater than the upper limit of the energy storage power of the nth 5G base station at the next moment then it means that the nth 5G base station at the current moment can be charged and discharged, and the discharge does not exceed the lower limit of the energy storage power at the next moment and the charge does not exceed the upper limit of the energy storage power at the next moment. The predicted output power of the nth 5G base station at the current moment after constraint needs to satisfy the following constraint conditions: and . The predicted output power of the nth 5G base station and .

[0058] The charging and discharging power of each 5G base station needs to satisfy the constraint of the energy storage power of the 5G base station, that is, for the energy storage power of the nth 5G base station at the current moment it needs to satisfy , where the upper limit of the energy storage power of the nth 5G base station at the next moment depends on the maximum battery capacity parameter of the nth 5G base station; the lower limit of the energy storage power of the nth 5G base station at the next moment depends on the integral value of the communication load power of the nth 5G base station at the next moment over the minimum power backup time , that is ; the energy storage power of the nth 5G base station at the current moment at the minimum power backup time is obtained by converting from the battery physical parameters of the nth 5G base station. ; the energy storage power of the nth 5G base station at the current moment is derived from the battery physical parameters of the nth 5G base station.

[0059] The lower limit of the energy storage power of the nth 5G base station at the next moment ​ Derived from the battery physical parameters of the th 5G base station and the communication load rate of the base station at the next moment .

[0060] Since the predicted output power of the th 5G base station in the original action vector output cannot fully ensure the safe operation of the 5G base station, through safety constraints, the predicted output power of the th 5G base station is safety-constrained to ensure the safe operation of the 5G base station.

[0061] The principle of safety constraint implementation is to use the predicted output power as the decision variable, and according to the upper and lower limits of the energy storage power of the th 5G base station at the next moment , inversely deduce the energy storage power of the th 5G base station at the current moment . Under the condition of the energy storage power of the th 5G base station at the current moment , the power boundaries of charging and discharging that can be taken are used as the safe feasible region of the decision variable. By minimizing the squared Euclidean distance between the predicted output power of the th 5G base station at the current moment and the predicted output power of the th 5G base station after constraint at the current moment . The predicted output power outside the safe feasible region can be moved to the nearest feasible region boundary in the way of the minimum moving distance, so as to obtain the predicted output power after constraint that meets the safety constraint conditions. Through this method, while ensuring the safety constraint boundary of the energy storage, the information of the original action is also retained as much as possible.

[0062] To verify the effectiveness of a base station energy storage scheduling method based on time and graph embedding reinforcement learning proposed by the present invention. Figure 3 Shows the comparison of the cumulative rewards obtained before and after adding time coding and graph convolution embedding. From the Figure 3 comparison, it can be seen that the improved T2V-GCN-DDPG model proposed in this patent has a significant improvement in the convergence speed and stability compared with the unimproved fully connected neural network model DDPG, indicating that the use of the Time2Vec time coding embedding model and the GCN graph convolution model can improve the feature extraction ability of the agent and can significantly accelerate the convergence process of the model.

[0063] Figure 4 Shows the comparison of the cumulative rewards obtained before and after safety constraint. Figure 4As can be seen, when safety constraints are imposed, the convergence speed of algorithm training will be significantly accelerated, the cumulative reward of the final convergence result will be higher, and the decision-making level will be stronger.

[0064] Embodiment 2: This embodiment relates to a base station energy storage scheduling device based on time and graph embedding reinforcement learning, including a memory and one or more processors. Executable code is stored in the memory. When the one or more processors execute the executable code, it is used for the method of base station energy storage scheduling based on time and graph embedding reinforcement learning in Embodiment 1 above. The device embodiment can be applied to any device with data processing capabilities, and the any device with data processing capabilities can be a device or apparatus such as a computer.

[0065] As Figure 5 , at the hardware level, the device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above Figure 1 shown method. Of course, in addition to the software implementation method, the present invention does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or a logic device.

[0066] Improvements to a technology can be clearly distinguished as hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many improvements to method flows today can be regarded as direct improvements to hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented with a hardware entity module. For example, a programmable logic device (PLD) (e.g., a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user programming the device. The designer can program by himself to "integrate" a digital system on a piece of PLD without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a hardware description language (HDL), and there is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that as long as the method flow is slightly logically programmed with the above-mentioned several hardware description languages and programmed into the integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.

[0067] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.

[0068] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0069] It should also be noted that the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity, or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such a process, method, commodity, or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of another identical element in the process, method, commodity, or device comprising the said element.

[0070] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0071] The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0072] Embodiment 3: The embodiment of the present invention further provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the method for base station energy storage scheduling based on time and graph embedding reinforcement learning in Embodiment 1 above.

[0073] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A base station energy storage scheduling method based on time and graph embedding reinforcement learning, characterized in that It includes the following steps: (1)Obtain the graph topology structure from the power grid system with 5G base station energy storage ; (2) Obtain the power grid training data of the power grid system at multiple moments and put them into the experience replay pool; (3) Randomly select B power grid training data from the experience replay pool to train the reinforcement learning agent decision network model, and obtain the trained reinforcement learning agent decision network model; (4) Analyze and make decisions on the state observation information of the power grid system at the current moment according to the trained reinforcement learning agent decision network model, obtain the original action vector at the current moment and perform safety constraints to obtain the safe action vector; finally, adjust the power grid system through the safe action vector to complete the scheduling of the power grid system.

2. The base station energy storage scheduling method based on time and graph embedding reinforcement learning according to claim 1, wherein The specific content of step (1) is: Based on the power grid system with 5G base station energy storage, determine the connection relationship between node branches, and obtain the access location of the superior power grid, the number of nodes , the number of new energy units , the access location, capacity parameters, and the number of 5G base stations and the access location as the graph topology of the power grid system .

3. The base station energy storage scheduling method based on time and graph embedding reinforcement learning according to claim 2, characterized in that, The power grid training data includes time State observation information , original motion vector , reward value and time State observation information ; The state observation information at the described moment includes the power grid node state information sets at moment , moment , the active power sets of new energy units , and the active power sets of the superior power grid ; The power grid node state information set at the described moment includes the active load, reactive load, voltage amplitude, phase angle of all nodes at moment , and the state of charge of the energy storage of 5G base stations; The active power set of new energy units at the described moment includes the active power injected by new energy units of all nodes at moment ; The active power set of the superior power grid at the described moment includes the active power injected by the superior network of all nodes at moment ;​​ The said moment of the reward value is calculated by normalizing the operating cost of the entire power grid system with 5G base station energy storage and then performing a subtraction operation 4. The base station energy storage scheduling method based on time and graph embedding reinforcement learning according to claim 3, wherein The said moment of the original action vector is obtained through the following sub-steps: (a.1) First, use the Time2Vec time encoding embedding model to process the moment to obtain the time encoding vector ; (a.2) Subsequently, the GCN graph convolutional model is used to process the power grid node status information set to obtain the node feature matrix ; (a.3) Concatenate the time-encoded vector , the node feature matrix , the set of output powers of all new energy units at time and the interactive power of the superior power grid at time . After concatenating them into a column vector and inputting it into the fully connected neural network model, the original action vector is obtained; the original action vector contains the predicted output powers of new energy units at time , the predicted output power of the superior network, and the predicted output powers of 5G base stations.

5. The base station energy storage scheduling method based on time and graph embedding reinforcement learning according to claim 3, characterized in that The state observation information at the moment is obtained through the following sub-steps: is obtained through the following sub-steps: (b.1) Input the original action vector into the power grid system with 5G base station energy storage, and adjust the output power of the superior power grid, new energy units, and 5G base stations in the power grid system; after adjustment, calculate the active load, reactive load, voltage amplitude, phase angle, active power injected by the energy unit, and active power injected by the superior network of all nodes at time through the active power flow equation relationship and reactive power flow equation relationship in the AC power flow; (b.2) The state of charge of the energy storage of the 5G base station for all nodes at the moment is calculated through the energy transfer relationship of the battery power; (b.3) Combine the active power load, reactive power load, voltage amplitude, phase angle, active power injected by new energy units, and active power injected by the superior network of all nodes at time with the state of charge of the 5G base station energy storage of all nodes at time to obtain the state observation information at time .

6. The base station energy storage scheduling method based on time and graph embedding reinforcement learning according to claim 3, characterized in that The decision-making network model of the reinforcement learning agent consists of an actor network and a critic network; the actor network includes a main policy network and a target policy network; the critic network includes an active action value network and a target action value network; the neural network parameters of the main policy network are , the neural network parameters of the target policy network are , the neural network parameters of the active action value network are , the neural network parameters of the target action value network are .

7. The base station energy storage scheduling method based on time and graph embedding reinforcement learning according to claim 6, wherein The specific content of step (3) is: (3.1) Randomly select B grid training data from the experience replay pool, and any grid training data is represented by a 4-tuple , where represents the state observation information at time , represents the original action vector at time , represents the reward value at time , represents the state observation information at time ; (3.2) The loss function is constructed from B grid training data and the policy gradient objective function ; The loss function has the following calculation formula: , where represents the evaluation value of the active action value network, represents the target value of the target action value network; The target value of the target action value network is , where represents the discount factor in the reinforcement learning algorithm; represents the output action predicted by the target policy network for the state observation information ​ The above-mentioned policy gradient objective function has the following calculation formula: ; Subsequently, the policy gradient objective function is calculated for the gradient : , where represents the gradient of the active value network with respect to the main policy network, represents the gradient of the main policy network with respect to its own neural network parameters ; (3.3) By minimizing the loss function and with respect to the gradient By means of gradient ascent, maximize the policy gradient objective function to optimize the neural network parameters of the active value network and the neural network parameters of the main policy network; (3.4) Usage Soft-update the neural network parameters of the target policy network, using Soft-update the neural network parameters of the target action-value network, where represents the neural network parameters after soft-update of the target policy network, represents the neural network parameters after soft-update of the target action-value network; (3.5) Repeat steps (3.1)-(3.4) until the training process of the reinforcement learning agent decision network model converges, and obtain the trained reinforcement learning agent decision network model.

8. A base station energy storage scheduling method based on time and graph embedding reinforcement learning according to claim 7, characterized in that The original action vector at the current moment and perform safety constraints to obtain the safe action vector, specifically: By minimizing the square Euclidean distance between the predicted output power of 5G base stations at the current moment and the predicted output power of 5G base stations after constraint at the current moment, and using the predicted output power of 5G base stations after constraint at the current moment to replace the corresponding content in the original action vector at the current moment, a corresponding safe action vector is obtained;​​​​​​ The predicted output power after constraint of the th 5G base station at the current moment needs to satisfy the following constraint conditions: c.1) When the current moment of the energy storage power of the th 5G base station is greater than the upper limit of the energy storage power of the th 5G base station at the next moment , the constraint condition is satisfied: and , where represents the upper limit of the discharge power of the th 5G base station; c.2) When the current moment of the energy storage power of the th 5G base station is less than the lower limit of the energy storage power of the th 5G base station at the next moment , the constraint condition is satisfied: and where , represents the upper limit of the charging power of the th 5G base station; c.3) Conversely, the constraint condition is satisfied: and .

9. A base station energy storage scheduling device based on time and graph embedding reinforcement learning, characterized in that, It includes a memory and one or more processors. Executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement a base station energy storage scheduling method based on time and graph embedding reinforcement learning according to any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, A program is stored thereon. When the program is executed by a processor, it implements a base station energy storage scheduling method based on time and graph embedding reinforcement learning according to any one of claims 1-8.

Citation Information

Patent Citations

  • Deep reinforcement learning economic dispatching method based on pre-training and knowledge guidance

    CN116468106A

  • Digital-analog combined drive graph depth reinforcement learning power system optimization scheduling method

    CN118523284A

  • Power grid topology optimization method based on search sorting

    CN118676903A

  • 5G base station group intra-day optimization operation method based on communication load migration and energy storage scheduling

    CN120186644A

  • Microgrid space-time perception energy management method based on secure deep reinforcement learning

    WO2024108817A1

Cited By

  • Control network training method and device of power distribution network and control method of power distribution network

    CN120999641A