5g base station energy storage and power distribution network coordinated optimization scheduling method considering power supply reliability

By establishing a dispatchable capacity assessment model and optimizing charging and discharging strategies in 5G base station energy storage systems, the problems of low energy storage utilization and volatility of new energy sources have been solved, achieving efficient utilization of energy storage resources and economically optimized scheduling of the power grid.

CN117748508BActive Publication Date: 2026-05-12SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SICHUAN UNIV
Filing Date
2023-09-28
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

The current low utilization rate of energy storage batteries in 5G base stations leads to a significant waste of flexible energy storage resources. Furthermore, the randomness and volatility of new energy power generation have an adverse impact on the power system, necessitating improvements in the utilization rate and absorption capacity of energy storage equipment.

Method used

By establishing a base station energy storage schedulable capacity assessment model, using the VAE model to encode the energy storage state and embedding the TD3 algorithm, a Markov decision process is constructed to optimize the charging and discharging strategy of base station energy storage. With the goal of minimizing system operating costs, the charging and discharging of base station energy storage equipment is reasonably controlled.

Benefits of technology

It has improved the utilization rate of energy storage in 5G base stations, reduced wind curtailment, lowered system operating costs, and enhanced the level of new energy consumption and the operational stability of the power grid.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117748508B_ABST
    Figure CN117748508B_ABST
Patent Text Reader

Abstract

The application provides a 5G base station energy storage and power distribution network collaborative optimization scheduling method considering power supply reliability. Through pre-training of a VAE model, real-time state of charge data dimension reduction and feature extraction of a multi-base station energy storage system are realized, so as to adapt to communication base stations with different regional characteristics, avoid dimension disaster caused by too large number of base stations, fully retain the characteristics of each base station, and facilitate the specification of individualized charging and discharging strategies. Meanwhile, the scheduling control model is configured as a Markov decision process with the minimum system operation cost as the target, and the VAE model is embedded into the TD3 algorithm, and the VAE state code is used as the agent for training. Since the model considers the power generation cost, energy storage battery loss cost and base station energy storage leasing cost of the system, the influence of wind power generation and energy storage equipment can be fully reflected, the energy storage charging and discharging strategy of each communication base station can be reasonably controlled, and the wind power consumption level and the system operation cost can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power dispatching technology, and in particular to a method for coordinated optimization dispatching of 5G base station energy storage and distribution network that takes into account power supply reliability. Background Technology

[0002] As the construction of new power systems deepens and the installed capacity of renewable energy generation continues to increase, the power grid urgently needs flexible energy storage devices to participate in system dispatch, in order to promote the consumption of new energy and alleviate the pressure of peak load electricity demand. However, the configuration of energy storage devices is limited by construction costs, resulting in insufficient current capacity. Therefore, fully exploring the idle flexible energy storage resources in the power grid has become the key to solving this problem.

[0003] In recent years, 5G communication technology has developed rapidly. As the core equipment of the fifth-generation mobile communication network, the number of 5G base stations has been growing rapidly year by year. By the end of 2022, the total number of 5G base stations built in my country had reached 2.312 million. To ensure uninterrupted power supply to communication equipment, mobile communication operators have installed battery energy storage as backup power for communication base stations. However, the generally high reliability of the current power grid leads to low utilization rates of base station energy storage batteries, resulting in a significant waste of flexible energy storage resources. Therefore, it is urgent to systematically evaluate the dispatchable capacity of base station energy storage based on the load characteristics of 5G base stations, and to participate idle base station energy storage resources in the coordinated operation of the power grid, thereby improving the utilization rate of base station energy storage and achieving a win-win situation for both communication operators and the power grid. Summary of the Invention

[0004] Based on the problems that need to be solved in the existing technology, the present invention provides a 5G base station energy storage and distribution network collaborative optimization scheduling method that takes into account power supply reliability. By treating the idle energy storage in a large number of 5G base stations as a flexible resource to participate in power system scheduling, the method can reduce the adverse effects of the randomness and volatility of new energy power generation on the power system.

[0005] The present invention provides a method for coordinated optimization scheduling of 5G base station energy storage and distribution network considering power supply reliability, which includes the following steps:

[0006] S1: After each decision cycle, obtain the state information S = [M] corresponding to the current time t. m,soc (t),L m,soc (t),U j [(t),W(t),t]; where M m,soc (t) represents the SOC[M] stored in N base stations at time t. 1,soc (t),···,M i,soc (t),···,M N,soc The VAE state encoding of [t], L m,soc(t) represents the lower limit of the SOC (System-on-Chip) capacity of N base stations at time t [L]. 1,soc (t),···,L i,soc (t),···,L N,soc The VAE state encoding of [(t)], where m is the dimension of the encoded state; U j W(t) represents the node voltage of node j in the distribution network at time t; W(t) represents the output of the wind turbine at time t; t is the time marker.

[0007] S2: Obtain the acquired status information S = [M m,soc (t),L m,soc (t),U j [(t), W(t), t] are input to the trained scheduling control model and output the corresponding scheduling control strategy; wherein, the scheduling control model is configured as a Markov decision process with the goal of minimizing the system operating cost, and the VAE model is embedded into the TD3 algorithm, and the VAE state code is used as the agent for training;

[0008] S3: According to the scheduling and control strategy, the distribution network executes the corresponding strategy actions to control the charging and discharging status of each 5G base station energy storage device.

[0009] According to a specific implementation method, in the 5G base station energy storage and distribution network collaborative optimization scheduling method considering power supply reliability provided by the present invention, the lower limit of the base station's energy storage SOC capacity at time t is... for: Among them, the load level of each base station The load rate characteristics of base stations in each load node area are as follows: The minimum backup power time t for base station energy storage at each load node is determined. s Minimum availability index of base stations Decide, The base station reserves a backup time t for the failure rate of the load point where the base station is located. s The probability P s With the probability distribution of power outage time Q s (t) correlation, i.e. Power outage time probability distribution Q s (t) is generated by fitting historical data of power outage time in the distribution network.

[0010] According to a specific implementation, the 5G base station energy storage and distribution network collaborative optimization scheduling method considering power supply reliability provided by the present invention uses a VAE model to encode the base station's state of charge. The VAE model includes an inference network and a generation network; wherein, the inference network is parameterized. of The probability distribution model encodes the input data into a latent variable z using variational inference, with p(z) following a standard normal distribution. The generative network uses the latent variable z to reconstruct the input data and generate p with parameter θ. θ (x|z) approximate probability distribution.

[0011] According to a specific implementation, in the 5G base station energy storage and distribution network collaborative optimization scheduling method considering power supply reliability provided by the present invention, the objective function is configured as follows:

[0012] F = min(C) gen +C BES +C Lease )

[0013]

[0014]

[0015]

[0016] Among them, C gen For the cost of generating electricity by the generator, a i b i c i Let P be the power generation cost coefficient of generator unit i. Git To generate electricity for the generator; C BES C represents the cost of charge / discharge degradation of base station energy storage batteries. a Where N is the cost discount factor for energy storage batteries, and N is the total number of energy storage units in the base station. C represents the real-time charging and discharging power of base station i, respectively, where Δt is the unit time interval; Lease For the cost of base station energy storage leasing, C b This is the rental cost coefficient; These are the upper and lower limits for real-time scheduling, respectively; ξ(t) is represented by (0,1), where 0 represents that the energy storage participates in the scheduling at time t, and 1 represents that the energy storage does not participate in the scheduling at time t.

[0017] According to one specific implementation, when training an agent based on the TD3 algorithm, the TD3 algorithm uses a Critic network and an Actor network as its framework. Furthermore, it parameterizes the Critic network using a neural network, approximates the value function by minimizing the Bellman residual, establishes the loss function L(θ) using the Monte Carlo sampling approximation, and then updates the parameters of the current Critic network by minimizing the loss function. The loss function L(θ) is specifically: N is the number of training samples; y i The target Q value; Q On This is the current Q value; Let these be the current Q-network parameters; and let's parameterize the Actor network using a neural network and update its parameters using gradient descent, i.e.

[0018] Furthermore, in the 5G base station energy storage and distribution network collaborative optimization scheduling method considering power supply reliability provided by this invention, the Critic network adopts two sets of target Q networks with the same network architecture, and selects the minimum value between the two as the target value, i.e. in, Here are the network parameters for two different target Q-networks, where a′ represents the action with random noise; and a′ = μ′(s t+1 )+ε,μ′ represents the target network strategy, and ε represents normally distributed random noise, which is obtained by truncating the sampled noise.

[0019] According to a specific implementation, in the 5G base station energy storage and distribution network collaborative optimization scheduling method considering power supply reliability provided by the present invention, during the process of training the agent based on the TD3 algorithm, the policy network and target value network parameters are also updated in a soft update manner, that is:

[0020]

[0021] θ μ′ =τθ μ +(1-τ)θ μ′ .

[0022] Furthermore, during the training of the agent based on the TD3 algorithm, the reward obtained by the agent through interaction with the environment is the sum of the generator cost, the battery charging and discharging loss cost, and the energy storage leasing cost:

[0023]

[0024] Where N represents the amount of energy stored in the base station, and T is the scheduling period.

[0025] Furthermore, during the training of the agent based on the TD3 algorithm, the penalty function used to correct inappropriate actions of the agent is configured as follows:

[0026] P b =||a as -a re ||2

[0027] Among them, a as Auxiliary vector for base station actions; a re This is the action vector of the agent under its current policy.

[0028] According to one specific implementation, when training the agent based on the TD3 algorithm, pre-training is also performed in conjunction with the behavior cloning algorithm:

[0029]

[0030] in, β represents the action strategy adopted by the agent when exploring the environment during training, ρ β (s) represents the distribution of the current state under policy μ; Q μ (s,μ(s)) represents the value of the action according to policy μ in state s, J β (μ) represents the expected return that can be obtained. μ i σ represents expectation; i Indicates variance.

[0031] Thus, the 5G base station energy storage and distribution network collaborative optimization scheduling method considering power supply reliability provided by this invention systematically evaluates the real-time schedulable capacity of base station energy storage by establishing a base station energy storage schedulable capacity assessment model. Furthermore, by pre-training a VAE model, it achieves dimensionality reduction and feature extraction of real-time state-of-charge data for multi-base station energy storage systems, thereby adapting to communication base stations with different regional characteristics. This avoids the dimensionality curse caused by an excessive number of base stations while preserving the characteristics of each base station, enabling the development of personalized charging and discharging strategies for each base station's energy storage. Simultaneously, the scheduling control model is configured as a Markov decision process with the goal of minimizing system operating costs, and the VAE model is embedded into the TD3 algorithm, using VAE state codes as agents for training. Since the model considers the system's power generation costs, energy storage battery loss costs, and base station energy storage leasing costs, it can comprehensively reflect the impact of wind power generation and energy storage equipment, rationally controlling the energy storage charging and discharging strategies of each communication base station. This effectively reduces system operating costs while improving wind power absorption. Attached image description:

[0032] Figure 1 This is a schematic diagram illustrating the relationship between power consumption and load rate of a 5G communication base station in this invention;

[0033] Figure 2 A schematic diagram of the probability density distribution curve of power outage time in a certain region's power distribution network;

[0034] Figure 3 This is a schematic diagram illustrating the operation of the scheduling and control model of the present invention, which consists of an environment model based on VAE state coding and a TD3 framework.

[0035] Figure 4 This is a flowchart illustrating the scheduling method of the present invention;

[0036] Figure 5This is a schematic diagram of the distribution network topology;

[0037] Figure 6 A comparison chart of VAE model pre-training performance;

[0038] Figure 7 A comparison chart showing the training results of the scheduling and control model;

[0039] Figure 8 This is a schematic diagram comparing the wind power absorption levels before and after dispatching.

[0040] Figure 9 This is a schematic diagram comparing the voltage levels of base station load nodes before and after scheduling. Detailed Implementation

[0041] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments. However, this should not be construed as limiting the scope of the present invention to the following embodiments; all technologies implemented based on the content of the present invention fall within the scope of the present invention.

[0042] To treat the idle energy storage in a large number of 5G base stations as a flexible resource for power system dispatch, it is necessary to establish a model to accurately assess the dispatchable capacity of base stations. Taking a typical distribution network topology as the operating scenario, a distribution network node reliability assessment model is established. By combining the probability distribution of power outage time in the distribution network, the minimum backup hours of each base station node can be obtained by joint solution. At the same time, taking into account the differences in load characteristics of different base stations, the minimum reserve capacity of each base station is assessed.

[0043] Specifically, a 5G base station mainly consists of a communication section and a power supply section. The communication section primarily includes active antenna units (AAU), baseband units (BBU), and network transmission equipment. The power supply section mainly includes an external power supply and a backup battery. The external power supply uses an AC power grid interface, converting AC power to DC power through a rectifier. The backup battery is prepared for potential power grid failures; when a failure occurs, the power to the communication equipment will be provided by the backup battery to ensure the safe and reliable operation of the base station. Therefore, a power consumption model for the communication equipment and a backup battery model can be established based on the configuration of the communication base station.

[0044] Taking a typical macro base station with 3 AAUs and 1 BBU as an example, base station power consumption includes two parts: dynamic power consumption and static power consumption. Static power consumption includes the power amplifier power consumption in the AAU, the power consumption of the BBU, and the power consumption of the transmission equipment. Dynamic power consumption is the power consumption of the radio frequency unit in the AAU, which is related to the base station communication load rate and has an approximately linear relationship, as shown in the characteristic curve. Figure 1 As shown, the power consumption of a single base station communication device can therefore be expressed as:

[0045]

[0046] In the above formula, Let t be the total power consumption of the communication equipment at time t for a single base station; This refers to the power amplifier power consumption in the AAU; This refers to the power consumption of the BBU. The power consumption of the transmission equipment; This refers to the power consumption of the radio frequency unit in the AAU; This represents the static power consumption of a single base station. This represents the maximum dynamic power consumption of a single base station; μ t This represents the communication load rate of a single base station, with a value range of (0,1).

[0047] The energy storage batteries for 5G base stations generally use lithium iron phosphate batteries that are designed for cascaded use. These batteries have high energy density, fast response capability, and can support a high frequency of charge and discharge cycles, giving them a good ability to participate in grid demand response. Therefore, the mathematical model for their charge and discharge is established as follows:

[0048] 1. Battery charge / discharge state constraints

[0049]

[0050] In the above formula, Let be the energy storage capacity at time t. η represents the energy storage margin at time t-1; ch η dis These are the charge and discharge efficiencies of the energy storage battery; These represent the charging and discharging power of the energy storage battery. This refers to the charge / discharge state of the energy storage battery, ensuring that the battery can only exist in one state at any given time.

[0051] 2. Battery charge and discharge upper and lower limit constraints

[0052]

[0053] In the above formula, This indicates the upper limit of the charging and discharging power of the energy storage battery; This represents the minimum capacity of backup energy storage required at time t to ensure the reliability of the base station. This indicates the upper limit of the base station's energy storage battery capacity.

[0054] Furthermore, the charging and discharging efficiency of energy storage batteries is related to the battery's current state of charge (SOC) and charging and discharging power, and can be approximated using a second-order polynomial:

[0055]

[0056] In the formula: a0~a5 and b0~b5 are fitting coefficients; S soc This represents the current state of charge of the battery; P ch P dis This refers to the current charging and discharging power of the battery.

[0057] Using a typical distribution network topology as the operating scenario, a reliability assessment model for distribution network nodes is established. Specifically, the states of components such as feeders, transformers, disconnectors, and circuit breakers in the distribution network are represented by (0,1), and the failure rate of component i is set to λ. i Each component in the circuit is numbered, and its state space is established as follows:

[0058]

[0059] For a radial power grid, based on its distribution network topology, a two-dimensional array can be used to store the connection relationships between nodes and generate its graph adjacency matrix, i.e.:

[0060]

[0061] In the above formula, v ij This indicates the connectivity between node i and node j; a connection is represented by 1, and a non-connection is represented by 0. n This represents the number of nodes in the distribution network. Based on the adjacency matrix above, a depth-first search algorithm can be used to obtain all paths from the root node to the load point n. The states of all elements along each path constitute a set {x1, x2, x3, ... xn}. k}, where x i Let represent the state of element i in the set during time period t. Considering that multiple elements rarely fail simultaneously in a power grid, this paper only discusses the reliability of each load node when there are a finite number of faulty elements or lines. For load point n, when there are a finite number of faulty elements in the system, if there is an available path from the root node to the load node (i.e., all elements along this path are working normally), then the load point is working normally; if there is no available path from the root node to the load point, then the load point is faulty. Based on the above, a fault state set S for the load points can be generated. i Then, based on the reliability data of the faulty components, the reliability index of the load point can be obtained.

[0062] When the load point is operating normally, the communication base station also operates normally, and the base station power supply reliability meets the standards. When the load point fails, it is necessary to assess whether the base station's reserved capacity meets the availability index. The assessment of the base station's reserved capacity is related to the power outage time of the distribution network, and the minimum backup hours for the base station can be determined based on the power outage time probability distribution of the distribution network; where the power outage time probability distribution Q of the distribution network... s (t) can be generated by fitting historical data of power outage times from the corresponding regional power distribution network, for example... Figure 2 As shown, the power outage time in a certain area is mainly concentrated between 0 and 3 hours, and the probability of power outage due to distribution network failure can be restored within 4 hours is higher than 95%. Since the normal backup power time of 5G base stations is about 4 hours, the backup energy storage of base stations still has a certain scheduling potential while ensuring its own communication reliability.

[0063] Therefore, based on the reliability indices of the distribution network load points and their outage time probability distribution curves generated above, the minimum backup power time for base station energy storage can be calculated according to the base station availability index. Based on the outage time probability distribution, it can be concluded that the time the distribution network remains in a fault state is less than the base station's reserved backup time t. s The probability P s As shown in the following formula:

[0064]

[0065] Based on the reliability assessment of the distribution network load points, the base station reliability index is calculated using the following formula:

[0066]

[0067] In the above formula, This represents the failure rate at the load point where the base station is located. Therefore, a minimum reliability specification for the base station can be specified. Therefore, the minimum backup power time t for base station energy storage at each load node under this index is derived. s .

[0068] Furthermore, since the minimum reserve capacity of a base station is related to the base station load level and the minimum backup power time, and considering the real-time fluctuations in base station load data, the real-time load rate of communication base stations can be predicted in the short term based on historical data of base station communication load in different regions. The inconsistent communication load levels of base stations in different regions result in significant differences in the minimum reserve capacity reserved for base station energy storage. This can be achieved by predicting the load rate characteristics of base stations in each node region. The load level of each base station is derived from equation (1). According to the base station's reserved backup time t s The lower limit of the capacity that the base station needs to reserve at time t is estimated as shown in the following formula:

[0069]

[0070] In the above formula, Reserve a lower limit for the capacity of base station energy storage in real time.

[0071] By treating the idle energy storage in massive 5G base stations as a flexible resource for power system dispatch, the main approach is to control the charging and discharging power of the base station energy storage batteries to absorb excess wind power during peak wind power output and discharge it during peak power output to reduce peak-shaving pressure, thereby achieving economic optimization of the distribution network dispatch and ultimately reducing wind curtailment and lowering system power generation costs. In essence, this is a problem of optimal charging and discharging control strategy for base station energy storage. Furthermore, if the power supply, load, and other quantities not controlled by the dispatch strategy within each decision cycle are approximated as constants, the model can be viewed as a sequential model and can be further transformed into a Markov decision process (MDP) model, thus enabling the solution using deep reinforcement learning algorithms.

[0072] The 5G base station energy storage and distribution network collaborative optimization scheduling method provided by this invention, which considers power supply reliability, utilizes reinforcement learning algorithms to achieve economic optimization scheduling of base station energy storage, enabling precise control of the charging and discharging power of each base station's energy storage in the distribution network. Specifically, it employs methods such as... Figure 3 The scheduling framework shown includes the environment model and the TD3 framework.

[0073] The environmental model includes the extraction of distribution network system state information and base station charge state information; the distribution network system state information mainly considers the node voltage U of node j at time t. j (t) and the output W(t) of the wind turbine at time t; In order to accurately schedule the energy storage of the base station, the agent must accurately perceive the state of charge of the base station. However, as the number of base stations increases, the input state dimension of the base station will increase, which greatly increases the optimization difficulty. Therefore, the real-time state of charge feature of the energy storage of the base station is encoded by the VAE model, and the dimension-reduced latent variables are used as the input of the agent in the TD3 algorithm, thereby enhancing the robustness of the algorithm and improving the learning performance of the agent.

[0074] In implementation, a VAE consists of two parts: an inference network and a generator network. The inference network is parameterized. of Probability distribution models encode input data into latent variables through variational inference. z p(z) is taken as a standard normal distribution; while the generative network uses the latent variable z to reconstruct the input data and generate p with parameter θ. θ (x|z) approximates the probability distribution. To ensure that the output data transformed from the latent variables is as equal as possible to the input data, log-maximum likelihood estimation of p is used. θ The parameters of (x|z), i.e. the optimization objective, are:

[0075]

[0076] Because the true posterior probability distribution p in the generative network θ (x|z) cannot be obtained directly, so VAE introduces a recognition model into the inference network. To replace the uncertain true posterior distribution p θ (x|z), and utilize the recognition model Approximating the true posterior probability distribution p θ (x|z). To make the true posterior distribution p θ (x|z) Approximate proximity recognition model The KL divergence is used to evaluate the similarity between two distributions, i.e., minimizing the divergence:

[0077]

[0078]

[0079] Visible div KL If ≥0 always holds true, then Then, we can use a deep learning network to maximize the likelihood probability logP. θ (X) requires maximizing the variational lower bound. The variational lower bound formula can be derived:

[0080]

[0081] In the first term of the above formula, Take it as a Gaussian distribution, that is p θ (z|x) follows a standard normal distribution, p θ (z|x)~N(0,1), therefore the KL divergence calculation formula is:

[0082]

[0083] In the last item, E qθ(z|x) [log(P θ [x|z)] is the log-likelihood of the posterior probability, which can be calculated through sampling:

[0084]

[0085] In the formula: P θ (x′ i |z i The distribution can be set as a Gaussian or Bernoulli distribution, and its mean and variance can be obtained through a neural network. P can then be calculated using the probability density formula. θ (x′ i |z i ).

[0086] In implementation, the TD3 algorithm uses a Critic network and an Actor network as its framework; its value network is updated by minimizing the difference between the evaluation value and the target value; and the policy network updates its network parameters by maximizing the cumulative expected return.

[0087] 1. Value Network

[0088] The Critic network is parameterized using a neural network. The value function is approximated by minimizing the Bellman residual, and the loss function L(θ) is established using the Monte Carlo sampling approximation. The parameters of the current Critic network are then updated by minimizing the loss function. The loss function is:

[0089]

[0090] In the formula: N is the number of training samples; y i The target Q value (i.e., the label); Q On This is the current Q value; These are the current Q network parameters.

[0091] During the strategy evaluation phase, y is calculated using the target Actor network and the Critic network. t ,Right now:

[0092]

[0093] In the formula: r t Q is the reward function; γ is the discount factor; Ta The target Q value; Target Q-network parameters.

[0094] Because Q-networks suffer from overestimation during training and large variance and inaccuracy in target estimates due to deterministic policies, the TD3 algorithm employs a dual-objective Q-network and target policy smoothing regularization to address these issues. The dual-objective Q-network constructs two identical target Q-networks, and the minimum value between the two is selected as the target value.

[0095]

[0096] In the formula: denoted as network parameters for two different target Q-networks; a′ represents the action with random noise.

[0097] Target policy smoothing regularization utilizes the actions surrounding the target action to calculate the target's Q-value, thus facilitating smooth estimation of the target. Therefore, in the next state s... t+1 The action is:

[0098] a′=μ′(s t+1 )+ε

[0099] In the formula: μ′ represents the target network policy; ε represents random noise.

[0100] To make the target action closer to the original action, a normal distribution is generally taken, and the sampling noise is truncated, i.e.: ε ~ clip(N(0,σ),-c,c); where: σ is the variance, and c is the upper and lower limits of the noise. This method can smooth the Q function, prevent the network update gradient from being too large, reduce the variance of the target estimation, improve the accuracy of the target estimation, and thus improve the stability of the network.

[0101] 2. Policy Network

[0102] The policy network maximizes the function J by optimizing the policy. β (μ), that is:

[0103]

[0104] In the formula: ρ β (s) represents the distribution of the current state under policy μ; Q μ (s,μ(s)) represents the value of the action according to policy μ in state s; J β (μ) represents the expected return that can be obtained.

[0105] Similarly, the Actor network is parameterized using a neural network, and its parameters are updated using gradient descent, i.e.:

[0106]

[0107] 3. Delayed updates

[0108] The policy network and target value network parameters are updated using a soft update method, namely:

[0109]

[0110] θ μ′ =τθμ+(1-τ)θ μ′

[0111] To reduce the error caused by updating the policy network before the Q-network has stabilized, the update of the Actor network parameters should be delayed. This means increasing the update frequency of the value network parameters and waiting as long as possible for the value network to converge before updating the policy network parameters, thus reducing accumulated error and consequently lowering variance.

[0112] Furthermore, since the agent continuously interacts with the active power distribution network environment to generate experience sample data and stores it in the experience replay pool to support the online training of the network, insufficient experience sample data can be generated in the early stages of training, leading to significant economic losses and system security issues. In practical engineering applications, to ensure the safety and economic benefits of base station energy storage participating in system scheduling in the early stages, the TD3+behavior cloning algorithm (TD3 with behavior cloning, TD3+BC) can be used for pre-training. A certain number of experience samples are imported into the agent during the early training stages, allowing the agent to learn from expert data, thus enabling the agent to achieve good results in the early stages of scheduling. To make the agent's actions as close as possible to the actions in the expert data, the TD3+behavior cloning algorithm adds a regularization term π(s)-a and λ to TD3, namely:

[0113]

[0114] In the formula: The state is then normalized, and the normalization process is as follows:

[0115]

[0116] In the formula: μ i σ represents expectation; i Indicates variance.

[0117] Specifically, in the 5G base station energy storage and distribution network collaborative optimization scheduling method for power supply reliability provided by this invention, the objective of the base station energy storage economic optimization scheduling based on the TD3 algorithm is to minimize the system operating cost. The system operating cost includes generator operating cost, base station energy storage battery charge / discharge discount cost, and base station energy storage leasing cost. Therefore, its objective function is configured as follows:

[0118] F = min(C) gen +C BES +C Lease )

[0119] Among them, C gen Cost of generating electricity by generator; C BES Cost of battery degradation during charging and discharging for base stations; C Lease The cost of leasing energy storage for base stations.

[0120] The power generation cost of the generator is as follows:

[0121]

[0122] In the formula: a i b i c iP is the power generation cost coefficient for generator set i; Git It generates electricity for the generator.

[0123] The degradation of energy storage batteries is a non-linear process affected by factors such as temperature, number of charge-discharge cycles, and depth of discharge. Frequent scheduling of base station energy storage will accelerate its lifespan degradation. The cost discount factor for energy storage batteries can be obtained from their degradation curves; therefore, the degradation loss cost of base station energy storage batteries is considered as follows:

[0124]

[0125] In the formula: C a The cost discount factor for energy storage batteries is N; N is the total number of energy storage units in the base station. Δt represents the real-time charging and discharging power of base station i, respectively; Δt is the unit time interval.

[0126] From the perspective of telecommunications operators, they primarily improve their economic efficiency by signing demand response cooperation agreements with distribution network operators. Distribution network operators pre-assess the minimum backup hours for each base station load point for the telecommunications operators' reference. Based on the dispatchable capacity assessment results, the telecommunications operators integrate the base station energy storage that can participate in distribution network dispatch and determine the real-time dispatchable capacity of the base station energy storage. Once a certain entry threshold is met, they can participate in electricity market bidding, thereby obtaining economic benefits. Therefore, the leasing cost of base station energy storage is considered as follows:

[0127]

[0128] In the formula: C b This is the rental cost coefficient; These are the upper and lower limits for real-time scheduling, respectively; ξ(t) is represented by (0,1), where 0 represents that the energy storage participates in the scheduling at time t, and 1 represents that the energy storage does not participate in the scheduling at time t.

[0129] Considering the similarities in the regions and load characteristic curves of some base stations, a pre-trained VAE model is used to perform feature encoding on the base station's energy storage SOC, thereby extracting key features and reducing computational complexity. The energy storage SOC of N base stations at time t is: [M 1,soc (t),···,M i,soc (t),···,M N,soc (t)] and the lower limit of reserved SOC capacity: [L 1,soc (t),···,L i,soc (t),···,L N,soc (t)] are encoded into M using an encoder respectively. m,soc (t), L m,soc (t), m The dimension is the state after encoding.

[0130] Therefore, the state space S of the intelligent agent can be defined as:

[0131] S = [M m,soc (t),L m,soc (t),U j (t),W(t),t]

[0132] Where: M m,soc (t), L m,soc (t) represents the VAE state code of the base station state; U j W(t) represents the node voltage of node j in the distribution network at time t; W(t) represents the output of the wind turbine; and t is a time stamp, represented by binary code.

[0133] Since the action space A represents the decision actions for model optimization, this invention uses the charging and discharging power of each base station's energy storage battery as the action. For N base station energy storage systems, the action space A can be defined as:

[0134] A=[a1(t),···,a i (t),···a N (t)]

[0135] In the formula: a i (t) represents the charging and discharging power of base station i at time t.

[0136] The reward R obtained by the intelligent agent from interacting with the environment is the sum of the generator cost, battery charging and discharging loss cost, and energy storage leasing cost.

[0137]

[0138] In the formula: N represents the amount of energy stored in the base station, and T is the scheduling period.

[0139] The penalty function is used to correct inappropriate actions by the agent. For base station energy storage devices, the agent's current action strategy may cause the base station's energy storage SOC state in the next state to exceed its allowed upper and lower limits. Therefore, it is necessary to set an over-limit penalty term:

[0140] P b =||a as -a re ||2

[0141] In the formula: a as Auxiliary vector for base station actions; a re This is the action vector of the agent under its current policy.

[0142] Based on the scheduling model established above, such as Figure 4As shown, the 5G base station energy storage and distribution network coordinated optimization scheduling method considering power supply reliability provided by the present invention includes the following steps:

[0143] S1: After each decision cycle, obtain the state information S = [M] corresponding to the current time t. m,soc (t),L m,soc (t),U j [(t),W(t),t]; where M m,soc (t) represents the SOC[M] stored in N base stations at time t. 1,soc (t),···,M i,soc (t),···,M N,soc The VAE state encoding of [t], L m,soc (t) represents the lower limit of the SOC (System-on-Chip) capacity of N base stations at time t [L]. 1,soc (t),···,L i,soc (t),···,L N,soc The VAE state encoding of [(t)], where m is the dimension of the encoded state; U j W(t) represents the node voltage of node j in the distribution network at time t; W(t) represents the output of the wind turbine at time t; t is the time marker.

[0144] S2: Obtain the acquired status information S = [M m,soc (t),L m,soc (t),U j The VAE model is input to the trained scheduling control model and outputs the corresponding scheduling control strategy. The scheduling control model is configured as a Markov decision process with the goal of minimizing the system operating cost. The VAE model is embedded into the TD3 algorithm and trained using VAE state codes as agents.

[0145] S3: According to the scheduling and control strategy, the distribution network executes the corresponding strategy actions to control the charging and discharging status of each 5G base station energy storage device.

[0146] To further verify the effectiveness of the 5G base station energy storage and distribution network collaborative optimization scheduling method considering power supply reliability provided by this invention, simulation experiments were conducted. Specifically, the IEEE 33-node distribution system was selected as the prototype for simulation calculations, and some adjustments were made based on it. For example... Figure 5As shown, a wind turbine generator with an installed capacity of 3MW is installed at node 14; 5G base stations with energy storage are set up at 16 nodes, including nodes 9-18 and 28-33, with 4 5G base stations under each node. The load of the base stations under each node is randomly selected, and the backup energy storage batteries for the base stations are lithium iron phosphate batteries that are used in a cascade manner. The equipment parameters of a single 5G base station are shown in the table below.

[0147] Table 1 Basic Parameters of 5G Base Station Equipment

[0148]

[0149] The output data for wind turbines and photovoltaic power generation systems are derived from Elia.be’s forecasts for the Aggregate Belgian Wind Farms and Belgium region from 01 / 06 / 2021 to 30 / 06 / 2021, and multiplied by an appropriate scaling factor to accommodate the capacity of the distribution system.

[0150] according to Figure 5 The distribution network topology and failure rate parameters of each component in the distribution network are shown in Table 1. The reliability of the load node with base station energy storage is evaluated.

[0151] Table 2 shows the load node reliability indicators (%) and minimum backup hours for base station energy storage.

[0152]

[0153] As shown in Table 2, the power distribution system is equipped with four tie lines. Load nodes 9, 12, 18, and 33, which are close to the network root node and the tie lines, have higher reliability, while load nodes 15 and 28-30, which are farther away from the tie lines, have lower reliability. The reliability indices of each load node differ significantly. Considering a base station availability index of 99.999%, the minimum backup time for each load node is calculated using the method in Section 2.2. Table 1 lists the reliability indices and corresponding backup power hours for load nodes containing 5G base station energy storage in this example. For high-reliability load nodes, their energy storage backup reserve time is shorter; for example, the minimum backup power times for nodes 12 and 33 are 1.68h and 1.70h, respectively. For low-reliability load nodes, their backup energy storage reserve time is longer; for example, the minimum backup power times for nodes 17 and 28 are 2.55h and 3.18h, respectively.

[0154] This experiment uses Tensorflow 2.0 as the framework and Pythoon 3.8 as the programming environment. The model was implemented on a machine equipped with an AMD Ryzen 7 4800H CPU @ 2.90GHz and an NVIDIA GeForce RTX 2060 graphics card.

[0155] First, for the State of Charge (SOC) of the energy storage of 64 base stations, this invention displays the state of charge of the energy storage of each base station in the form of a 100×64 pixel image, and constructs its VAE model using fully connected neural networks and convolutional neural networks respectively; as shown Figure 6 The VAE model pre-training results shown allow for a comparative analysis of the reconstruction performance of the two networks. The average error ratio of the fully connected neural network in decoding and reconstructing the image is approximately 18.6%, while the average error ratio of the convolutional neural network is approximately 3.2%, both showing smaller errors compared to the fully connected neural network. This demonstrates that the convolutional neural network constructs the probability distribution of the VAE inference and generation network models more accurately, enabling it to accurately and quickly extract features from high-dimensional state information and achieve dimensionality reduction, thus approximating the original image to the greatest extent possible.

[0156] Next, for training the scheduling control model, i.e., the TD3 network, the parameters θ of the policy network and the value network are first initialized. μ =θ μ′ , Initialize the experience pool; at each iteration time t, use the distribution network operating state and VAE charge state encoding as the input state s of the policy network. t The Actor network generates the energy storage actions of each base station based on the current state and adds noise 'a'. t =μ(s) t The distribution network environment executes the current control strategy and performs state transitions, generating the state s for the next time step. t+1 At this point, based on the current system environment, calculate the reward r at time t. t And feed it back to the agent, the agent will use the experience at time t {s} t ,a t ,r t ,s t+1} are stored in the experience pool. Simultaneously, m samples are drawn from the experience pool each time. The target policy network calculates the next action a. t+1 =μ′(s t+1 )+ε, using two target value networks to calculate the target value for m samples. And take y t =min{Q′1,Q′2} to prevent overestimation of the target value network. Update the value network parameters using the loss function L(θ). Considering delayed updates, the network parameters θ are updated every d steps using gradient descent. μ The above process is repeated cyclically to achieve adaptive learning of the model until the reward is maximized and the model converges. In this invention, a decoupled training mode between the VAE model and the TD3 algorithm is adopted, which can stabilize the input state of the TD3 algorithm and accelerate the convergence speed of the algorithm.

[0157] like Figure 7 As shown, the dispatch control model underwent continuous simulation training for 7 days. After approximately 2500 training rounds, the model converged and the corresponding training results were obtained. After approximately 2500 training rounds, the power distribution system was able to save approximately 950 yuan in system costs through the dispatch of 5G base stations, and the monthly economic benefits of the power distribution network operation were approximately 4000 yuan. Since the simulation system includes 64 base stations, the economic benefit per base station per month is approximately 62 yuan. In addition, to alleviate the overfitting of early experience caused by the small amount of data in the experience pool and the multiple sampling and learning of data in the early stage of exploration, the network reset technique was used in the 250th and 500th training rounds of the above model simulation. The results show that the network reset can bring positive learning effects and provide the possibility of generating higher quality data for future updates.

[0158] Optimizing the charging and discharging strategy of base station energy storage using a trained scheduling model can improve the economic efficiency of system operation, optimize the voltage levels of each base station load node, reduce real-time voltage deviation, and improve the operational stability of the distribution network. For example... Figure 8 As shown, the utilization of energy storage to absorb wind power is mainly concentrated between 00:00 and 08:00 each day, while during other times, wind power tends to operate at full capacity due to limitations in the distribution network load. Figure 9 As shown in the shaded area, the daily absorption of renewable energy generation is approximately 0.5MW, representing an increase of about 2% in renewable energy absorption compared to before base station energy storage was involved in dispatching. For a single base station energy storage system, its daily flexibly dispatchable power is around 8kW. Currently, my country has approximately 2.3 million 5G base stations. Assuming 50% of these can participate in system dispatching, if this method can be effectively implemented, the daily absorption of renewable energy could reach approximately 8.6GW. Therefore, base station energy storage has enormous potential in improving renewable energy absorption.

[0159] like Figure 9 As shown, during the period from 10:00 to 12:00, the distribution network load level is high, resulting in a lower overall voltage level. At this time, the participation of base station energy storage in dispatching can alleviate the pressure of peak power consumption to some extent and raise the voltage level of base station load nodes. Figure 8 Analysis shows that the overall voltage deviation of the system decreased by 4.8% at 10:00 and by 4.0% at 11:00. This indicates that the improvement in the voltage level at the base station load points also led to an overall increase in the voltage levels of all nodes in the power grid.

[0160] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for coordinated optimization scheduling of 5G base station energy storage and distribution network considering power supply reliability, characterized in that, Includes the following steps: S1: After each decision cycle, obtain the state information S = [M] corresponding to the current time t. m,soc (t),L m,soc (t),U j [(t),W(t),t]; where M m,soc (t) represents the SOC[M] stored in N base stations at time t. 1,soc (t),···,M i,soc (t),···,M N,soc The VAE state encoding of [t], L m,soc (t) represents the lower limit of the energy storage SOC capacity of N base stations at time t [L] 1,soc (t),···,L i,soc (t),···,L N,soc The VAE state encoding of [(t)], where m is the dimension of the encoded state; U j W(t) represents the node voltage of node j in the distribution network at time t; W(t) represents the output of the wind turbine at time t. t For time stamps; S2: Obtain the acquired status information S = [M m,soc (t),L m,soc (t),U j [(t), W(t), t] are input to the trained scheduling control model and output the corresponding scheduling control strategy; wherein, the scheduling control model is configured as a Markov decision process with the goal of minimizing the system operating cost, and the VAE model is embedded into the TD3 algorithm, and the VAE state code is used as the agent for training; S3: According to the scheduling and control strategy, the distribution network executes the corresponding strategy actions to control the charging and discharging status of each 5G base station energy storage device.

2. The 5G base station energy storage and distribution network collaborative optimization scheduling method considering power supply reliability as described in claim 1, characterized in that, Base station energy storage SOC capacity lower limit at time t for: Among them, the load level of each base station The load rate characteristics of base stations in each load node area are as follows The minimum backup power time t for base station energy storage at each load node is determined. s Minimum availability index of base stations Decide, The base station reserves a backup time t for the failure rate of the load point where the base station is located. s The probability P s With the probability distribution of power outage time Q s (t) correlation, i.e. Power outage time probability distribution Q s (t) is generated by fitting historical data of power outage time in the distribution network.

3. The 5G base station energy storage and distribution network collaborative optimization scheduling method considering power supply reliability as described in claim 2, characterized in that, A VAE model is used to encode the base station's state of charge. The VAE model includes an inference network and a generator network; wherein the inference network is parameterized. of The probability distribution model encodes the input data into a latent variable z using variational inference, with p(z) following a standard normal distribution. The generative network uses the latent variable z to reconstruct the input data and generate p with parameter θ. θ (xz) approximate probability distribution.

4. The 5G base station energy storage and distribution network collaborative optimization scheduling method considering power supply reliability as described in claim 3, characterized in that, With the goal of minimizing system operating costs, the objective function is configured as follows: F=min(C gen +C BES +C Lease ) Among them, C gen For the cost of generating electricity by the generator, a i b i c i Let P be the power generation cost coefficient of generator unit i. Git To generate electricity for the generator; C BES C represents the cost of charge / discharge degradation of base station energy storage batteries. a Where N is the cost discount factor for energy storage batteries, and N is the total number of energy storage units in the base station. C represents the real-time charging and discharging power of base station i, respectively, where Δt is the unit time interval; Lease For the cost of base station energy storage leasing, C b This is the rental cost coefficient; These are the upper and lower limits for real-time scheduling, respectively; ξ(t) is represented by (0,1), where 0 represents that the energy storage participates in the scheduling at time t, and 1 represents that the energy storage does not participate in the scheduling at time t.

5. The 5G base station energy storage and distribution network collaborative optimization scheduling method considering power supply reliability as described in claim 4, characterized in that, When training an agent based on the TD3 algorithm, the TD3 algorithm uses a Critic network and an Actor network as its framework. Furthermore, it parameterizes the Critic network using a neural network, approximates the value function by minimizing the Bellman residual, establishes the loss function L(θ) using the Monte Carlo sampling approximation, and then updates the parameters of the current Critic network by minimizing the loss function. The loss function L(θ) is specifically: N is the number of training samples; y i The target Q value; Q On This is the current Q value; Let these be the current Q-network parameters; and let's parameterize the Actor network using a neural network and update its parameters using gradient descent, i.e.

6. The 5G base station energy storage and distribution network collaborative optimization scheduling method considering power supply reliability as described in claim 5, characterized in that, The Critic network employs two identical target Q-networks, and the minimum value between the two is selected as the target value. in, Here are the network parameters for two different target Q-networks, where a′ represents the action with random noise; and a′ = μ′(s t+1 )+ε,μ′ represents the target network strategy, and ε represents normally distributed random noise, which is obtained by truncating the sampled noise.

7. The 5G base station energy storage and distribution network collaborative optimization scheduling method considering power supply reliability as described in claim 6, characterized in that, During the training of the agent based on the TD3 algorithm, the parameters of the policy network and the target value network are also updated using a soft update method, namely: i μ′ =tθ μ +(1-τ)θ μ′ 。 8. The 5G base station energy storage and distribution network collaborative optimization scheduling method considering power supply reliability as described in claim 7, characterized in that, During the training of the agent based on the TD3 algorithm, the reward obtained by the agent through interaction with the environment is the sum of generator cost, battery charging and discharging loss cost, and energy storage leasing cost: Where N represents the amount of energy stored in the base station, and T is the scheduling period.

9. The 5G base station energy storage and distribution network collaborative optimization scheduling method considering power supply reliability as described in claim 8, characterized in that, During the training of the agent based on the TD3 algorithm, the penalty function used to correct the agent's inappropriate actions is configured as follows: P b =||a as -a re ||2 Among them, a as Auxiliary vector for base station actions; a re This is the action vector of the agent under its current policy.

10. The 5G base station energy storage and distribution network collaborative optimization scheduling method considering power supply reliability as described in claim 9, characterized in that, When training an agent based on the TD3 algorithm, a behavior cloning algorithm is also used for pre-training: in, β represents the action strategy adopted by the agent when exploring the environment during training, ρ β (s) represents the distribution of the current state under policy μ; Q μ (s,μ(s)) represents the value of the action according to policy μ in state s, J β (μ) represents the expected return that can be obtained. μ i σ represents expectation; i Indicates variance.