Antenna and bandwidth resource planning methods to ensure power grid observability

By combining effective capacity theory, deep reinforcement learning, and simulated annealing algorithm, the power grid antenna and bandwidth resources under MIMO technology are optimized, solving the real-time and channel fading problems in the communication environment of smart grid systems and achieving efficient utilization of bandwidth and antenna resources.

CN116647920BActive Publication Date: 2026-05-26SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SOUTHEAST UNIV
Filing Date
2023-05-30
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

The existing power grid communication environment is insufficient to meet the real-time requirements of PMU data for smart grid systems, and the channel fading effect of wireless communication technology has time-varying characteristics. Therefore, it is necessary to optimize the topology and communication parameters to improve the performance indicators of the power system.

Method used

Combining effective capacity theory, deep reinforcement learning algorithm, and simulated annealing algorithm, this paper proposes an antenna and bandwidth resource planning method to ensure power grid observability under MIMO technology. The method uses the deep reinforcement learning algorithm DQN for bandwidth allocation, the simulated annealing algorithm for antenna allocation, and the alternating optimization algorithm for joint planning.

Benefits of technology

While achieving the same observability metrics, the total bandwidth and number of antennas of the system should be minimized to meet the planning requirements of the smart grid system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure QLYQS_9
    Figure QLYQS_9
  • Figure QLYQS_15
    Figure QLYQS_15
  • Figure QLYQS_16
    Figure QLYQS_16
Patent Text Reader

Abstract

This invention discloses a method for planning antenna and bandwidth resources under MIMO technology, prioritizing smart grid observability while simultaneously satisfying relevant quality of service (QoS) indicators. Observability is a prerequisite for the normal operation of a smart grid, and smart grids also have high real-time requirements for measurement data generated by PMU devices. This invention is based on deep reinforcement learning algorithms, simulated annealing, and the bisection method, and introduces effective capacity theory to measure the statistical QoS requirements corresponding to wireless fading channels, addressing strict latency constraints. The problem model is planned as follows: under constraints of total bandwidth and total number of antennas, an alternating optimization algorithm allocates the total bandwidth and total number of antennas to the transmitters corresponding to each PMU device, generating a resource allocation scheme with the smart grid observability index as the objective function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of mobile communication system technology and power grid control system, and in particular to an antenna and bandwidth resource planning method for ensuring power grid observability under MIMO technology. Background Technology

[0002] The rapid development of 5G communication technology has yielded excellent performance in terms of bandwidth, large-scale IoT, and latency, making it suitable for solving the "last mile" information interconnection problem in the ubiquitous power IoT. However, as the scale of the ubiquitous power IoT increases, so does the demand for communication. Research on optimizing topology and communication parameters to improve power system performance is still limited. The integration of smart grid systems and wireless communication technologies is a future trend. Power Management Units (PMUs) measure the real-time operating status of smart grid systems to ensure their normal and stable operation. However, smart grid systems have extremely high real-time requirements for PMU data, which the existing power grid communication environment struggles to meet. Furthermore, the channel fading effect of wireless communication technologies has time-varying characteristics, necessitating cross-layer analysis of indicators such as latency violation probability, latency, and channel capacity. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide an antenna and bandwidth resource planning method for ensuring power grid observability under MIMO technology, which can use the least amount of bandwidth and antenna resources to achieve the same observability index.

[0004] To address the aforementioned technical problems, this invention combines grid observability, effective capacity theory, the bisection method, deep reinforcement learning algorithms, and simulated annealing algorithms to propose an antenna and bandwidth resource planning method for ensuring grid observability under MIMO technology, comprising the following steps:

[0005] (1) Determine the expression for the effective capacity and the probability that the data arrives on time;

[0006] (2) Define the objective function and analyze the variables to be optimized in the objective function;

[0007] (3) Generate an initial scheme for bandwidth allocation and antenna allocation;

[0008] (4) Set the relevant parameters of the deep reinforcement learning algorithm DQN, including learning rate α, discount factor γ, and exploration probability ∈;

[0009] (5) Set the relevant parameters for the simulated annealing algorithm, including the initial temperature T. s Termination temperature T f Temperature decay coefficient β;

[0010] (6) Calculate the numerical value of the objective function based on the initial scheme of bandwidth allocation and antenna allocation;

[0011] (7) Use the deep reinforcement learning algorithm DQN to allocate bandwidth, use the simulated annealing algorithm to allocate antennas, and use the alternating optimization algorithm. That is, repeat steps (7) to (8) until the condition for the end of the loop is met.

[0012] (10) Output the corresponding bandwidth allocation and antenna allocation scheme.

[0013] Preferably, step (1) includes the following specific steps:

[0014] (11) Analyze the transmission model, assuming the device PMU k Each base station is configured with n Tk and n R One antenna, then PMU k Data transmission on the k-th channel between the base station and the ground station is represented as follows:

[0015] y k =H k x k +n k k = 1, 2, ... K

[0016] Where K represents the number of channels, x k For n Tk ×1D PMU k The emitted signal vector, y k For n R A 1×1 dimensional receive vector, H k For PMU k The channel matrix between the base station and the ground station has a dimension of n. R ×n Tk In addition, n k n R A 1×1 dimensional Gaussian noise vector, where the variance of each element is σ. 2 Equipment PMU k The average signal-to-noise ratio is defined as:

[0017]

[0018] Where P k Indicates device PMU k The upper bound of the average power. The normalized input covariance matrix of the k-th channel is defined as:

[0019]

[0020] in x represents k The conjugate transpose of B kIndicates allocation to PMU k The bandwidth value. Based on the above definition, assume the input covariance matrix is... When the Shannon capacity is reached, the Shannon capacity of the channel under MIMO technology is expressed as follows:

[0021]

[0022] (12) Based on the definition of effective capacity and the properties of MIMO transmission technology, the expression for effective capacity is as follows:

[0023]

[0024]

[0025] In the above formula, θ k The QoS index is represented by T, the coherence time is represented by Γ(·), and f represents the Gamma function. k =min(n R ,n Tk ), and b k =max(n R ,n Tk )-min(n R ,n Tk In addition, G k It is f k ×f k A dimensional Hankel matrix, the elements of which are:

[0026]

[0027] (13) Calculate PMU based on effective capacity theory and queuing theory. k The probability that the generated measurement data arrives at the control center on time is as follows:

[0028]

[0029] Among them, D max This indicates the maximum time delay limit.

[0030] Preferably, step (2) includes the following specific steps:

[0031] (21) Assume that the smart grid system has N buses. Define the connection matrix of the smart grid topology as L. The elements of this matrix L m,n Defined as: if bus m is connected to bus n, or m = n, then L m,n =1, otherwise L m,n =0, m,n =1,2,…,N. Define the PMU device installation vector X = {x1,x2,…,x...} N}T If a PMU device is deployed on bus i, then x i =1, otherwise x i =0. Then define the diagonal probability matrix Λ P for:

[0032] Λ P =diag{P1,P2,…,P N}

[0033] Among them, if PMU k If installed on bus i, then P i =p k Otherwise P i =0. This yields the expected power grid observability vector. Represented as:

[0034]

[0035] Observability vector for:

[0036]

[0037] Among them Λ Q The diagonal communication constraint matrix is ​​defined as follows:

[0038] Λ Q =diag{Q1,Q2,…,Q N}

[0039] Q i It is a binary random variable, and its probability distribution is as follows:

[0040] Pr{Q i =1}=P i

[0041] Pr{Q i =0}=1-P i

[0042] Observability redundancy (OR) is defined as the expected redundancy of PMU devices that observe all buses in a smart grid system, i.e.:

[0043]

[0044] Observability sensitivity OS is defined as the expected power grid observability vector. The smallest value among them. The observability probability OP is defined as the observability vector. The probability that all elements in the set are greater than or equal to a certain set threshold.

[0045] (22) The objective function is set as the observability redundancy OR, observability sensitivity OS, or observability probability OP defined in step (21);

[0046] (32) Analyze the variables to be optimized in the objective function. First, according to the definitions of OR, OS, and OP, the objective function is composed of p k The values ​​of k = 1, 2, ..., K are calculated, while p k For θ k B k and n Tk The objective function is a function of the PMU device. Therefore, the variables to be optimized in the objective function are the number of antennas, bandwidth, and QoS index of each PMU device.

[0047] Preferably, step (6) includes the following specific steps:

[0048] (61) Assume the device PMU k by The constant rate generates measurement data, and the optimal θ k The corresponding effective capacity should be equal to the data generation rate. Right now

[0049]

[0050] (62) Combining the initial scheme of bandwidth allocation and antenna configuration, and based on the equation in step (61), solve θ using the bisection method. k ;

[0051] (63) Based on θ obtained in step (62) k Calculate p k ;

[0052] (64) Repeat steps (62) to (63) to find θ and p for all PMU devices;

[0053] (65) Calculate the values ​​of the objective function OR, OS, or OP based on the results of step (64);

[0054] Preferably, step (7) includes the following specific steps:

[0055] (71) Define the state space and action space of the agent in the DQN algorithm. Let f D The numerical value of the objective function is used to define the state s of the agent. D For a 1×(K+1) dimensional vector, its expression is as follows:

[0056] s D =[n T1 ,n T2 ,…,n TK ,fD ]

[0057] The set of all possible states of the agent constitutes the state space S. Let Δu represent the unit change in bandwidth, then the agent's action design is as follows:

[0058] a1=((K-1)Δu,-Δu,…,-Δu)

[0059] a2=(-Δu,(K-1)Δu,…,-Δu)

[0060]

[0061] a K =(-Δu,-Δu,…,(K-1)Δu)

[0062] The set consisting of the above K actions is called the action space A.

[0063] (72) Set up the neural network for the DQN algorithm. The input of the neural network is the state of the agent, and the output is the value of each action in that state. The neural network is designed as a three-layer neural network containing an input layer, a hidden layer and an output layer. According to the setting of the state space S and the action space A in step (71), the number of neurons in the input layer is (K+1) and the number of neurons in the output layer is K.

[0064] (73) Set the reward function of this DQN algorithm as follows:

[0065] r = δ(f D,t+1 -f D,t )

[0066] Where f D,t f represents the numerical value of the objective function in the current state; D,t+1 This represents the value of the objective function after the action is performed, i.e., the value of the objective function in the next state. The parameter δ is the coefficient of the reward function, and the value of this coefficient is greater than 0.

[0067] (74) The exploration / exploitation strategy of the DQN algorithm is set as follows: the agent selects a random action with a probability of ∈ and the agent selects a greedy action with a probability of 1-∈, that is, selects the action with the highest value in this state.

[0068] (75) Combining the relevant settings of the DQN algorithm in steps (71) to (74), execute the DQN algorithm to allocate bandwidth to the smart grid system.

[0069] The beneficial effects of this invention are as follows: This invention first combines the channel uncertainty of wireless communication with the observability of the power grid using effective capacity theory and queuing theory, and proposes an algorithm that can jointly plan the bandwidth and antenna resources of a smart grid system by employing the bisection method, deep reinforcement learning algorithm, and simulated annealing algorithm. This algorithm can ensure that the power grid minimizes the total bandwidth and total number of antennas of the system while achieving the same observability index, effectively meeting the planning requirements of smart grid systems. Detailed Implementation

[0070] This embodiment uses the IEEE 14 bus power system, which is widely used as a standard test case, to verify the invention. It assumes that the channel fading mode is block fading and Rayleigh fading, and that OFDMA technology is used for multiple access between each PMU device and the base station. It also assumes that all PMU devices generate measurement data at a rate of 60 kbps, the coherence time T is 0.005 s, and the maximum delay limit for data transmission from the PMU to the control center is 10 ms. The PMU installation vector is expressed as:

[0071] X={0,1,0,1,1,1,1,1,1,0,1,0,1,0} T

[0072] The average signal-to-noise ratio of each PMU device is set between 10dB and 15dB.

[0073] The planning process is as follows, including the following steps:

[0074] (1) Cross-layer statistical time delay analysis is performed using effective capacity theory, which combines communication delay and power grid observability indicators. Using effective capacity theory and queuing theory, the effective capacity value and the probability of PMU data being transmitted to the control center within the maximum time delay limit are calculated, and the objective function value is calculated based on this.

[0075] (2) A bandwidth allocation scheme is generated based on the deep reinforcement learning algorithm DQN and the binary search method. During the optimization of the objective function, the value of the objective function under the current bandwidth allocation scheme is calculated in conjunction with step (1). The optimal bandwidth allocation scheme is obtained by iteratively applying the DQN algorithm.

[0076] (3) Antenna configuration schemes are generated based on simulated annealing and the bisection method. During the optimization of the objective function, the numerical value of the objective function under the current antenna configuration scheme is calculated in step (1), and the Metropolis algorithm is used to determine whether to adopt a new antenna configuration scheme. Through repeated iterations using the simulated annealing algorithm, the optimal antenna configuration scheme is obtained.

[0077] (4) Execute the alternating optimization algorithm, that is, repeat steps (2) to (3) to jointly optimize the bandwidth and antenna resources. When step (3) cannot find a better antenna configuration scheme, the loop ends.

[0078] In step (1), cross-layer statistical time delay analysis is performed using effective capacity theory, and the objective function is calculated. Specifically, the steps include:

[0079] (11) Analyze the transmission model, assuming the device PMU k Each base station is configured with n Tk and n R One antenna, then PMU k Data transmission on the k-th channel between the base station and the ground station is represented as follows:

[0080] y k =H k x k +n k k = 1, 2, ... K

[0081] Where, x k For n Tk ×1D PMU k The emitted signal vector, y k For n R A 1×1 dimensional receive vector, H k For PMU k The channel matrix between the base station and the ground station has a dimension of n. R ×n Tk In addition, n k n R A 1×1 dimensional Gaussian noise vector, where the variance of each element is σ. 2 Equipment PMU k The average signal-to-noise ratio is defined as:

[0082]

[0083] Where P k Indicates device PMU k The upper bound of the average power. The normalized input covariance matrix of the k-th channel is defined as:

[0084]

[0085] in x represents k The conjugate transpose of B k Indicates allocation to PMU k The bandwidth value. Based on the above definition, assume the input covariance matrix is... When the Shannon capacity is reached, the Shannon capacity of the channel under MIMO technology is expressed as follows:

[0086]

[0087] Based on the definition of effective capacity and the properties of MIMO transmission technology, the expression for effective capacity is as follows:

[0088]

[0089] In the above formula, Γ(·) represents the Gamma function, f k =min(n R ,n Tk ), and b k =max(n R ,n Tk )-min(n R ,n Tk In addition, G k It is f k ×f k A dimensional Hankel matrix, the elements of which are:

[0090]

[0091] Based on effective capacity theory and queuing theory, calculate PMU. k The probability that the generated measurement data arrives at the control center on time is as follows:

[0092]

[0093] Among them, D max This indicates the maximum time delay limit.

[0094] (12) Assume that the smart grid system has N buses. Define the connection matrix of the smart grid topology as L. The elements of this matrix L m,n Defined as: if m = n or bus m is connected to bus n, then L m,n =1, otherwise L m,n =0, m,n =1,2,…,N. Define the PMU device installation vector x = {x1,x2,…,x...} N} T If a PMU device is deployed on bus i, then x i =1, otherwise x i =0. Then define the diagonal probability matrix Λ P for:

[0095] Λ P =diag{P1,P2,…,P N}

[0096] Among them, if PMU k If installed on bus i, then P i =p k Otherwise P i =0. This yields the expected power grid observability vector. Represented as:

[0097]

[0098] Observability vector for:

[0099]

[0100] Among them Λ Q The diagonal communication constraint matrix is ​​defined as follows:

[0101] Λ Q =diag{Q1,Q2,…,Q N}

[0102] Q i It is a binary random variable, and its probability distribution is as follows:

[0103] Pr{Q i =1}=P i

[0104] Pr{Q i =0}=1-P i

[0105] Observability redundancy (OR) is defined as the expected redundancy of PMU devices that observe all buses in a smart grid system, i.e.:

[0106]

[0107] Observability sensitivity OS is defined as the expected power grid observability vector. The smallest value among them. The observability probability OP is defined as the observability vector. The objective function is defined as the probability that all elements in the objective function are greater than or equal to a certain set threshold. The objective function is set as the OR, OS, or OP size, and the variables to be optimized in the objective function are the number of antennas, bandwidth, and QoS index of each PMU device.

[0108] (13) Assume the device PMU k by The constant rate generates measurement data, and the optimal θ k The corresponding effective capacity should be equal to the data generation rate. Right now

[0109]

[0110] Combining bandwidth allocation and antenna configuration, θ is solved using the bisection method based on this equation. k And further calculate the probability p that the data arrives on time. k After finding θ and p for all PMU devices, calculate the objective function OR, OS or OP according to the definition of observability index in step (12).

[0111] Step (2) generates a bandwidth allocation scheme based on the deep reinforcement learning algorithm DQN and the binary search method, specifically including the following steps:

[0112] (21) Define the state space and action space of the DQN algorithm. Let f D The numerical value of the objective function is used to define the state s of the agent. D For a 1×(K+1) dimensional vector, its expression is as follows:

[0113] s D =[n T1 ,n T2 ,…,n TK ,f D ]

[0114] The set of all possible states of the agent constitutes the state space S. Let Δu represent the unit change in bandwidth, then the agent's action design is as follows:

[0115] a1=((K-1)Δu,-Δu,…,-Δu)

[0116] a2=(-Δu,(K-1)Δu,…,-Δu)

[0117]

[0118] a K =(-Δu,-Δu,…,(K-1)Δu)

[0119] The set consisting of the above K actions is called the action space A.

[0120] (22) Set up the neural network for the DQN algorithm. The input of the neural network is the state of the agent, and the output is the value of each action in that state. The structure of the neural network is designed as a three-layer neural network containing an input layer, a hidden layer and an output layer. According to the setting of the state space S and the action space A in step (21), the number of neurons in the input layer is (K+1) and the number of neurons in the output layer is K.

[0121] (23) Set the reward function of this DQN algorithm as follows:

[0122] r = δ(f D,t+1 -f D,t )

[0123] Where f D,t f represents the numerical value of the objective function in the current state; D,t+1 This represents the value of the objective function after the action is performed, i.e., the value of the objective function in the next state. The parameter δ is the coefficient of the reward function, and the value of this coefficient is greater than 0.

[0124] (24) The exploration / exploitation strategy of the DQN algorithm is set as follows: the agent selects a random action with a probability of ∈ and the agent selects a greedy action with a probability of 1-∈, that is, selects the action with the highest value in this state.

[0125] (25) Combining the relevant settings of the DQN algorithm in steps (21) to (24), execute the DQN algorithm to allocate bandwidth to the smart grid system.

[0126] Step (3) generates an antenna configuration scheme based on the simulated annealing algorithm and the binary search method, specifically including the following steps:

[0127] (31) The initial input to the simulated annealing algorithm is the current bandwidth allocation and antenna configuration scheme, denoted by [B1, B2, ..., B K ] and [n T1 ,n T2 ,…,n TK Using the dichotomy method and combining it with the following expression:

[0128]

[0129] Calculate the optimal QoS index θ for each PMU device. * Then, calculate the value of the objective function under the initial input, denoted as H.

[0130] (32) A new antenna configuration scheme is randomly generated according to certain rules, denoted as [n′]. T1 ,n′ T2 ,…,n′ TK This ensures that the sum of the number of antennas in each PMU device equals the given total number of antennas.

[0131] (33) Calculate the optimal QoS index θ for each PMU device again using the bisection method. * Then, the objective function value corresponding to the new antenna configuration scheme is calculated, denoted as H′.

[0132] (34) Determine whether to adopt the new antenna configuration scheme according to the Metropolis algorithm. Let ΔE = H′ - H. If ΔE > 0, then accept the new solution [n′]. T1 ,n′ T2 ,…,n′ TK ], H = H ′If ΔE < 0, then with probability Accept the new solution, i.e., [n T1 ,n T2 ,…,n TK ]=[n′ T1 ,n′ T2 ,…,n w TK ], H = H ′ .

[0133] (35) Repeat steps (32) to (34) for a total of M iterations.

[0134] (36) Update the temperature by letting T = βT, where β is the temperature decay coefficient.

[0135] (37) Repeat steps (35) to (36) until the termination temperature is reached.

[0136] Although the present invention has been illustrated and described with reference to preferred embodiments, those skilled in the art should understand that various changes and modifications can be made to the present invention without departing from the scope defined by the claims.

Claims

1. A method for planning antenna and bandwidth resources to ensure power grid observability, characterized in that, Includes the following steps: Determine the expression for effective capacity and the probability of data arriving on time; Establish the objective function and the variables to be optimized within the objective function; Generate an initial scheme for bandwidth allocation and antenna allocation; Set the relevant parameters for the Deep Reinforcement Learning algorithm DQN, including the learning rate. Discount factor Exploring Probability ; Set the relevant parameters for the simulated annealing algorithm, including the initial temperature. Termination temperature Temperature decay coefficient ; The objective function is calculated based on the initial scheme of bandwidth allocation and antenna allocation. Bandwidth allocation is performed using the deep reinforcement learning algorithm DQN, and the specific steps include: Defining the state space and action space of the agent in the DQN algorithm: The state of the agent in the deep reinforcement learning algorithm DQN For one A dimensional vector, its expression is as follows: ; This represents the numerical value of the objective function. The state space is the set of all possible states of an agent, representing the number of channels. ; Indicates the first Each PMU is configured with an antenna, k=1,2,…K; The action design of the intelligent agent is as follows: ; ; ; ; From the above The set of actions is called the action space. , This represents the unit change in bandwidth value; The neural network for the DQN algorithm is defined as follows: the input to the neural network is the agent's state, and the output is the value of each action in that state; the neural network contains an input layer, hidden layers, and an output layer, based on the state space. and action space The configuration specifies that the number of neurons in the input layer is [number missing]. The number of neurons in the output layer is [number missing]. indivual; The reward function for this DQN algorithm is defined as follows: ; in This represents the numerical value of the objective function in the current state; This represents the value of the objective function after the action is performed, i.e., the value of the objective function in the next state; parameters This is the coefficient of the reward function, and its value is greater than 0; The exploration / exploitation strategy of this DQN algorithm is set as follows: the agent uses... The agent selects random actions based on probability, and the agent... The probability-based greedy action is to choose the current state. Take the most valuable action; Antenna allocation is performed using a simulated annealing algorithm; Bandwidth and antenna allocation are performed using an alternating optimization algorithm until the loop termination condition is met; Output the corresponding bandwidth allocation and antenna allocation scheme.

2. The antenna and bandwidth resource planning method for ensuring power grid observability according to claim 1, characterized in that, The expression for the effective capacity is as follows: ; In the above formula, This represents the QoS index, where T represents the coherence time. and Representing the equipment Antennas configured with base stations Indicates assignment to The bandwidth value, This represents the Shannon capacity of the channel under MIMO technology. for Channel matrix between and base stations Indicates equipment The average signal-to-noise ratio, Represents the Gamma function. ,and ;also, yes A dimensional Hankel matrix, the elements of which are: ; The probability that the generated measurement data arrives at the control center on time is as follows: ; in, This indicates the maximum time delay limit.

3. The antenna and bandwidth resource planning method for ensuring power grid observability according to claim 2, characterized in that, The objective function sets the observability redundancy (OR), observability sensitivity (OS), or observability probability (OP). Observability redundancy (OR) is defined as the expected redundancy of PMU devices that observe all buses in a smart grid system. Observability sensitivity OS is defined as the expected power grid observability vector. The smallest value in the observability vector; the observability probability OP is defined as the observability vector. The probability that all elements in the set are greater than or equal to a certain set threshold; The variables to be optimized in the objective function are the number of antennas, bandwidth, and QoS index of each PMU device.

4. The antenna and bandwidth resource planning method for ensuring power grid observability according to claim 3, characterized in that, The specific steps for calculating the objective function based on the initial bandwidth and antenna allocation schemes include: Based on the initial scheme of bandwidth allocation and antenna configuration, the bisection method is used to solve the problem. ; get ; Indicates equipment A constant rate at which measurement data is generated; according to calculate ; Repeat the above steps to find the corresponding PMU devices. and ; according to and Calculate the values ​​of the objective function OR, OS, or OP.