A beam scheduling method and system based on peak information age optimization in millimeter-wave communication networks

By constructing an information age update model in millimeter-wave communication networks and transforming the beam scheduling problem into a Markov decision process, a deep reinforcement learning algorithm is used to generate the optimal beam scheduling scheme. This solves the problem of insufficient information timeliness under dynamic channel conditions and improves the real-time response capability of the network.

CN121037863BActive Publication Date: 2026-03-06GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511273524.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-08
Publication Date
2026-03-06
Estimated Expiration
2045-09-08

AI Technical Summary

Technical Problem

When millimeter-wave communication networks face dynamic channel conditions and multi-user scheduling requirements, existing scheduling methods struggle to optimize information timeliness while ensuring transmission reliability, leading to extended information update intervals and impacting the system's real-time response capabilities.

Method used

An information age update model is constructed and the beam scheduling problem is transformed into a Markov decision process. Deep reinforcement learning algorithms, such as the actor-critic algorithm, are used to generate the optimal beam scheduling scheme to minimize the peak information age.

Benefits of technology

It significantly improves the timeliness of network information, adapts to complex and ever-changing network environments, and provides reliable support for application scenarios with high real-time requirements such as industrial control and intelligent transportation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121037863B_ABST
    Figure CN121037863B_ABST
Patent Text Reader

Abstract

This invention relates to a beam scheduling method and system based on peak information age optimization in millimeter-wave communication networks. The method includes: acquiring environmental parameters and network parameters of the millimeter-wave network; constructing an information age update model based on the environmental parameters and network parameters; constructing a beam scheduling problem based on the information age update model; transforming the beam scheduling problem into a Markov decision process; solving the Markov decision process using a deep reinforcement learning algorithm to obtain the beam scheduling scheme with the minimum peak information age. This invention can automatically select the optimal beam scheduling scheme based on dynamic channel conditions and network load, effectively reducing the peak information age and improving the timeliness of network information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of millimeter-wave communication technology, and in particular to a beam scheduling method and system based on peak information age optimization in millimeter-wave communication networks. Background Technology

[0002] In recent years, millimeter-wave communication technology has been widely used in industrial automation, short-range mobile communication, and satellite communication due to its significant advantages of high bandwidth and high speed. However, in practical deployments, this technology faces technical bottlenecks such as severe signal attenuation, limited transmission distance, and sensitivity to obstacles. These factors lead to limited network coverage and reduced data transmission reliability. Especially in industrial control scenarios that require continuous updates to system status information, these problems significantly extend the information update time interval, severely impacting the system's real-time response capability. Current mainstream scheduling algorithms perform poorly in dealing with dynamically changing channel conditions and complex multi-user scheduling requirements, making it difficult to optimize information timeliness while ensuring transmission reliability. Therefore, developing a novel scheduling method that can effectively improve the timeliness of millimeter-wave network information has significant theoretical and practical value.

[0003] Currently, directional beamforming and multicast transmission are commonly used communication architectures in millimeter-wave network scheduling. Directional beamforming suffers from limited coverage, while the fixed scheduling mode of multicast transmission cannot adapt to the optimal transmission requirements under dynamic channel conditions. Most existing scheduling methods rely on static parameter configuration, failing to fully consider the impact of channel dynamics and multi-user interference on scheduling performance. Summary of the Invention

[0004] The purpose of this invention is to provide a beam scheduling method and system based on peak information age optimization in millimeter-wave communication networks. It is applicable to application scenarios with high requirements for information timeliness, such as automated production lines in industrial environments and intelligent transportation. It can automatically select the optimal beam scheduling scheme according to dynamic channel status and network load, effectively reducing peak information age and improving network information timeliness.

[0005] To achieve the above objectives, the present invention provides the following solution:

[0006] A beam scheduling method based on peak information age optimization in millimeter-wave communication networks includes:

[0007] Obtain environmental and network parameters of the millimeter-wave network;

[0008] Based on the environmental and network parameters, an information age update model is constructed;

[0009] Based on the information age update model, construct the beam scheduling problem;

[0010] The beam scheduling problem is transformed into a Markov decision process, and a deep reinforcement learning algorithm is used to solve the Markov decision process to obtain the beam scheduling scheme with the minimum peak information age.

[0011] Optionally, the environmental parameters include the channel gain between each node and the base station, and the set of beamgroup associated nodes;

[0012] The network parameters include network bandwidth, received power, noise power, data packet size, and time slot duration.

[0013] Optionally, constructing an information age update model based on the environmental and network parameters includes:

[0014] Calculate the node communication rate based on the channel gain and the network bandwidth;

[0015] Calculate the beam group transmission rate based on the node communication rate and the set of nodes associated with the beam group;

[0016] The number of transmission time slots is obtained based on the beam group transmission rate, the data packet size, and the time slot duration;

[0017] An information age update model is constructed based on the beam group scheduling decision, the number of transmission time slots, and the number of remaining time slots.

[0018] Optionally, constructing an information age update model includes:

[0019]

[0020] Among them, a i (t+1) represents the information age of node i in time slot t+1, a i z(t) represents the information age of the node in time slot t, where z(t-τ) is the information age of the node. j ) represents time slot t-τ j The number of remaining transmission time slots at that time, x j (t-τ j ) indicates that beamgroup j is in time slot t-τ j The scheduling decision variable, τ j Y represents the number of transmission time slots in beamgroup j. j Let j be the set of nodes associated with beam group j.

[0021] Optionally, based on the information age update model and the beam scheduling constraints, the beam scheduling problem is constructed as follows:

[0022] Based on the information age update model, the peak information age optimization objective is constructed by minimizing the maximum information age among all nodes, wherein the peak information age optimization objective is:

[0023] MinA peak ;

[0024]

[0025] Based on the peak information age optimization objective and beam scheduling constraints, the beam scheduling problem is constructed, wherein the beam scheduling constraints are:

[0026]

[0027] Among them, A peak A represents the maximum information age among all nodes. i peak (t) represents the maximum information age of node i in time slot t, T is the total length of the observed time slots, and N is the total number of nodes in the millimeter-wave communication system. This indicates the total number of beam groups in a millimeter-wave network.

[0028] Optionally, transforming the beam scheduling problem into a Markov decision process includes:

[0029] Construct a Markov decision process that includes a state space, action space, transmission probability, reward function, and discount factor;

[0030] Wherein, in the state space S, s∈S, the state s(t) in time slot t is:

[0031] s(t)={(a i (t),z i (t),h i (t))} i∈N ;

[0032] The action space X, where x∈X, has the following selectable action x(t) in time slot t:

[0033]

[0034] The reward function c(t) in time slot t is:

[0035]

[0036] Among them, a i (t) is the information age of node i in time slot t, z i (t) represents the number of remaining transmission time slots, h i (t) is the channel gain between node i and the base station, N is the total number of nodes, β is the weighting coefficient, and τ j x is the number of transmission time slots in beamgroup j. j (t) represents the beam scheduling decision. This represents the total number of beamgroups in the millimeter-wave network, s represents the current state, and x represents the action.

[0037] Optionally, solving the Markov decision process using a deep reinforcement learning algorithm includes:

[0038] The actor-critic algorithm employs an actor network that outputs a probability value vector, a constraint processing module that generates a Boolean mask vector with the same dimension as the probability value vector, a mask module that calculates the probability distribution of legal actions using the probability value vector and the Boolean mask vector, and a critic network that outputs an estimate of the current state value function.

[0039] Based on the legal action probability distribution, the optimal strategy is output through the near-end policy optimization algorithm to obtain the beam scheduling scheme with the minimum peak information age.

[0040] Optionally, calculating the probability distribution of legal actions includes:

[0041]

[0042] Where, π θ (X,S) represents the legal action probability distribution, π′ θ (X,S) is a probability value vector. is a Boolean mask vector, and ° represents the Hadamard product.

[0043] Optionally, the loss function of the actor-critic algorithm is:

[0044] L(θ,ω)=L actor (θ)+L critic (ω);

[0045]

[0046]

[0047] Where L(θ,ω) is the total loss function of the actor-critic algorithm, L actor (θ) is the actor network loss function, L critic (ω) is the commentator network loss function. For the gradient loss of the pruning strategy, γ e L is the entropy regularization coefficient. entropy (θ) represents the entropy regularization term, where θ is the parameter of the actor network. For the trajectory length, To observe the state of trajectory t, Let c(l) be the state-value function, c(l) be the reward, and l be the traversal path from the observed trajectory. To the future moment, Discount factor Power of 1 This represents the state of the trajectory's endpoint.

[0048] This invention also provides a system for implementing a beam scheduling method based on peak information age optimization in millimeter-wave communication networks, comprising:

[0049] The parameter acquisition module is used to acquire environmental and network parameters of the millimeter-wave network.

[0050] The first construction module is used to construct an information age update model based on the environmental parameters and network parameters;

[0051] The second construction module is used to update the model based on the information age and construct the beam scheduling problem;

[0052] The optimization solution module is used to transform the beam scheduling problem into a Markov decision process, and to solve the Markov decision process using a deep reinforcement learning algorithm to obtain the beam scheduling scheme with the minimum peak information age.

[0053] The beneficial effects of this invention are as follows: by constructing a Markov decision process model and solving it using a deep reinforcement learning algorithm, this invention can automatically generate the optimal beam scheduling scheme based on environmental parameters; by minimizing the peak information age, it significantly improves the timeliness of network information; and by adapting to complex and ever-changing network environments through an intelligent learning mechanism, it provides reliable protection for application scenarios with high real-time requirements such as industrial control and intelligent transportation. Attached Figure Description

[0054] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0055] Figure 1 This is a system model diagram of millimeter-wave data transmission with an adaptive antenna according to an embodiment of the present invention;

[0056] Figure 2 This is a framework diagram of a beam scheduling method based on peak information age optimization in a millimeter-wave communication network according to an embodiment of the present invention.

[0057] Figure 3 This is a flowchart illustrating a beam scheduling method based on peak information age optimization in a millimeter-wave communication network according to an embodiment of the present invention. Detailed Implementation

[0058] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0059] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0060] like Figure 3 As shown, this embodiment provides a beam scheduling method based on peak information age optimization in millimeter-wave communication networks, including:

[0061] Obtain environmental and network parameters of the millimeter-wave network;

[0062] Based on environmental and network parameters, an information age update model is constructed;

[0063] Based on the information age update model, construct the beam scheduling problem;

[0064] The beam scheduling problem is transformed into a Markov decision process, and a deep reinforcement learning algorithm is used to solve the Markov decision process to obtain the beam scheduling scheme with the minimum peak information age.

[0065] Furthermore, environmental parameters include the channel gain between each node and the base station, as well as the set of beamgroup-associated nodes;

[0066] Network parameters include network bandwidth, received power, noise power, data packet size, and time slot duration.

[0067] Furthermore, based on environmental and network parameters, the information age update model is constructed as follows:

[0068] Calculate the node communication rate based on channel gain and network bandwidth;

[0069] Calculate the beam group transmission rate based on the node communication rate and the set of nodes associated with the beam group;

[0070] The number of transmission time slots is obtained based on the beam group transmission rate, data packet size, and time slot duration;

[0071] An information age update model is constructed based on beamgroup scheduling decisions, the number of transmission time slots, and the number of remaining time slots.

[0072] Furthermore, constructing the information age update model includes:

[0073]

[0074] Among them, a i (t+1) represents the information age of node i in time slot t+1, a i z(t) represents the information age of the node in time slot t, where z(t-τ) is the information age of the node. j ) represents time slot t-τ j The number of remaining transmission time slots at that time, x j (t-τ j ) indicates that beamgroup j is in time slot t-τ j The scheduling decision variable, τ j Y represents the number of transmission time slots in beamgroup j. j Let j be the set of nodes associated with beam group j.

[0075] Furthermore, based on the information age update model and beam scheduling constraints, the beam scheduling problem is constructed as follows:

[0076] Based on the information age update model, the peak information age optimization objective is constructed by minimizing the maximum information age among all nodes, where the peak information age optimization objective is:

[0077] MinA peak ;

[0078]

[0079]

[0080] Based on the peak information age optimization objective and beam scheduling constraints, a beam scheduling problem is constructed, where the beam scheduling constraints are:

[0081]

[0082] Among them, A peak A represents the maximum information age among all nodes. i peak (t) represents the maximum information age of node i in time slot t, T is the total length of the observed time slots, and N is the total number of nodes in the millimeter-wave communication system. This indicates the total number of beam groups in a millimeter-wave network.

[0083] Furthermore, transforming the beam scheduling problem into a Markov decision process includes:

[0084] Construct a Markov decision process that includes a state space, action space, transmission probability, reward function, and discount factor;

[0085] Wherein, in the state space S, s∈S, the state s(t) in time slot t is:

[0086] s(t)={(ai (t),z i (t),h i (t))} i∈N ;

[0087] The action space X, where x∈X, has the following selectable action x(t) in time slot t:

[0088]

[0089] The reward function c(t) in time slot t is:

[0090]

[0091] Among them, a i (t) is the information age of node i in time slot t, z i (t) represents the number of remaining transmission time slots, h i (t) is the channel gain between node i and the base station, N is the total number of nodes, β is the weighting coefficient, and τ j x is the number of transmission time slots in beamgroup j. j (t) represents the beam scheduling decision. This represents the total number of beamgroups in the millimeter-wave network, s represents the current state, and x represents the action.

[0092] Furthermore, the use of deep reinforcement learning algorithms to solve the Markov decision process includes:

[0093] The actor-critic algorithm is used to output a probability value vector from the actor network, generate a Boolean mask vector with the same dimension as the probability value vector from the constraint processing module, calculate the probability distribution of legal actions from the probability value vector and the Boolean mask vector from the mask module, and output an estimate of the current state value function from the critic network.

[0094] Based on the legal action probability distribution, the optimal strategy is output through the near-end policy optimization algorithm to obtain the beam scheduling scheme with the minimum peak information age.

[0095] Specifically, the model is solved using deep reinforcement learning algorithms based on A2C and Proximal Policy Optimization (PPO) to obtain the beam scheduling scheme with the minimum peak information age, including the following steps:

[0096] S1: Input a Markov decision process (MDP) quintuple (S,X,P,C,γ) constructed from network parameters and environmental parameters as the basic framework;

[0097] S2: Design a PPO algorithm based on an actor-critic framework to train the scheduling strategy. The actor network outputs action probabilities, and the critic network evaluates state values. To handle action constraints, an action masking mechanism is introduced to ensure that only legal actions (i.e., actions that satisfy x) are selected. j(actions where t)≤1);

[0098] S3: Based on the Proximal Policy Optimization (PPO) training mechanism, it collects interaction trajectories, calculates temporal difference advantages, iteratively optimizes the strategy by editing the objective function, updates the value network synchronously, uses the reward function to guide the peak information age penalty, and combines network and environmental parameter driving to output the minimum peak information age beam scheduling scheme adapted to the scenario.

[0099] Furthermore, calculating the probability distribution of legal actions includes:

[0100]

[0101] Where, π θ (X,S) represents the legal action probability distribution, π′ θ (X,S) is a probability value vector, and mask(S) is a Boolean mask vector. Let ω represent the Hadamard product, where ω is a parameter of the commentator network.

[0102] Furthermore, the loss function of the actor-critic algorithm is:

[0103] L(θ,ω)=L actor (θ)+L critic (ω);

[0104]

[0105] Where L(θ,(ω)) is the total loss function of the actor-critic algorithm, L actor (θ) is the actor network loss function, L critic (ω) is the commentator network loss function. γ is used to prune the policy gradient loss and limit the magnitude of policy updates. e L is the entropy regularization coefficient. entropy (θ) is the entropy regularization term, used to encourage exploration and prevent the strategy from converging to a local optimum too early. θ represents the parameters of the actor network, and ω represents the parameters of the critic network. For the trajectory length, For the observed trajectory, π θ For output strategy, For observation trajectory state, Let c(l) be the state-value function, and c(l) be the traversal function from the observed trajectory. Rewards in the future Discount factor Power of 1 Let x be the state at the endpoint of the trajectory, l be the decision variable, and l be the traversal from the currently observed trajectory. The future moments are used to collect the actual costs that will occur in the future.

[0106] The method of this embodiment will be further described below with reference to the accompanying drawings:

[0107] A beam scheduling method based on peak information age optimization in millimeter-wave communication networks includes:

[0108] S1: Obtain environmental and network parameters of the millimeter-wave network;

[0109] S2: Construct a millimeter-wave single-hop multicast network system model based on parameter definition, and simultaneously define beam scheduling constraints and information age (AoI) model based on this system model;

[0110] S3: Constructing a Markov Decision Process (MDP) quintuple based on environmental and network parameters.

[0111] S4: A deep reinforcement learning model is constructed based on the actor-critic (A2C) framework, and the model is solved using a deep reinforcement learning algorithm based on A2C and proximal policy optimization (PPO) to obtain the beam scheduling scheme with the minimum peak information age.

[0112] Furthermore, network parameters include: network bandwidth W, received power P. r Noise power P0, data packet size M, and duration t of a single time slot o .

[0113] Furthermore, environmental parameters include: the channel gain h between each node i and the base station. i (t), and the set of nodes Y associated with beam group j. j .

[0114] Consider a single-hop wireless millimeter-wave network, such as Figure 1 As shown, the network consists of a base station (BS) equipped with a millimeter-wave directional antenna and N receiving nodes. The receiving nodes are randomly distributed within a circular area of ​​finite radius r, denoted as {1,…,N}, with a total of N nodes. The BS is located at the center of this area and is responsible for transmitting data to all receiving nodes within the area. In this network, the BS employs a simulated beamforming system with a single radio frequency link and uses a single-beam transmission mode for data distribution. Due to the high directionality of millimeter-wave communication, the BS can only cover a subset of the target nodes in each transmission; therefore, the beam direction must be adjusted sequentially in the time domain, and multicast data across the entire network must be achieved through multiple rounds of transmission.

[0115] Furthermore, the step of constructing beam scheduling calculations and information age updates based on network parameters and channel states, and building a model, specifically includes:

[0116] Based on the current channel gain h i (t) and network bandwidth W, computing node communication rate r i ;

[0117] Based on node communication rate r i Beamgroup associated node set Y j Calculate the beam group transmission rate r j ;

[0118] Based on beam group transmission rate r j Data packet size M and time slot duration t o Construct the number of transmission time slots τ j ;

[0119] Beamgroup scheduling decision x j (t), number of transmission time slots τ j Construct an information age update model using the remaining time slot number z(t);

[0120] Based on the age a of each node information i (t) Construct a peak information age optimization target.

[0121] The problem definition is constructed based on all current dynamic channel environments.

[0122] Furthermore, the step of constructing the Markov Decision Process (MDP) quintuple based on environmental and network parameters is as follows:

[0123] Combining environmental and network parameters, a quintuple MDP is constructed, comprising a state space (integrating node information, age, etc.), an action space (beam scheduling decision), transmission probability, a reward function, and a discount factor. Here, S is a finite state space, and X is a finite action space. P(s'|s,α) represents the state transition probability, i.e., the probability of transitioning to state s'∈S after performing action x∈X in state s∈S. C(s,x) is the immediate reward function, representing the immediate reward obtained by performing action x in state s. γ∈[0,1]

[0124] This is a discount factor that reflects the degree to which the impact of current rewards on future decisions diminishes.

[0125] Furthermore, a Markov Decision Process (MDP) quintuple is constructed based on environmental and network parameters, and a deep reinforcement learning model is built based on the actor-critic (A2C) framework. The model is then solved using a deep reinforcement learning algorithm based on A2C and Proximal Policy Optimization (PPO) to obtain the beam scheduling scheme with the minimum peak information age. The steps include:

[0126] S1: Input a Markov Decision Process (MDP) quintuple constructed from network parameters and environmental parameters as the basic framework;

[0127] S2: Design a PPO algorithm based on an actor-critic framework to train the scheduling strategy. The actor network outputs action probabilities, and the critic network evaluates state values. To handle action constraints, an action masking mechanism is introduced to ensure that only legal actions (i.e., actions that satisfy x) are selected. j (actions where t)≤1);

[0128] S3: Based on the Proximal Policy Optimization (PPO) training mechanism, it collects interaction trajectories, calculates temporal difference advantages, iteratively optimizes the strategy by editing the objective function, updates the value network synchronously, uses the reward function to guide the peak information age penalty, and combines network and environmental parameter driving to output the minimum peak information age beam scheduling scheme adapted to the scenario.

[0129] Furthermore, in millimeter-wave network communication scenarios, the data transmission requirements of individual nodes (s) i Data transmission rate r i Constructed based on Shannon's capacity formula:

[0130]

[0131] Where, r i This represents the communication rate of node i, measured in bits per second. This rate is determined by Shannon's formula and reflects the theoretical maximum transmission rate that the node can achieve under current channel conditions and network configuration.

[0132] Furthermore, regarding the data transmission characteristics of beamgroup coverage nodes based on millimeter-wave communication, the transmission rate of beamgroup j is determined by the worst-case channel conditions of the nodes within the group, expressed as:

[0133] r j =min{r i}, i∈Υ j ;

[0134] Where, r j Y represents the transmission rate of beam group j, measured in bits per second. j Let j be the set of nodes associated with beam group j. This formula ensures that all nodes within the beam group can successfully receive transmitted data, so the transmission rate is determined by the node with the worst channel conditions within the group.

[0135] Furthermore, regarding the time slot requirements for beamgroup data transmission in millimeter-wave communication, the number of time slots required for beamgroup j to transmit data is calculated by dividing the data packet size by the number of bits transmitted per time slot and rounding up.

[0136]

[0137] Where, τ jThe number of transmission time slots for beam group j is a positive integer. The minimum number of time slots required to complete the transmission is determined by dividing the data packet size by the number of bits that can be transmitted per time slot and rounding up.

[0138] Furthermore, considering the dynamic changes in node data freshness required in millimeter-wave communication, the update rule for the Information Age (AoI) of a single node i is as follows:

[0139]

[0140] Among them, a i (t) represents the information age of node i in time slot t, and its unit is the number of time slots. Two update scenarios are specified: when a node successfully receives new data, the information age is reset to the number of transmission time slots; otherwise, the information age is incremented by one time slot.

[0141] Furthermore, the update rule for the remaining transmission slot number z(t) is expressed as follows:

[0142]

[0143] Where z(t) represents the number of remaining transmission time slots at time slot t, x j (t) represents the scheduling decision variable for beamgroup j in time slot t, τ j This represents the number of transmission time slots for beamgroup j. Two update scenarios are specified: when a new beamgroup starts transmitting, the remaining time slots are initialized to the number of transmission time slots minus 1; otherwise, the remaining time slots are decremented by 1 until they reach zero.

[0144] Furthermore, the optimization objective for peak information age is to minimize the maximum information age among all nodes, and its expression is as follows:

[0145]

[0146] Among them, a i (t) represents the instantaneous information age of the i-th node at time slot t; T is the total length of the observed time slots; N is the total number of nodes in the millimeter-wave communication system. The optimization objective is to ensure that the oldest information age in the network is as small as possible, thereby improving the information timeliness of the entire system.

[0147] Furthermore, the current peak information age optimization method based on millimeter-wave communication defines the problem as follows:

[0148] MinA peak ;

[0149]

[0150] in x represents the total number of beamgroups in a millimeter-wave network. j(t) is the beam scheduling decision variable, where t represents the discrete time slot and takes values ​​in the range {1,…,T}. It is stipulated that within any single time slot t, a base station can schedule at most one beam group to avoid conflicts caused by scheduling multiple beam groups simultaneously.

[0151] Furthermore, the state space S of the step of constructing a Markov Decision Process (MDP) quintuple based on environmental and network parameters is used to characterize the key states of the network and nodes at the scheduling time, where s∈S. The expression for the state s(t) in time slot t is:

[0152] s(t)={(α i (t), z i (t), hi(t))} i∈N ;

[0153] Among them, a i (t) is the information age of node i in time slot t, z i (t) represents the number of remaining transmission time slots, h i (t) is the channel gain between node i and the base station, and N is the total number of nodes.

[0154] Furthermore, based on environmental and network parameters, the action space X of the Markov Decision Process (MDP) quintuple step is constructed to describe the beam group scheduling decision for time slot t. Given x∈X, the expression for the optional action x(t) in time slot t is:

[0155]

[0156] in, It is the total number of beam groups, x j (t) = 1 indicates that time slot t schedules beamgroup j, x j (t) = 0 indicates no scheduling.

[0157] The reward function (instant reward c(t)) in the step of constructing the Markov Decision Process (MDP) quintuple based on environmental and network parameters is used to quantify the quality of scheduling decisions and guide the optimization of the peak information age objective. Its expression is:

[0158]

[0159] Where β is the weighting coefficient, τ j x is the number of transmission time slots in beamgroup j. j (t) represents the beam scheduling decision, a i (t) is the information age of node i.

[0160] This embodiment models the beam scheduling problem as an MDP, and can use deep reinforcement learning algorithms (such as DQN, PPO, DDPG) to optimize beam selection and scheduling strategies, so as to dynamically adapt to channel changes and minimize the peak information age of the system.

[0161] In millimeter-wave wireless network environments, the action space expands exponentially with the increase in the number of nodes. Furthermore, the network environment is dynamically changing, making it difficult to find the optimal scheduling policy using traditional optimization methods when channel conditions change. To address this issue, a model-free scheduling method based on deep reinforcement learning is proposed. The core of this method is a policy gradient approach, employing the Proximal Policy Optimization (PPO) algorithm within the Actor-Critic (A2C) framework.

[0162] Generally, policy-based methods aim to dynamically learn an optimal policy π. * At the same time, maximize the expected value of collecting rewards. To reduce sampling complexity, PPO achieves optimization by maximizing the following alternative objective function:

[0163]

[0164] Where E{} represents the expected value. For the strategy ratio, π θ This represents the strategy parameterized by parameter θ. old This indicates the strategy parameters before the most recent update. Indicating the dominance function The estimator. clip(ρ) θ (1-∈, 1+∈) is the clipping function: when When ρ θ Restricted to 1+∈, when The time limit is 1-∈.

[0165] As attached Figure 2 As shown, a multicast scheduling method based on deep reinforcement learning is proposed, comprising an A2C (Advantage Actor-Critic) framework consisting of an actor network and a critic network. Both networks adopt the same structure, including a state input layer, multiple fully connected layers as hidden layers, and a ReLU function as the activation function. The actor network outputs a probability vector representing the policy, denoted as π'. θ(X,S) represents the process by which the actor network formulates policies based on different system states; its dimension is equal to the size of the action space. The critic network outputs an estimate of the current state value function, used to evaluate the expected reward obtained from performing an action from the current state. The PPO reinforcement learning algorithm is used to learn through interaction with the environment, improving the accuracy of the policy so that the final policy can be used for accurate beam scheduling in millimeter-wave networks.

[0166] To better adapt to the specific multicast scheduling problem scenario of this invention, this embodiment improves upon the A2C framework by introducing an action masking module to handle complex constraints. This module effectively masks illegal actions that violate constraints, ensuring that only feasible actions are considered during training, thereby improving the model's training efficiency and decision quality. Under strict constraints, this design allows the controller to still learn efficiently and make reasonable decisions. In this model, the output of the ActorNetwork is a probability vector π'. θ (X,S) represents the probability distribution of taking different actions X in the current state S. Simultaneously, the constraint processing module generates a probability vector π'. θ A Boolean mask vector mask(S) of the same dimensions (X,S). Each element of mask(S) corresponds to a possible action; a value of 1 indicates that the action is feasible, and a value of 0 indicates that the action violates the constraints. The core function of the constraint processing module is to dynamically generate mask(S) based on the current state of the system and predefined constraints. In the current scenario, the main constraints include: the millimeter-wave base station can only schedule one node group per time slot; and at any given transmission time, only one node group is allowed to be active.

[0167] After illegal actions are masked by the masking module, a very small value ω needs to be added to the masked action location to ensure smooth gradient propagation. Furthermore, to ensure the neural network can correctly identify and select the optimal action, the remaining probability distribution needs to be renormalized. Therefore, the masking module outputs the probability distribution π of legal actions. θ (X,S) ensures that the model always optimizes and makes decisions within the feasible action space.

[0168]

[0169] Where, π' θ (X,S) and mask(S) are two matrices of the same dimension. This represents their Hadamard product (element-by-element product).

[0170] The commentator network is a neural network consisting of parameters ω, which can be used to estimate the state value function. The estimated value is denoted as This estimate will be used in the Generalized Dominance Estimation (GAE) method to construct the dominance function. Specifically, GAE utilizes the old strategy The dominance function A is estimated by sampling within a finite number of times (i.e., one batch). π (·)

[0171] Assuming a batch of R samples is collected before each policy update, for simplicity, the initial slot index will be omitted in the following steps. Mark the trajectory length. GAE uses... A linear combination of steps is used to estimate the action value function Q. π (s(t),x(t)), its estimated value The calculation is as follows:

[0172]

[0173] Accordingly, the estimator of the advantage function can be calculated as follows:

[0174]

[0175] in This refers to the timing difference (TD) error.

[0176] For actor networks, an entropy term is added to their loss function to prevent them from getting trapped in local optima too early. This is denoted as L. entropy (θ), to encourage exploration:

[0177]

[0178] Define the adjustable coefficient γ e Then the final loss function L of the policy network actor (θ) takes the following form:

[0179]

[0180] in, It is L clip (θ) By observing the trajectory It approximates the expected value.

[0181] For a critic network, its output state-value function The estimation error (i.e., the loss function to be minimized) can also be obtained through the trajectory length. An approximation is made, in the following form:

[0182]

[0183] In this embodiment, the actor network and the critic network are treated as a unified neural network, whose parameters are represented by θ and ω, and which can output policy π. θ and state value function Therefore, the final loss function is L actor (θ) by L critic (ω) together form the following:

[0184] L(θ,ω)=L actor (θ)+L critic (ω);

[0185] The scheduling strategy output by this algorithm can serve as the basis for base stations to make scheduling decisions in each time slot, guiding the selection of scheduling multicast beams and real-time updates of system status.

[0186] This embodiment also provides a system for implementing a beam scheduling method based on peak information age optimization in a millimeter-wave communication network, including:

[0187] The parameter acquisition module is used to acquire environmental and network parameters of the millimeter-wave network.

[0188] The first building module is used to build an information age update model based on environmental and network parameters;

[0189] The second building module is used to update the model based on the information age and construct the beam scheduling problem;

[0190] The optimization solution module is used to transform the beam scheduling problem into a Markov decision process. A deep reinforcement learning algorithm is used to solve the Markov decision process to obtain the beam scheduling scheme with the minimum peak information age.

[0191] Specifically, an actor-commentator (A2C) framework with action masking is designed, integrating relevant networks and action masking modules, configuring experience replay cache interaction data, and training it using a near-end policy optimization (PPO) mechanism based on A2C. The trained A2C framework solves the MDP model, and the policy is output to the base station controller through the real-time decision module to obtain the beam scheduling scheme with the minimum peak information age.

[0192] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A method for optimizing beam scheduling based on peak information age in a millimeter wave communication network, characterized in that, The method comprises the following steps: obtaining environment parameters and network parameters of a millimeter wave network; constructing an information age update model according to the environment parameters and the network parameters; constructing a beam scheduling problem according to the information age update model and beam scheduling constraints, comprising: constructing a peak information age optimization target by minimizing the maximum information age in all nodes according to the information age update model, wherein the peak information age optimization target is ; ; ; constructing the beam scheduling problem according to the peak information age optimization target and the beam scheduling constraints, wherein the beam scheduling constraints are ; ; wherein, is the maximum information age in all nodes, is the maximum information age in a node, is the maximum information age in a slot, is the maximum information age in a slot, is the total slot length observed, is the total number of nodes in the mmWave communication system, denotes the total number of beam groups in the mmWave network; transforming the beam scheduling problem into a Markov decision process, and solving the Markov decision process by using a deep reinforcement learning algorithm to obtain a beam scheduling scheme with the minimum peak information age.

2. The method of claim 1, wherein, The environment parameters comprise channel gains between each node and a base station, and a node set associated with a beam group; The network parameters comprise network bandwidth, received power, noise power, packet size, and time slot length.

3. The method of claim 2, wherein, The information age update model is constructed according to the environment parameters and the network parameters, comprising: calculating a node communication rate according to the channel gains and the network bandwidth; calculating a beam group transmission rate according to the node communication rate and the node set associated with the beam group; obtaining a transmission time slot number according to the beam group transmission rate, the packet size, and the time slot length; constructing an information age update model according to a beam group scheduling decision, the transmission time slot number, and a remaining time slot number.

4. The method of claim 3, wherein, The information age update model is constructed, comprising: ; in, Represents a node In the time slot +1 information age, Indicates the node in the time slot Information age, Indicates time slot The number of remaining transmission time slots at that time. Indicates beam group In the time slot Scheduling decision variables, Indicates beam group The number of transmission time slots, For beam groups The set of associated nodes.

5. The method of claim 1, wherein, transforming the beam scheduling problem into a Markov decision process, comprising: constructing a Markov decision process comprising a state space, an action space, a transmission probability, a reward function, and a discount factor; Wherein the state space , there are , in time slots of states : ; The action space , there are , in the time slot of the optional action : ; The reward function at time slot t is: ; in, It is a node In the time slot Information age, It is the number of remaining transmission time slots. It is a node With the channel gain of the base station, This represents the total number of nodes in a millimeter-wave communication system. These are weighting coefficients. It is a beam group The number of transmission time slots, It is a beam scheduling decision. This represents the total number of beamgroups in the millimeter-wave network, s represents the current state, and x represents the action.

6. The method of claim 5, wherein, solving the Markov decision process by using a deep reinforcement learning algorithm, comprising: outputting a probability value vector by an actor network in an actor-critic algorithm, generating a Boolean mask vector with the same dimension as the probability value vector by a constraint processing module, calculating a legal action probability distribution by a mask module through the probability value vector and the Boolean mask vector, and outputting an estimated value of a current state value function by a critic network; outputting an optimal strategy according to the legal action probability distribution by a proximal policy optimization algorithm to obtain the beam scheduling scheme with the minimum peak information age.

7. The method of claim 6, wherein, The legal action probability distribution is calculated, comprising: ; wherein, is a legal action probability distribution, is a probability value vector, is a Boolean mask vector, denotes a Hadamard product, are parameters of the critic network.

8. The method of claim 7, wherein, The loss function of the actor-critic algorithm is ; ; ; ; wherein, Lactor-critic is the total loss function for the actor-critic algorithm, Lactor-net is the actor network loss function, Lcritic-net is the critic network loss function, Ltrunc-pg is the truncated policy gradient loss, a is the entropy regularization coefficient, H is the entropy regularization term, p is the parameter of the actor network, T is the trajectory length, s is the state of the observed trajectory t ~, V is the state value function, r is the reward, l is the iteration from the observed trajectory to the future time, g is the discounted factor to the power of l, s is the state of the trajectory end.

9. The system implemented by the method for beam scheduling based on peak information age optimization in millimeter wave communication network according to any one of claims 1-8, characterized in that, The method comprises the following steps: a parameter acquisition module is configured to obtain environment parameters and network parameters of a millimeter wave network; a first construction module is configured to construct an information age update model according to the environment parameters and the network parameters; a second construction module is configured to construct a beam scheduling problem according to the information age update model; an optimization solving module is configured to transform the beam scheduling problem into a Markov decision process, and solve the Markov decision process by using a deep reinforcement learning algorithm to obtain a beam scheduling scheme with the minimum peak information age.

Citation Information

Patent Citations

  • High-throughput hopping beam scheduling method for differentiated services

    CN114629547A

  • RIS-assisted MISO system optimization method based on deep reinforcement learning

    CN118900143A