Intelligent dynamic resource allocation method for industrial 5G networks based on branched D3QN
By adopting the D3QN method with improved branch structure in industrial 5G networks, dynamic allocation of physical resource blocks and power is solved, and the problem of resource allocation complexity and reliability constraints in industrial 5G networks is achieved, and efficient resource utilization and low power consumption are achieved.
Patent Information
- Application Number
- CN202310076271.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-01-18
AI Technical Summary
In industrial 5G networks, how to realize intelligent joint allocation of physical resource blocks and power, meet the differentiated service quality requirements of each node, and balance resource utilization and power consumption.
Using an intelligent dynamic resource allocation method based on branch D3QN, in a single base station multi-slicing and multi-node system model, the physical resource blocks and power are dynamically allocated through the improved branch structure of D3QN, satisfying the reliability constraints between slices, and minimizing the weighted sum of physical resource block usage and power.
By reducing the size of the action space, the learning speed of the allocation strategy is accelerated, the service quality of each node is guaranteed, and resource utilization and power consumption are balanced.
Smart Images

Figure CN116095757B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of mobile communications and relates to an industrial 5G network intelligent dynamic resource allocation method based on branch D3QN. Background Art
[0002] With the widespread adoption of 5G technology in the Industrial Internet of Things (IIoT), ensuring the real-time and reliable transmission of industrial data within 5G networks has become a critical challenge for 5G technology development. Industrial 5G networks contain a large number of nodes with varying reliability and quality of service (QoS) requirements. This requires a rational resource allocation approach to ensure that each node meets these differentiated QoS requirements. Furthermore, resource utilization efficiency and total power consumption are both key performance indicators. For optimal system performance, a balance must be struck between improving resource utilization and reducing power consumption.
[0003] In industrial 5G networks, a single-base-station, multi-slice, multi-node model is a commonly used network slicing radio access network downlink system architecture. The base station allocates network resources for the communication links between the base station and the nodes to maintain quality of service. Compared to the resource allocation problem in traditional IoT systems, resource allocation in industrial 5G network scenarios is more complex. First, resource allocation must meet reliability constraints, making resource allocation strategies more complex. Second, as the number of nodes in industrial 5G networks increases, the number of possible physical resource block allocation combinations increases exponentially, resulting in a very complex action space. Finally, the joint allocation of physical resource blocks and power allocation further increases the size of the action space.
[0004] Therefore, there is an urgent need for an intelligent dynamic resource allocation method for network slicing suitable for industrial 5G networks to meet reliability constraints and achieve a balance between resource utilization and power consumption. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide an intelligent dynamic resource allocation method for industrial 5G networks based on branched D3QN, which performs intelligent joint allocation of physical resource blocks and power in industrial 5G network scenarios. Specifically, in the downlink scenario of the network slicing radio access network, the improved D3QN with a branched structure is used to solve the weighted sum minimization problem of physical resource block usage and power. The branched structure effectively decouples the physical resource block and power allocation combination, reducing the action space for problem solving. Through the proposed resource allocation method, the present invention reduces the action space expressed by the deep reinforcement learning method and accelerates the learning speed of the allocation strategy, ensuring the service quality requirements of each node.
[0006] In order to achieve the above object, the present invention provides the following technical solutions:
[0007] An intelligent dynamic resource allocation method for industrial 5G networks based on branched D3QN is proposed. In a single-base station multi-slice multi-node system model, a D3QN allocation method improved by a branched structure is used to dynamically allocate physical resource blocks and power to each node in each time slot. While ensuring mutual isolation and reliability between slices, the weighted sum of physical resource block usage and power is minimized.
[0008] The method specifically comprises the following steps:
[0009] S1: Obtain parameter information in the industrial 5G network system, construct a D3QN with improved branch structure, and initialize network parameters;
[0010] S2: Establish the state space, action space, and reward function based on the system model and the optimization goal of minimizing the weighted sum of physical resource block usage and power;
[0011] S3: Updates the latency, backlog, reliability information, and branch structure of each node to explore the action space and output the optimal action;
[0012] S4: Train the network and calculate the loss function, then use the gradient descent algorithm to update the network parameters until the network converges to obtain the trained network.
[0013] Furthermore, in step S1, a D3QN with improved branch structure is constructed, which specifically includes the following steps:
[0014] S101: The improved branch structure of D3QN is a network structure based on deep reinforcement learning, which includes a shared decision module and multiple network branches. It is assumed that in a single base station cell downlink scenario, the industrial 5G network contains N nodes. The shared decision module is shared by all branches, allowing for shared learning experience. Based on the number of nodes, the number of branches can be determined to be N+2, including a shared value output branch, a physical resource block allocation branch, and N power allocation branches. The shared value output module can output a shared value V(s) and use it in the advantage function of the dueling system. Because the allocation of physical resource blocks to each node is constrained by the isolation between slices, that is, a physical resource block can only be allocated to one node in the same time slot, the allocation of physical resource blocks must be completed in a single branch. Since the power allocation of each node is independent, the power allocation of N nodes can be divided into N branches.
[0015] S102: The duel structure is adopted in the improved branch structure D3QN because the duel structure can more quickly identify redundant actions, distinguish the impact of state on Q value and the impact of action on Q value, and thus accelerate learning. The advantage function of each branch in the duel structure is:
[0016] Qd (s,a d )=V(s)+A(s,a d )
[0017] Among them, Q d (s,a d ) is the state-action value of branch d, V(s) is the shared value, A(s,a d ) is in state s and action a d When the action advantage value of branch d;
[0018] S103: For the calculation of the target value in the D3QN with improved branch structure, the calculation is performed for each individual branch. The target value y of branch d d The calculation formula is:
[0019]
[0020] Where r represents the reward, γ represents the discount factor, and the target network corresponding to the branch d is defined as s ′ and a ′ d They represent the state of the next time slot and the action in branch d respectively.
[0021] Furthermore, step S2 specifically includes the following steps:
[0022] S201: Construct an optimization objective based on minimizing the weighted sum of physical resource block usage and power, expressed as:
[0023]
[0024] Where t represents the time slot, K represents the number of physical resource blocks, N represents the number of nodes, k and n represent the indexes of physical resource blocks and nodes respectively; ρ k,n (τ) is an indicator function. If physical resource block k is allocated to node n in time slot τ, then ρ k,n (τ)=1, otherwise, ρ k,n (τ)=0;ω n The weight representing the power of node n, p n represents the power of node n;
[0025] The constraints are as follows:
[0026] The number of allocated physical resource blocks is less than or equal to the number of physical resource blocks K owned:
[0027]
[0028] Ensure that slices are isolated from each other:
[0029]
[0030] Assume that there are M slices in total, m represents the index of the slice, and the reliability constraint of the slice is:
[0031]
[0032] Among them, D m represents the delay of slice m, represents the maximum tolerable delay of slice m, χ m Represents the reliability requirement; the power selection range is less than or equal to the maximum power p max :
[0033] 0≤p n (t)≤p max
[0034] S202: Latency refers to the end-to-end latency in a 5G-enabled Industrial IoT network slice, including transmission delay, propagation delay, queuing delay, and processing delay. Assuming the transmission rate is fast enough and there are sufficient computing resources to process the data to simplify the problem, transmission delay and processing delay can be ignored. Furthermore, propagation delay depends on the speed of light, so the propagation time between the node and the base station is very small and can be ignored. Therefore, in the case of limited bandwidth, the total delay D can be expressed as:
[0035] D=d queue
[0036] Among them, d queue Indicates queue delay;
[0037] S203: In order to store incoming traffic data to be sent and calculate the queuing delay, a queue is maintained for each node; the backlog U of the queue of node n at time slot t+1 is n (t+1) is expressed as:
[0038] U n (t+1)=max{U n (t)+α n (t)-r n (t),0}
[0039] Among them, α n (t), r n (t) represents the amount of data received and sent, respectively.
[0040] S204: Establish the state space s of the system t for:
[0041] s t ={q t ,dt ,α t ,h t}
[0042] Among them, q t d t , α t 、h t They are the set of backlogs, delays, arrival flows, and channel gains of each node;
[0043] S205: The action space of the physical resource block is all possible allocation methods based on the slice division. Taking two slices as an example, assuming there are a total of K physical resource blocks, the action space a of the physical resource block can be divided into: {{0,0}, {0,1}, …, {K,0}}. In order to better coordinate scheduling with the physical resource block, the power is discretized. The specific power action space can be expressed as:
[0044]
[0045] Among them, p max represents the maximum power, and |p| is the number of power discretizations.
[0046] In summary, the action space can be defined as
[0047]
[0048] S206: The allocation of physical resource blocks in S205 is based on the allocation of slices, so it is necessary to further allocate the physical resource blocks allocated to the slices to the nodes in the slices. The specific method is to determine the priority of allocating physical resource blocks to the nodes in the slice based on the latency and backlog of the current node;
[0049]
[0050] in, Indicates the priority of the current node. denote the delay and backlog of node n respectively. The backlog is defined as the amount of data that has not timed out and is waiting to be sent. ζ denotes the weight of the importance of delay and backlog.
[0051] Then, physical resource blocks are allocated according to the priority of the nodes. The amount of physical resource blocks that node n can obtain in slice m is:
[0052]
[0053] in, They represent the amount of physical resource blocks allocated to node n and slice m respectively;
[0054] S207: In order to meet the optimization goal of minimizing the weighted sum of the used physical resource blocks and power, a reward function is constructed as follows:
[0055]
[0056] Among them, K allo is the physical resource block used, ω n is the weight to measure the importance of physical resource blocks and power; δ n is the indicator function, which is 0 if node n satisfies reliability, otherwise it is 1; is the penalty weight, L n,t is the packet loss of node n at time slot t.
[0057] Furthermore, step S3 specifically includes the following steps:
[0058] S301: In the random action selection phase, to speed up training, further constraints are imposed on the action selection. For the physical resource block action space, if there are nodes in the slice that do not meet the quality of service constraint, the allocation actions with a value of 0 are removed, thereby obtaining a new action space a′. For the power action space, if a node does not meet the quality of service constraint, the allocation actions with a value of 0 are removed, thereby obtaining a new action space p′.
[0059] S302: In the process of obtaining the optimal action, a total of N+1 actions are obtained by selecting the action corresponding to the maximum action value function in each branch except the branch that outputs the shared value; assuming that the optimal behavior of each branch is a * , then a * It can be obtained by the following formula;
[0060]
[0061] The set of optimal actions is:
[0062]
[0063] in, represents the optimal action of branch i.
[0064] Further, step S4 specifically includes the following steps:
[0065] S401: The loss function Loss is obtained based on the estimated value and the target value returned by the target value network:
[0066]
[0067] Where Z represents the number of branches;
[0068] S402: Based on the loss function of step S401, calculate the gradient loss function for:
[0069]
[0070] in, represents the gradient vector of the estimation network, θ d represents the network parameters of branch d;
[0071] S403: Every several steps, the network parameters of the estimated network are updated to the target network until the network converges.
[0072] The beneficial effects of the present invention are:
[0073] 1) The intelligent allocation method for network slicing resources in industrial 5G network scenarios provided by the present invention autonomously learns allocation strategies through training, and then dynamically allocates physical resource blocks and power to each node based on time slots, which can not only meet service quality requirements but also balance network resource utilization and power consumption.
[0074] 2) Based on the deep reinforcement learning framework and the improved branch structure of D3QN, the present invention decouples the action space through the branch structure, divides the allocation of physical resource blocks and the power of each node into different branches, effectively reducing the size of the action space, making the intelligent allocation method applicable to multi-slice and multi-node scenarios.
[0075] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings, in which:
[0077] Figure 1 This is a training diagram of the improved D3QN based on the deep reinforcement learning framework branch structure of the present invention;
[0078] Figure 2 This is a flow chart of the method for joint allocation of physical resource blocks and power according to the present invention. DETAILED DESCRIPTION
[0079] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.
[0080] See also Figures 1 and 2 The present invention improves a network slicing intelligent dynamic resource allocation method suitable for industrial 5G networks. For a single-base station multi-node network slicing wireless access network scenario, physical resource blocks and power are jointly allocated to ensure that nodes meet reliability constraints, while minimizing the number of physical resource blocks used and the weighted sum of power. The D3QN method with an improved branch structure is used to divide the large action space caused by the coupling of multi-node resource allocation into multiple independent small action spaces, reducing the dimension of the action space to speed up the training speed, and obtaining a network slicing intelligent dynamic resource allocation method. The present invention reduces the complexity of problem solving through the D3QN method with an improved branch structure, ensures that reliability constraints are met, and achieves a balance between resource utilization and power consumption.
[0081] Figure 1 This is a training diagram of D3QN with improved branch structure based on deep reinforcement learning framework. Figure 1 As shown in the figure, during the interaction between the estimation network and the environment, the system state and action are used as inputs to obtain the state-action Q-value, reward, and next system state for each branch. The experience set is then stored in an experience replay pool. A batch of experience sets is randomly selected from the experience replay pool to form a training segment. The estimated value and target value for each training segment are then calculated. The loss function for the current state is calculated from the estimated and target values, and the gradient loss function is further derived. The estimation network updates its parameters using gradient descent. After a certain number of steps, the parameters of the estimation network are updated to the target network.
[0082] Figure 2 The flowchart of the hybrid update industrial wireless sensor network scheduling method based on information age of the present invention is as follows: Figure 2 As shown, the specific steps include:
[0083] V1-V4: Starting with the joint allocation method of physical resource blocks and power, obtaining parameter information in the industrial 5G network, constructing and initializing the D3QN with improved branch structure, and initializing the network state set, action space and reward function.
[0084] V5-V8: Update the latency, backlog, and reliability of each node. Based on the current system state, obtain actions, obtain rewards and the next state through environmental feedback, and input the experience collection into the experience replay pool.
[0085] V9~V14: Randomly select experience sets from the experience replay pool to form training segments, calculate the estimated value and target value of each branch, calculate the loss function based on the estimated value and target value, calculate the gradient and update the network weights, and update the estimated network parameters to the target network every few steps.
[0086] V15~V17: Determine whether the system is stable. If not, execute V5. If stable, perform joint allocation of physical resource blocks and power according to the trained network.
[0087] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.
Claims
1. An industrial 5G network intelligent dynamic resource allocation method based on branch D3QN, characterized in that: In the single-base station multi-slice multi-node system model, the D3QN allocation method with improved branch structure dynamically allocates physical resource blocks and power to each node in each time slot, minimizing the weighted sum of physical resource block usage and power while ensuring mutual isolation and reliability between slices. The method specifically comprises the following steps: S1: Obtain parameter information in the industrial 5G network system, construct a D3QN with improved branch structure, and initialize network parameters; wherein, constructing the D3QN with improved branch structure specifically includes the following steps: S101: The D3QN with improved branch structure includes a shared decision module and multiple network branches; it is set in a single base station cell downlink scenario, where the industrial 5G network contains N nodes; the shared decision module is shared by all branches, and the learning experience is shared; the number of branches is determined to be N+2 according to the number of nodes, including a shared value output branch, a physical resource block allocation branch, and N power allocation branches; the shared value output module outputs the shared value V(s), which is used in the advantage function in the duel system; the allocation of physical resource blocks is completed in one branch; the power allocation of each node is independent of each other, and the power allocation of N nodes is divided into N branches; S102: A duel structure is adopted in the D3QN with improved branch structure; the advantage function of each branch in the duel structure is: Q d (s,a d )=V(s)+A(s,a d ) Among them, Q d (s,a d ) is the state-action value of branch d, V(s) is the shared value, A(s,a d ) is in state s and action a d When the action advantage value of branch d is S103: For the calculation of the target value in the D3QN with improved branch structure, the calculation is performed for each separate branch. The target value y of branch d d The calculation formula is: Where r represents the reward, γ represents the discount factor, and the target network of the branch corresponding to branch d is defined as s′ and a′ d They represent the state of the next time slot and the action in branch d respectively; S2: Establish the state space, action space and reward function according to the system model and the optimization goal based on minimizing the weighted sum of physical resource block usage and power, which specifically includes the following steps: S201: construct an optimization objective based on minimizing the weighted sum of physical resource block usage and power, expressed as: Where t represents the time slot, K represents the number of physical resource blocks, N represents the number of nodes, k and n represent the indexes of physical resource blocks and nodes respectively; ρ k,n (t) is an indicator function. If physical resource block k is allocated to node n in time slot t, then ρ k,n (t)=1, otherwise, ρ k,n (t) = 0; ω n The weight representing the power of node n, p n represents the power of node n; The constraints are as follows: The number of allocated physical resource blocks is less than or equal to the number of physical resource blocks K owned: Ensure that slices are isolated from each other: Assume that there are M slices in total, m represents the index of the slice, and the reliability constraint of the slice is: Among them, D m represents the delay of slice m, represents the maximum tolerable delay of slice m, χ m Reliability requirements; The power selection range is less than or equal to the maximum power p max : 0≤p n (t)≤p max S202: In the case of limited bandwidth, the total delay D is expressed as: D=d queue Among them, d queue Indicates queue delay; S203: In order to store the incoming traffic data to be sent and count the queuing delay, a queue is maintained for each node; the backlog U of the queue of node n at time slot t+1 n (t+1) is expressed as: U n (t+1)=max{U n (t)+α n (t)-r n (t),0} Among them, α n (t), r n (t) represents the amount of data received and sent, respectively; S204: Establish the state space s of the system t for: s t ={q t ,d t ,α t ,h t } Among them, q t d t , α t 、h t They are the set of backlogs, delays, arrival flows, and channel gains of each node; S205: The action space of the physical resource block is all possible allocation methods based on slice division; S206: Determine the priority of allocating physical resource blocks to nodes in the slice according to the latency and backlog of the current node; in, Indicates the priority of the current node. They represent the delay and backlog of node n respectively. The backlog is defined as the amount of data that has not timed out and is waiting to be sent. ζ represents the weight of the importance of delay and backlog. Then, physical resource blocks are allocated according to the priority of the nodes. The amount of physical resource blocks that node n can obtain in slice m is: in, They represent the amount of physical resource blocks allocated to node n and slice m respectively; S207: Construct the reward function as: Among them, K allo is the physical resource block used, ω n is the weight to measure the importance of physical resource blocks and power; n is the indicator function, which is 0 if node n satisfies reliability, otherwise it is 1; is the penalty weight, L n,t is the packet loss of node n at time slot t; S3: Update the delay, backlog, reliability information and branch structure of each node to explore the action space of the improved D3QN and output the optimal action, which includes the following steps: S301: In the random action selection stage, further constraints are imposed on the action selection. For the physical resource block action space, if there are nodes in the slice that do not meet the service quality constraint condition, the allocation action with a value of 0 is removed, thereby obtaining a new action space a′; for the power action space, if the node does not meet the service quality constraint condition, the allocation action with a value of 0 is removed, thereby obtaining a new action space p′; S302: In the process of obtaining the optimal action, a total of N+1 actions are obtained by selecting the action corresponding to the maximum action value function in each branch except the branch that outputs the shared value; assuming that the optimal behavior of each branch is a * , then a * It is obtained by the following formula; The set of optimal actions is: in, represents the optimal action of branch i; S4: Train the network, calculate the loss function, and then use the gradient descent algorithm to update the network parameters until the network converges to obtain a trained network. Specifically, the following steps are included: S401: The loss function Loss is obtained based on the estimated value and the target value returned by the target value network: Where Z represents the number of branches; S402: Based on the loss function of step S401, calculate the gradient loss function for: in, represents the gradient vector of the estimation network, θ d represents the network parameters of branch d; S403: Every several steps, the network parameters of the estimated network are updated to the target network until the network converges.