Resource management method and system based on 5g power virtual private network
By employing an adaptive hierarchical resource management strategy that combines dual deep Q networks and multi-agent priority experience replay with the action actor critic algorithm in 5G power virtual private networks, the issues of service quality and performance stability in 5G power virtual private network resource management are resolved, achieving high efficiency and stability in dynamic resource allocation.
Patent Information
- Application Number
- CN202411771157.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-04
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2044-12-04
AI Technical Summary
Existing 5G power virtual private network resource management methods fail to effectively meet users' quality of service requirements and ensure the stability of service performance, neglecting the impact of network dynamism on service performance.
A resource management approach based on 5G power virtual private network is adopted. By establishing a 5G power virtual private network-slicing system, the dual deep Q network algorithm is used to allocate bandwidth resources to slices in the long time domain. Combined with the multi-agent priority experience playback and action actor commentator algorithm, resource blocks and transmission power are allocated to power end users in the short time domain, forming an adaptive hierarchical resource management strategy.
It achieves efficient allocation of dynamically scheduled resources, meeting users' service quality requirements while ensuring service performance stability.
Smart Images

Figure CN119815549B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of power virtual private network, and particularly relates to a resource management method and system based on 5G power virtual private network. BACKGROUND
[0002] 5G power virtual private network refers to that in a 5G network, physical facilities are divided into multiple logically isolated soft slices through software-defined network technology, in addition, hard slices can be formed through physical isolation, and then a special network for the power industry is virtually formed in the wireless access network, transmission network and core network. With the development of new services such as enhanced mobile broadband, low latency and high reliability communication and massive machine type communication, the power virtual private optical network communication technology has gradually been unable to meet the increasingly differentiated business needs of emerging businesses. On this basis, it is urgent to explore new communication technologies to meet the complex and diverse needs of new power grid businesses. 5G power virtual private network includes two types of virtual private networks, namely local virtual private network and wide area virtual private network. The local virtual private network mainly faces local tasks in a closed scene, and realizes the needs of power business through 5G power virtual private network and base station and mobile edge computing. The wide area virtual private network faces tasks in an open scene, and usually realizes end-to-end differentiated services through power slice private network.
[0003] Network slicing is one of the key technologies of 5G, which is a kind of on-demand networking method. Multiple virtual end-to-end networks are separated on a unified infrastructure, and each network slice is logically isolated in the wireless access network, bearer network and core network to adapt to various types of applications. With the expansion of the power system and the introduction of new applications, different applications have different needs for the network. Network slicing can provide dedicated network resources for each application to ensure its service quality and performance, and provide ultra-low latency communication services for tasks to ensure timely response and operation.
[0004] However, the existing work only allocates resources according to real-time network information to ensure the QoS (Quality of Service) requirements of services, ignoring the influence of network dynamics on service performance. Real-time allocation of resources according to the service request at the current time cannot guarantee the stability of service performance within a certain time range. SUMMARY
[0005] The technical problem to be solved by the present application is to provide a resource management method and system based on 5G power virtual private network, which can meet the service quality requirements of users while ensuring the stability of service performance.
[0006] To solve the above technical problems, a technical solution adopted by the present application is:
[0007] A resource management method based on 5G power virtual private network, comprising the steps of:
[0008] establishing a 5G power virtual private network-slice system, the 5G power virtual private network-slice system comprising a plurality of base stations, a plurality of power terminal users and a plurality of slices;
[0009] establishing an objective function for maximizing system benefits, and establishing constraint conditions of the objective function;
[0010] splitting the objective function into an upper objective function and a lower objective function;
[0011] allocating bandwidth resources for the slices by the base stations in a long time domain according to resource requirements based on the upper objective function by using a double deep Q network algorithm, obtaining an overall resource allocation strategy, and allocating resource blocks and transmission power for the power terminal users in a short time domain by the slices based on the lower objective function by using a multi-agent priority experience replay combined with an actor-critic algorithm, obtaining a fine-grained resource division strategy.
[0012] To solve the above technical problems, another technical solution adopted by the present application is:
[0013] A resource management system based on a 5G power virtual private network, comprising a memory, a processor and a computer program stored on the memory and executable on the processor, wherein the processor implements the following steps when executing the computer program:
[0014] establishing a 5G power virtual private network-slice system, the 5G power virtual private network-slice system comprising a plurality of base stations, a plurality of power terminal users and a plurality of slices;
[0015] establishing an objective function for maximizing system benefits, and establishing constraint conditions of the objective function;
[0016] splitting the objective function into an upper objective function and a lower objective function;
[0017] allocating bandwidth resources for the slices by the base stations in a long time domain according to resource requirements based on the upper objective function by using a double deep Q network algorithm, obtaining an overall resource allocation strategy, and allocating resource blocks and transmission power for the power terminal users in a short time domain by the slices based on the lower objective function by using a multi-agent priority experience replay combined with an actor-critic algorithm, obtaining a fine-grained resource division strategy.
[0018] The beneficial effects of the present application are that a 5G power virtual private network-slice system including multiple base stations, multiple power terminal users and multiple slices is established, a system benefit maximization objective function is established, constraint conditions of the objective function are established, the objective function is split into an upper objective function and a lower objective function, bandwidth resources are allocated for the slices by the base stations in a long time domain based on the upper objective function by using a double deep Q network algorithm according to resource requirements, an overall resource allocation strategy is obtained, and resource blocks and transmission power are allocated for the power terminal users by the slices in a short time domain based on the lower objective function by using a multi-agent priority experience replay combined with an actor critic algorithm, a fine-grained resource division strategy is obtained, the present application adopts adaptive hierarchical resource management, the double-network structure formed by the target network and the training network in the upper strategy DDQN algorithm solves the problem that the DQN estimated reward is higher than the actual reward, and makes the network structure more stable and makes a better overall resource allocation strategy, in the lower strategy, since discrete and continuous actions are involved at the same time, the multi-agent priority experience replay combined with the actor critic algorithm solves the problem that both discrete actions and continuous actions are involved, can consider and allocate in fine granularity, and realizes dynamic scheduling of resources, so that the service performance is stable while meeting the service quality requirements of users. BRIEF DESCRIPTION OF DRAWINGS
[0019] Figure 1 A step flowchart of a resource management method based on a 5G power virtual private network for an embodiment of the present application;
[0020] Figure 2 A structural schematic diagram of a resource management system based on a 5G power virtual private network for an embodiment of the present application;
[0021] Figure 3 A 5G power virtual private network-slice system in a resource management method based on a 5G power virtual private network for an embodiment of the present application;
[0022] Figure 4 A double deep Q network algorithm and a multi-agent priority experience replay combined with an actor critic algorithm in a resource management method based on a 5G power virtual private network for an embodiment of the present application. DETAILED DESCRIPTION
[0023] To explain the technical content, purposes and effects of the present application in detail, the following will be described in combination with the embodiments and the accompanying drawings.
[0024] Please refer to Figure 1 A resource management method based on a 5G power virtual private network, comprising the steps of:
[0025] A 5G power virtual private network-slice system is established, and the 5G power virtual private network-slice system comprises a plurality of base stations, a plurality of power terminal users and a plurality of slices;
[0026] A target function is established to maximize system benefits, and constraint conditions of the target function are established;
[0027] The target function is split into an upper target function and a lower target function;
[0028] Based on the upper target function, bandwidth resources are allocated for the slices by the base stations in a long time domain according to resource requirements by using a double deep Q network algorithm, an overall resource allocation strategy is obtained, and based on the lower target function, resource blocks and transmission power are allocated for the power terminal users by the slices in a short time domain by using a multi-agent priority experience replay combined with an action actor critic algorithm, and a fine-grained resource division strategy is obtained.
[0029] From the above description, the beneficial effects of the present application are that a 5G power virtual private network-slice system is established, and the 5G power virtual private network-slice system comprises a plurality of base stations, a plurality of power terminal users and a plurality of slices, a target function is established to maximize system benefits, and constraint conditions of the target function are established, the target function is split into an upper target function and a lower target function, bandwidth resources are allocated for the slices by the base stations in a long time domain according to resource requirements by using a double deep Q network algorithm based on the upper target function, an overall resource allocation strategy is obtained, and resource blocks and transmission power are allocated for the power terminal users by the slices in a short time domain by using a multi-agent priority experience replay combined with an action actor critic algorithm based on the lower target function, and a fine-grained resource division strategy is obtained, the present application adopts adaptive hierarchical resource management, the double-network structure formed by the target network and the training network in the upper-layer strategy DDQN algorithm solves the problem that the DQN estimated reward is higher than the actual reward, and makes the network structure more stable, and makes a better overall resource allocation strategy, in the lower-layer strategy, since discrete and continuous actions are involved at the same time, the multi-agent priority experience replay combined with the action actor critic algorithm solves the problem that both discrete actions and continuous actions are involved, fine-grained consideration and allocation can be performed, dynamic scheduling resources are realized, and therefore, the quality of service demand of users is met while ensuring stable service performance.
[0030] Further, the establishment of the target function to maximize system benefits comprises:
[0031]
[0032] In the formula, The long-time-domain bandwidth reservation set is represented as The short-time-domain resource block allocation set is represented as The user power allocation set is represented as T represents the number of long-time-domain, U lindicates the system benefit of the lth long time domain.
[0033] It can be known from the above description that optimizing resource allocation with the goal of maximizing system benefit can reduce system cost while providing optimal service performance.
[0034] Further, the constraint condition for establishing the target function comprises:
[0035] The resource block allocation constraint, the transmission queue constraint, the allocation constraint of each resource block in a single time slot, the data packet transmission constraint, the slice occupied frequency band resource sum constraint, the power terminal user bandwidth sum constraint, the quality of service requirement constraint, the power terminal user transmission rate constraint and the power terminal user transmission power constraint are established.
[0036] It can be known from the above description that the resource block allocation constraint, the transmission queue constraint, the allocation constraint of each resource block in a single time slot, the data packet transmission constraint, the slice occupied frequency band resource sum constraint, the power terminal user bandwidth sum constraint, the quality of service requirement constraint, the power terminal user transmission rate constraint and the power terminal user transmission power constraint ensure the feasibility of the finally obtained resource management strategy.
[0037] Further, the upper target function is specifically:
[0038]
[0039] In the formula, indicates the allocation set of user power after converting continuous variables into discrete variables, ω2 indicates the payment resource cost function weight, ω3 indicates the overload penalty weight, indicates the cost function of resources paid by the slice in the lth long time domain, indicates the overload penalty of the slice n in the lth long time domain, indicates the slice set;
[0040] The lower target function is specifically:
[0041]
[0042] In the formula, ω1 indicates the system total benefit weight, ΔT indicates the multiple relationship of time span, and REtotal(t) indicates the system total benefit.
[0043] It can be known from the above description that since there is coupling between different levels, the target function is divided into two-level sub-problems, i.e., the upper target function and the lower target function, so as to meet the dynamic traffic demand and the differentiated service demand of the power terminal user, ensure the stability of service performance and meet the QoS.
[0044] Further, the base station allocates bandwidth resources for the slices in a long time domain according to resource requirements based on the upper target function, obtains a general resource allocation strategy, and allocates resource blocks and transmission power for the power terminal users in a short time domain based on the lower target function using the slices, to obtain a fine-grained resource division strategy, which includes:
[0045] determining a state set, an action set, and a reward function based on the upper target function and the lower target function;
[0046] defining a loss function;
[0047] minimizing the gradient of the loss function using an Adam algorithm, and updating Q network parameters of a current Q network based on the state set, the action set, and the reward function to obtain a target Q network;
[0048] allocating bandwidth resources for the slices in a long time domain according to resource requirements by the base station based on the target Q network, to obtain a general resource allocation strategy;
[0049] regarding each base station as an agent, and determining a discrete action that maximizes a Q function value of the agent in a current state;
[0050] for training of an actor-critic network, the agent combines a state and the general resource allocation strategy, and takes the state and the general resource allocation strategy as inputs of the actor network, and the actor-critic network outputs selection of power and observation frequency according to the inputs;
[0051] the agent combines the state, the selected continuous action, and the made fine-grained resource division strategy as inputs of the critic network, and the critic network outputs an estimated Q value function of the fine-grained resource division strategy according to the inputs, and determines an optimal fine-grained resource division strategy according to the obtained estimated Q value function;
[0052] the agent takes a corresponding action according to maximization of expected rewards, and updates the actor-critic network according to a priority experience replay mechanism according to a policy gradient;
[0053] for training of the critic network, an error between a target Q value and an actual Q value of the critic network is minimized as a target, and parameters of the critic network are updated according to a priority experience replay mechanism by a gradient descent method;
[0054] allocating resource blocks and transmission power for the power terminal users in a short time domain using the slices according to the trained actor-critic network, to obtain a fine-grained resource division strategy.
[0055] As can be known from the above description, the DDQN algorithm uses the double-network structure formed by the target network and the training network to solve the problem that the DQN estimates a reward higher than the actual reward, and makes the network structure more stable, while the single DDQN algorithm is not suitable for resource scheduling in a short time domain, so the multi-agent priority experience replay is used in the lower layer to combine with the actor-critic algorithm to dynamically schedule resources, support diversified power services with different QoS requirements, and guarantee fine-grained division of resources, meet the dynamic traffic demand and differentiated service demand of power terminal users.
[0056] Further, the determining the state set, the action set and the reward function based on the upper-layer target function and the lower-layer target function comprises:
[0057] determining a first state set composed of a communication resource remaining condition, a resource block allocation state and a slice set based on the upper-layer target function, and determining a second state set composed of a slice reserved resource condition, a user quantity and a user arrival rate based on the lower-layer target function;
[0058] determining a first action set as completing slice reserved communication resource division at an l-th long time domain moment based on the upper-layer target function, and determining a second action set composed of t time slot slice frequency band resource block scheduling and user equipment transmit power allocation based on the lower-layer target function;
[0059] determining a target value of the upper-layer target function as a first reward function, and determining a target value of the lower-layer target function as a second reward function.
[0060] As can be known from the above description, the Markov process is constructed by determining the state set, the action set and the reward function, which can effectively simplify the complexity of the optimization problem, improve the resource management efficiency, and timely allocate resources.
[0061] Further, the defining the loss function comprises:
[0062]
[0063] In the formula, L(μ) represents the loss function, Q(s,a,μ) represents the current Q network, s represents the state, a represents the selected action, μ represents the neural network parameter of the training network, r represents the obtained reward, γ S represents the weight, represents the target Q network, s' represents the state of the target Q network, and μ' represents the neural network parameter of the target Q network.
[0064] As can be known from the above description, the loss function using the double-network design can measure the policy, train the policy towards the optimal policy direction, and ensure the reliability of the overall resource allocation policy.
[0065] Further, the determining the discrete action of the agent that maximizes the Q function value in the current state comprises:
[0066]
[0067] wherein a m (t) represents the discrete action of the agent m that maximizes the Q function value in the current state s(t), represents the optimal overall resource allocation strategy, a s,m (t) represents the action of the agent m in the state s, μ m represents the parameter of the actor network, represents the parameter of the critic network, represents the maximum Q value.
[0068] As can be seen from the above description, the determining the discrete action of the agent that maximizes the Q function value in the current state can maximize the expected return of the agent, optimize the policy function, and improve the decision efficiency.
[0069] Further, the method further comprises:
[0070] In the priority experience replay mechanism, the importance of the rth experience is measured by the absolute value of the TD-error, specifically:
[0071]
[0072] wherein δ r represents the importance of the rth experience, represents the target Q value, represents the actual Q value, represents the rth state of the agent m, represents the rth action of the agent m, represents the parameter of the critic network;
[0073] The sampling probability is determined according to the importance of the rth experience, specifically:
[0074]
[0075] p r = |δ r | + ζ;
[0076] wherein P(r) represents the sampling probability, a represents the priority decision parameter for execution, p r represents the priority weight of the rth experience, and ζ represents a preset positive value.
[0077] From the above description, by introducing the priority experience replay mechanism to increase the sampling probability of important experiences, these experiences are used more frequently in the training process, the agent can focus more on those experiences that have an important influence on its decision-making, thereby improving the learning efficiency, and sampling according to the priority of the experience makes the data distribution more stable, which helps to improve the stability and effect of training.
[0078] Please refer to Figure 2 Another embodiment of the present application provides a resource management system based on a 5G power virtual private network, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement each step of the above-mentioned resource management method based on a 5G power virtual private network.
[0079] The above-mentioned resource management method and system based on a 5G power virtual private network can be applied to the resource management scenario of a 5G power virtual private network, which will be described in detail through the specific embodiments below:
[0080] Please refer to Figure 1 , Figure 3 and Figure 4 The first embodiment of the present application is:
[0081] A resource management method based on a 5G power virtual private network, comprising the steps of:
[0082] S1, establishing a 5G power virtual private network-slice system, the 5G power virtual private network-slice system comprising a plurality of base stations, a plurality of power terminal users and a plurality of slices.
[0083] Specifically, a plurality of power terminal users and a plurality of base stations are determined, the base stations are responsible for communication with the power terminal users within their coverage range, and the power terminal users all have their own specific service requirements. The basic network infrastructure is divided into a plurality of virtual networks, i.e. slices, to support different user services. A plurality of groups of associated power terminal users are determined, and each group of associated power terminal users is allocated a base station and a slice. The 5G power virtual private network-slice system is established according to the plurality of base stations, the plurality of power terminal users and the plurality of slices, as shown in Figure 3 .
[0084] Among them, the plurality of base stations constitute a base station set, denoted as m represents the mth base station, and M represents the total number of base stations. The plurality of power terminal users constitute a power terminal set, denoted as i represents the ith power terminal user, and I represents the total number of power terminal users. Each power terminal user has a base station providing service, and each base station has a group of associated power terminal users, denoted as . The plurality of slices constitute a slice set, denoted as n represents the nth slice, N represents the total number of slices, each slice has a group of associated power terminal users, and the total number of power terminal users is represented by U.
[0085] S2, a target function is established to maximize system benefit, and a constraint condition of the target function is established.
[0086] The establishment of the target function to maximize system benefit includes:
[0087]
[0088] In the formula, indicates a long-time domain bandwidth reservation set, indicates a short-time domain resource block allocation set, indicates a user power allocation set, T indicates a long-time domain quantity, and U l indicates the system benefit of the lth long-time domain.
[0089] The constraint condition of the target function includes:
[0090] The resource block allocation constraint, the transmission queue constraint, the allocation constraint of each resource block in a single time slot, the data packet transmission constraint, the slice occupied frequency band resource sum constraint, the power terminal user bandwidth sum constraint, the quality of service requirement constraint, the power terminal user transmission rate constraint, and the power terminal user transmission power constraint are established.
[0091] The resource block allocation constraint indicates that the resource block allocation variable is a binary variable, and specifically:
[0092]
[0093] The transmission queue constraint ensures the stability of the transmission queue, and specifically:
[0094]
[0095] The allocation constraint of each resource block in a single time slot ensures that each resource block is only allocated to a unique power terminal user in a single time slot, and specifically:
[0096]
[0097] The data packet transmission constraint ensures that the data packet transmission process meets the AoI requirement, and specifically:
[0098]
[0099] The slice occupied frequency band resource sum constraint ensures that the total sum of the slice occupied frequency band resources does not exceed the total frequency band resources, and specifically:
[0100]
[0101] In the formula, The sum of the power terminal user bandwidths represents the total frequency band resource.
[0102] The sum of the power terminal user bandwidths represents the total frequency band resource.
[0103]
[0104] The quality of service requirement constraint ensures that all devices meet the QoS requirement, specifically:
[0105]
[0106] The power terminal user transmission rate constraint is specifically:
[0107]
[0108] The power terminal user transmission rate constraint is specifically:
[0109]
[0110] In the formula, P max The power terminal user transmission rate constraint is specifically:
[0111] Due to the coupling between different levels, the optimization problem of the above objective function is divided into double-level sub-problems. In the upper-level strategy, the system makes slice resource allocation strategy in a long time domain. For such problems of simultaneously optimizing discrete and continuous variables, they need to be simplified before they can be solved using deep reinforcement learning. The focus is on the overall resource allocation strategy, which only needs to reserve certain resources for slices without considering and allocating in detail, and can accept the small loss caused by the conversion of continuous variables to discrete variables. The continuous variables are converted to discrete variables The set of which is Therefore, the above objective function is divided into an upper-level objective function and a lower-level objective function, as described in S3.
[0112] S3, the target function is divided into an upper-level target function and a lower-level target function.
[0113] The upper-level target function is specifically:
[0114]
[0115] In the formula, The set of the user power allocation after converting the continuous variables to discrete variables is ω1, ω2 represents the payment resource cost function weight, and ω3 represents the overload penalty weight, a cost function representing the cost of resource payment for slice n in the lth long time domain, a cost function representing the overload penalty of slice n in the lth long time domain, a set of slices;
[0116] The lower-layer objective function is specifically:
[0117]
[0118] In the formula, ω1 represents the total system revenue weight, ΔT represents the multiple relationship of the time span, and REtotal(t) represents the total system revenue.
[0119] The resource management process is divided into two levels of inter-slice and intra-slice. At the inter-slice level, the base station allocates bandwidth resources for the slices according to the resource demand. At the intra-slice level, the slice performs a resource block (RB) allocation process to allocate necessary RBs for each end power terminal user, and each slice allocates a certain number of RBs for each power terminal user on each slice from the pre-allocated bandwidth resources to meet various QoS requirements of each end power terminal user in terms of data rate and delay.
[0120] In order to adapt to the dynamic topology change of the network, fine-grained resource scheduling is required, and the cost of frequent slice reconfiguration is too high and the problem of resource waste is avoided. The present application performs resource management under a double time domain framework, as shown in the formula: Figure 4 The short time domain is defined as STI, represented by t∈{1,2,...}, and the fixed duration between time slot t and the next time t+1 is τ; the long time domain is defined as LTI, represented by l∈{1,2,...}, and the duration is ΔT STIs. In each LTI, the system re-determines the association strategy of the power terminal user and the base station, and the bandwidth resources and computing resources reserved for each slice. In each STI, in order to ensure that the power terminal user meets the stability and dynamics in each time slot t, the system will perform resource scheduling within the slice according to the terminal user state.
[0121] In the long time domain, is the bandwidth resource reserved for slice n in the lth LTI, and the bandwidth resources reserved for different slices are obtained The number of RBs in slice n can be calculated as In the short time domain, the remaining situation of the resource block in slice n is represented as R n ; The allocation of RBs is represented by a binary variable used to represent the allocation of the kth RB in slice n in time slot t, if it is allocated to the ith device, then otherwise
[0122] Assuming that the communication between the power end user i and the base station m includes the uplink communication of transmitting the user dynamic service request to the base station and the downlink communication of sending the result to the user, the present application only considers the uplink communication delay in the communication model, then the SINR (Signal to Interference plus Noise Ratio) that the power end user i and the base station m can achieve on the time slot t slice n is:
[0123]
[0124] In the formula, SINR i (t) represents the SINR that the power end user i and the base station m can achieve on the time slot t slice n, p i (t) represents the transmission power of the power end user i, represents the channel power gain of the power end user i, represents the path loss, and a represents the exponent of the path loss, represents the interference noise, and σ 2 represents the additive Gaussian white noise power, d i,m (t) represents the distance between the power end user i and the base station m, x u represents the horizontal coordinate of the power end user i, y u represents the vertical coordinate of the power end user i, x m represents the horizontal coordinate of the base station m, y m represents the vertical coordinate of the base station m.
[0125] B represents the bandwidth of one RB, and the transmission rate that the user equipment can achieve on one resource block is r B , which is represented as:
[0126] r B (t) = Blog2(1 + SINR i (t));
[0127] Therefore, the wireless communication rate of the power end user is represented as:
[0128]
[0129] In the formula, K represents the total number of resource blocks possessed by the slice n.
[0130] The transmission rate requirement of the i-th power end user accessing to the slice n is represented as The number of RBs required by the i-th power end user is In order to satisfy that the number of required RBs is an integer, it is defined that is the floor function, i.e. if x is an integer, then is x itself, if is a non-integer, then
[0131] Suppose the size of the data packet composed of application and data sent to the base station is D i (t), the base station receives D i (t) and returns a feedback signal ACK, after the power terminal user equipment receives the feedback signal, it transmits the next data packet. If the received data packet has an error code, it will discard the data packet and feedback to the terminal to retransmit the error data packet; if SINR is less than the threshold value Γ th , a binary variable γ i (t) = 1 is used, i.e. the information transmitted by the power terminal user equipment i at time slot t can withstand interference, so the data can be successfully decoded at the base station side, otherwise γ i (t) = 0, which is:
[0132]
[0133] When the data packet is transmitted, the transmission rate also directly affects whether the data packet can successfully arrive at the base station. When the transmission rate is too small, the data cannot be completely transmitted within a unit time interval τ, which may cause data loss or the need to split the data packet for transmission, so this situation will affect whether the data packet can successfully arrive at the base station. A binary variable ο i (t) = 1 is used to indicate that the sensor data is successfully transmitted within a unit time slot, otherwise ο i (t) = 0, which is:
[0134]
[0135] Suppose the tasks between the power terminal users in the virtual power private network are independent, the task arrival rate of the power terminal user i at time slot t follows a Poisson process, denoted as λ i (t). The power terminal user has different task arrival rates in the present application, so the overall arrival rate can be represented as In order to evaluate the average transmission delay, the process of transmitting information from the power terminal user equipment to the base station is modeled as an M / M / 1 queue. Due to the randomness of the wireless channel, it is assumed that the average service rate follows an exponential distribution, and the average service rate can be represented as:
[0136]
[0137] According to the M / M / 1 queuing theory, the average transmission delay can be defined as: Tup(t) = 1 / (μ(t)-λ(t)). In order to ensure the stability of the transmission queue, the following condition must be met:
[0138]
[0139] AoI is a measure of information timeliness, defined as the time elapsed since the latest received data packet by the target node since its generation. With Δ i (t) measures the AoI of the data packet generated by the power end-user equipment i. When the power end-user equipment i transmits the data packet to the base station, the AoI value Δ i (t) of the power end-user equipment at time slot t is represented as:
[0140]
[0141] The average AoI of the power end-user equipment i at the DT synchronization is defined as: is: The AoI of all users should meet the target constraint: wherein, represents the AoI of the user , and represents the upper limit value of the AoI of the user on slice n.
[0142] Assuming that the revenue of the system with respect to the transmission AoI is related to the transmission AoI, the revenue is proportional to the difference between the required and obtained AoI of the system, and the revenue increases with the increase of the AoI difference, represented as:
[0143]
[0144] wherein, RE age (t) represents the system revenue related to the transmission AoI, χ in represents the unit income of the system with respect to the delay, and ξ max represents the upper limit value of the AoI of all users.
[0145] Similarly, the revenue of the system with respect to the transmission error is related to the transmission process of the user data packet, assuming that η in and ε in are the unit incomes of the data decoding and the successful transmission of the data packet in the transmission process of the system, specifically:
[0146]
[0147] wherein, RE tr (t) represents the system revenue with respect to the transmission error, D i (t) represents the data packet.
[0148] Therefore, the total revenue of the system is: REtotal(t) = β1REage(t) + β2REtr(t), wherein β1 represents a first weight coefficient, β2 represents a second weight coefficient, and β1 + β2 = 1.
[0149] The cost function of the resource payment of the slice in the lth long time domain, i.e., the resource allocation cost, is: In the formula, c sl represents the use cost of a unit bandwidth, l sl represents the unit cost of slice scaling.
[0150] It is assumed that is the bandwidth resource reservation rate of the slice n in the lth LTI time, and the calculation formula is:
[0151]
[0152] It is assumed that α E represents a slice resource reservation alert value, and a part of the resource reservation rate that is less than the threshold is punished, and the more the difference from the alert value is, the more the punishment is, and the overload penalty of the slice n in the lth LTI time is Specifically, it is:
[0153]
[0154] In the formula, ε E represents a unit penalty of the part of the resource reservation rate that is less than the threshold, α B represents a minimum threshold of the slice resource reservation rate. The slice resource reservation rate of each slice must be greater than or equal to α B , and when α B is 0, it means that no reservation allocation is allowed to the resources in the slice without considering future conditions.
[0155] The networked power terminal user transmits tasks by accessing the base station in the coverage range. In the upper layer strategy, the base station allocates bandwidth resources for the slice in the long time domain according to the resource demand, dynamically adjusts the slice resources reserved for the service, and establishes resource constraints; in the lower layer strategy, the information age (Age of Information, AoI) is used to ensure the determinacy of the transmission delay, and the slice allocates a certain number of resource blocks for the user in the short time domain, and further provides reliable communication for the user through more flexible resource scheduling, as described below.
[0156] S4, based on the upper layer objective function, the base station allocates bandwidth resources for the slices in a long time domain according to resource requirements by using a double deep Q network algorithm, obtains an overall resource allocation strategy, and based on the lower layer objective function, allocates resource blocks and transmit power for the power terminal users in a short time domain by using the slices by using a multi-agent priority experience replay combined with an actor-critic algorithm, and obtains a fine-grained resource division strategy, as shown in Figure 4 Specifically, S4.1-S4.10 are included.
[0157] S4.1, based on the upper layer objective function and the lower layer objective function, determine a state set, an action set, and a reward function, specifically including S4.1.1-S4.1.3:
[0158] S4.1.1, based on the upper layer objective function, determine a first state set composed of communication resource remaining conditions, resource block allocation states, and slice sets, and based on the lower layer objective function, determine a second state set composed of slice reserved resource conditions, user numbers, and user arrival rates.
[0159] The first state set is The second state set is
[0160]
[0161] S4.1.2, based on the upper layer objective function, determine a first action set as completing slice reserved communication resource division at the lth long time domain moment, and based on the lower layer objective function, determine a second action set composed of t time slot slice frequency band resource block scheduling and user equipment transmit power allocation.
[0162] The first action set is The second action set is
[0163]
[0164] S4.1.3, determine the target value of the upper layer objective function as a first reward function R l , that is and determine the target value of the lower layer objective function as a second reward function R s , that is R s = ω1RE total (t).
[0165] S4.2, define the loss function, DDQN uses the double network structure formed by the target network and the training network, solves the problem that the DQN estimates the reward higher than the actual reward, and makes the network structure more stable, so the loss function using the double network design can measure the strategy, and train the strategy towards the optimal strategy direction, specifically:
[0166]
[0167] In the formula, L(mu) represents the loss function, Q(s,a,mu) represents the current Q network, s represents the state, a represents the selected action, mu represents the neural network parameters of the training network, r represents the obtained reward, and gamma S represents the weight, represents the target Q network, s' represents the state of the target Q network, and mu' represents the neural network parameters of the target Q network.
[0168] S4.3, adopt Adam algorithm to minimize the gradient of the loss function, and update the Q network parameters of the current Q network based on the state set, action set and reward function, to obtain the target Q network.
[0169] The gradient of the loss function is specifically:
[0170]
[0171] In the formula, represents the gradient of the loss function, represents the gradient of the current Q network.
[0172] S4.4, using the target Q network, the base station allocates bandwidth resources to the slice in a long time domain according to resource requirements, and obtains an overall resource allocation strategy.
[0173] Specifically, using the target Q network, the SDN controller controls the base station to allocate bandwidth resources to the slice in a long time domain according to resource requirements, and obtains an overall resource allocation strategy.
[0174] Single DDQN algorithm is no longer applicable to resource scheduling in short time domain, so multi-agent priority experience replay combined with action actor critic algorithm PER-MACA2C is adopted, as follows.
[0175] S4.5, each base station is regarded as an agent, and the discrete action that makes the Q function value maximum in the current state of the agent is determined.
[0176] The determination of the discrete action that makes the Q function value maximum in the current state of the agent includes:
[0177]
[0178] where a m (t) denotes the discrete action of the agent m that maximizes the Q function value at the current state s(t), denotes the optimal overall resource allocation policy, a s,m (t) denotes the action of the agent m at state s, μ m denotes the parameters of the actor network, denotes the parameters of the critic network, denotes the maximum Q value.
[0179] S4.6, for the training of the actor network, the agent m combines the state s(t) and the overall resource allocation policy as inputs of the actor network, and the actor network outputs the selection of power and observation frequency according to the inputs.
[0180] S4.7, the agent m combines the state, the selected continuous action and the made fine-grained resource partitioning policy as inputs of the critic network, and the critic network outputs the estimated Q value function of the fine-grained resource partitioning policy according to the inputs, and determines the optimal fine-grained resource partitioning policy according to the obtained estimated Q value function.
[0181] S4.8, the agent m takes the corresponding action according to the maximization of the expected reward, and updates the actor network according to the priority experience replay mechanism according to the policy gradient.
[0182] wherein the policy gradient is specifically:
[0183]
[0184] denotes the policy gradient, and r denotes the index of experience, denotes the gradient of the policy function, denotes the gradient of the action value function Q with respect to the action a m .
[0185] S4.9, for the training of the critic network, the parameters of the critic network are updated according to the priority experience replay mechanism by the gradient descent method, aiming at minimizing the error between the target Q value and the actual Q value of the critic network.
[0186] The output y of the target critic network is specifically: m
[0187]
[0188] wherein r m represents the reward of the agent m, and ε represents a weight value, represents the Q value function of the target network of the critic network, represents the target network parameter of the actor network, represents the target network parameter of the critic network, so as to reduce the tendency of overestimation of the Q value.
[0189] the error between the target Q value and the actual Q value of the critic network Specifically,
[0190]
[0191] In the formula, represents the target Q value, represents the actual Q value.
[0192] In addition,
[0193] In the priority experience replay mechanism, the importance of the rth experience is measured by the absolute value of the TD-error, and specifically,
[0194]
[0195] In the formula, δ r represents the importance of the rth experience, represents the rth state of the agent m, represents the rth action of the agent m, represents the parameter of the critic network;
[0196] The greater the TD-error, the greater the possibility of experience replay, and the sampling probability is determined according to the importance of the rth experience, and specifically,
[0197]
[0198] p r = |δ r | + ζ;
[0199] In the formula, P(r) represents the sampling probability, α represents the priority decision parameter of execution, p r represents the priority weight of the rth experience, and ζ represents a preset positive value, that is, a preset positive number, which is used to prevent the probability of sample r from being 0.
[0200] S4.10, according to the actor network trained, the slice is used to allocate resource blocks and transmit power in a short time domain for the power terminal user, and a fine-grained resource allocation strategy is obtained.
[0201] Please refer to Figure 2 Embodiment two of the present application is:
[0202] A resource management system based on a 5G power virtual private network, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements each step of the above-mentioned resource management method based on a 5G power virtual private network when executing the computer program.
[0203] To sum up, the present application provides a resource management method and system based on a 5G power virtual private network, establishes a 5G power virtual private network-slice system including multiple base stations, multiple power terminal users and multiple slices, establishes an objective function to maximize system efficiency, establishes the constraint conditions of the objective function, splits the objective function into an upper objective function and a lower objective function, allocates bandwidth resources for the slices in the long time domain based on the upper objective function through the base stations using a double deep Q network algorithm according to resource requirements, obtains an overall resource allocation strategy, and allocates resource blocks and transmit power for the power terminal users in the short time domain based on the lower objective function using the slices using a multi-agent priority experience replay combined with an actor-critic algorithm, to obtain a fine-grained resource allocation strategy. The present application uses adaptive hierarchical resource management, the double-network structure formed by the target network and the training network in the upper strategy DDQN algorithm solves the problem that the DQN estimated reward is higher than the actual reward, and makes the network structure more stable and makes a better overall resource allocation strategy. In the lower strategy, since both discrete and continuous actions are involved, the multi-agent priority experience replay combined with the actor-critic algorithm solves the problem of containing both discrete and continuous actions, can consider and allocate in fine granularity, and realizes dynamic scheduling of resources, so as to meet the service quality requirements of users while ensuring stable service performance. In addition, the DDQN algorithm uses the double-network structure formed by the target network and the training network to solve the problem that the DQN estimated reward is higher than the actual reward, and makes the network structure more stable. However, a single DDQN algorithm is not suitable for resource scheduling in the short time domain, so the multi-agent priority experience replay combined with the actor-critic algorithm is used in the lower layer to dynamically schedule resources, support diversified power services with different QoS requirements, and ensure fine-grained division of resources, meet the dynamic traffic demand and differentiated service requirements of power terminal users.
[0204] The above only describes the embodiments of the present application, and does not limit the patent scope of the present application, any equivalent transformation or direct or indirect application in related technical fields based on the content of the present application specification and drawings are also included in the patent protection scope of the present application.
Claims
1. A resource management method based on a 5G power virtual private network, characterized in that, Including the following steps: A 5G power virtual private network-slicing system is established, which includes multiple base stations, multiple power terminal users, and multiple slices. An objective function is established to maximize system benefits, and constraints on the objective function are also established. The objective function is split into an upper-level objective function and a lower-level objective function; Based on the upper-layer objective function, the base station allocates bandwidth resources to the slice in the long-term domain using a dual deep Q network algorithm according to resource requirements, thus obtaining an overall resource allocation strategy. Based on the lower-layer objective function, the slice is used in the short-term domain to allocate resource blocks and transmission power to the power terminal user using a multi-agent priority experience replay combined with the action actor critic algorithm, thus obtaining a fine-grained resource partitioning strategy. The objective function established to maximize system benefits includes: ; In the formula, Represents the long-term bandwidth reservation set. Represents the short-time domain resource block allocation set. This represents the set of user power allocations, where T represents the number of long-time domains. Represents the system benefit in the l-th long time domain; The upper-level objective function is specifically as follows: ; In the formula, This represents the set of user power allocations after converting continuous variables into discrete variables. This represents the weight of the payment resource cost function. Indicates the weight of the overload penalty. This represents the cost function of resources required for a slice in the l-th long-term domain. This represents the overload penalty for slice n in the l-th long time domain. Represents a set of slices; The lower-level objective function is specifically as follows: ; In the formula, Indicates the weight of the total system revenue. Indicates a multiple relationship between time spans. This represents the total revenue of the system.
2. The resource management method based on a 5G power virtual private network according to claim 1, characterized in that, The constraints for establishing the objective function include: Establish constraints on resource block allocation, transmission queue, allocation of each resource block within a single time slot, data packet transmission, total bandwidth occupied by slices, sum of bandwidth for power end users, quality of service requirements, power end user transmission rate, and power end user transmit power.
3. The resource management method based on a 5G power virtual private network according to claim 1, characterized in that, The above-layer objective function is used by the base station to allocate bandwidth resources to the slice in the long-term domain according to resource requirements using a dual deep Q network algorithm, resulting in an overall resource allocation strategy. Then, based on the lower-layer objective function, the slice is used in the short-term domain to allocate resource blocks and transmit power to the power terminal user using a multi-agent priority experience replay combined with an action actor / critic algorithm, resulting in a fine-grained resource partitioning strategy, including: The state set, action set, and reward function are determined based on the upper-level objective function and the lower-level objective function. Define the loss function; The gradient of the loss function is minimized using the Adam algorithm, and the Q-network parameters of the current Q-network are updated based on the state set, action set, and reward function to obtain the target Q-network. By utilizing the target Q network and the base station to allocate bandwidth resources to the slice in the long-term time domain according to resource requirements, an overall resource allocation strategy is obtained. Each base station is treated as an agent, and the discrete actions of the agent in the current state that maximize the Q function value are determined. For training the actor's network, the agent combines the state and the overall resource allocation strategy, using the state and the overall resource allocation strategy as inputs to the actor's network, and the actor's network selects based on the input-output power and observation frequency; The agent combines the state, the selected continuous action, and the fine-grained resource allocation strategy as input to the critic network. The critic network outputs the estimated Q-value function of the fine-grained resource allocation strategy based on the input, and determines the optimal fine-grained resource allocation strategy based on the obtained estimated Q-value function. The agent takes corresponding actions based on maximizing expected reward, and updates the actor's network according to the policy gradient and the priority experience replay mechanism. The training of the critic network aims to minimize the error between the target Q-value and the actual Q-value of the critic network. The parameters of the critic network are updated using gradient descent based on a priority experience replay mechanism. Based on the trained actor home network, the slices are used to allocate resource blocks and transmission power to the power terminal users in the short time domain to obtain a fine-grained resource partitioning strategy.
4. A resource management method based on a 5G power virtual private network according to claim 3, characterized in that, The process of determining the state set, action set, and reward function based on the upper-level objective function and the lower-level objective function includes: Based on the upper-layer objective function, a first state set consisting of the remaining communication resources, resource block allocation status, and slice set is determined, and based on the lower-layer objective function, a second state set consisting of slice reserved resources, number of users, and user arrival rate is determined. Based on the upper-layer objective function, the first action set is determined to be the completion of the slice reserved communication resource allocation at the l-th long time domain time, and based on the lower-layer objective function, the second action set is determined to be composed of the t-slot slice frequency band resource block scheduling and user equipment transmit power allocation; The target value of the upper-level objective function is determined as the first reward function, and the target value of the lower-level objective function is determined as the second reward function.
5. A resource management method based on a 5G power virtual private network according to claim 3, characterized in that, The defined loss function includes: ; In the formula, Represents the loss function. This represents the current Q-network, where s represents the state and a represents the selected action. Let r represent the neural network parameters used to train the network, and r represent the reward obtained. Indicates weight, Indicates the target Q-network, Indicates the state of the target Q-network. This represents the neural network parameters of the target Q-network.
6. A resource management method based on a 5G power virtual private network according to claim 3, characterized in that, The discrete actions that maximize the Q-function value of the agent in the current state include: ; In the formula, This indicates the current state of agent m. The discrete action that maximizes the Q-function value. This represents the optimal overall resource allocation strategy. This represents the action of agent m in state s. This refers to the parameters of the actor's home network. This represents the parameters of the critic network. This represents the maximum Q value.
7. A resource management method based on a 5G power virtual private network according to claim 3, characterized in that, Also includes: In the aforementioned priority experience replay mechanism, the importance of the r-th experience is measured by the absolute value of the TD-error, specifically as follows: ; In the formula, This indicates the importance of the r-th experience. Indicates the target Q value. This represents the actual Q value. Let r represent the r-th state of agent m. This represents the r-th action of agent m. The parameters representing the critic network; The sampling probability is determined based on the importance of the r-th experience, specifically as follows: ; ; In the formula, Indicates the sampling probability. The parameter indicates the execution priority. Indicates the first Priority weight of secondary experience This indicates a preset positive value.
8. A resource management system based on a 5G power virtual private network, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements each step of the resource management method based on a 5G power virtual private network as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Power business matching method, electronic equipment and storage medium
CN116708181A
Dynamic slicing and resource allocation method for 5G power internet of things
CN117880898A