Network slice resource allocation method and device, storage medium and electronic equipment
By using a hierarchical reinforcement learning method, the allocation of spectrum resources between and within network slices is dynamically optimized, solving the complexity of resource allocation in 5G application scenarios and achieving efficient resource utilization and improved user satisfaction.
Patent Information
- Application Number
- CN202510325490.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-12-12
AI Technical Summary
How to achieve effective allocation of network slice resources, especially in complex and ever-changing 5G application scenarios, to meet the needs of different service types, reduce the overhead of resource reconfiguration between slices, and improve resource utilization.
A hierarchical reinforcement learning approach is adopted to decompose the spectrum resource allocation problem between and within network slices. By using reinforcement learning models such as the actor-critic algorithm or the reinforcement algorithm, combined with the hierarchical reinforcement learning framework, the resource allocation is dynamically optimized, thereby improving resource utilization and reducing the reconfiguration overhead between slices.
It achieves efficient resource allocation under complex environments and changing needs, improves user satisfaction and resource utilization, and reduces the overhead of resource reconfiguration between slices.
Smart Images

Figure CN121125498A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and particularly relates to a network slice resource allocation method and device, a storage medium and an electronic device. BACKGROUND
[0002] Network slicing (or simply referred to as slicing, etc.) is a key technology in 5G, which constructs multiple virtualized and mutually isolated logical networks on a physical network to meet the different service quality requirements of specific business types or industry users. At the same time, network slicing can also achieve effective isolation between resources to ensure that services within different slices do not affect each other.
[0003] However, how to effectively allocate network slice resources is a technical problem to be solved. SUMMARY
[0004] Therefore, the present application provides a network slice resource allocation method and device, a storage medium and an electronic device, which mainly aims to solve the technical problem of how to effectively allocate network slice resources.
[0005] In a first aspect, the present application provides a network slice resource allocation method, comprising:
[0006] obtaining first state information of each network slice, second state information of each terminal corresponding to each network slice, and third state information of spectrum resources of an access network device;
[0007] allocating spectrum resources of each network slice according to the first state information and the third state information, and allocating spectrum resources of each terminal corresponding to each network slice according to the spectrum resources allocated to each network slice and the second state information.
[0008] Optionally, allocating spectrum resources of each network slice according to the first state information and the third state information, and allocating spectrum resources of each terminal corresponding to each network slice according to the spectrum resources allocated to each network slice and the second state information, comprises:
[0009] determining spectrum resources of each network slice and spectrum resources of each terminal corresponding to each network slice by using a reinforcement learning model according to the first state information, the second state information and the third state information, and allocating spectrum resources according to the determination result.
[0010] Optionally, the training process of the reinforcement learning model comprises:
[0011] determining first parameter information, the first parameter information comprising at least parameters of a first action network, parameters of a second action network and parameters of a critic network, wherein the first action network is configured to perform actions of inter-network slice spectrum resource allocation, the second action network is configured to perform actions of intra-network slice spectrum resource allocation, and the critic network is configured to guide the actions of the first action network and the second action network;
[0012] performing iterative training on the reinforcement learning model based on the first parameter information using an actor-critic algorithm.
[0013] Optionally, the process of performing iterative training on the reinforcement learning model based on the first parameter information using an actor-critic algorithm comprises:
[0014] performing inter-network slice spectrum resource allocation using the first action network based on the first parameter information, and calculating a corresponding first reward value by a first reward function, wherein the first reward function is determined based on satisfaction of terminal users, network slice reconfiguration overhead, spectrum resource utilization and decision cost of option switching;
[0015] performing intra-network slice spectrum resource allocation using the second action network based on the first parameter information and spectrum resources allocated to the network slice, and calculating a corresponding second reward value by a second reward function, wherein the second reward function is determined based on satisfaction of terminal users;
[0016] calculating action value and bias of the second action network using the critic network, and updating parameters of the second action network according to the action value and the bias;
[0017] updating an experience pool based on actions of the first action network and the second action network, and combining the first reward value, the second reward value, state of inter-network slice after spectrum resource allocation and state of corresponding terminal;
[0018] when data in the experience pool reaches a preset storage threshold, updating parameters of the critic network and the first action network using the data in the experience pool, and emptying the experience pool after the updating is completed.
[0019] Optionally, the training process of the reinforcement learning model comprises:
[0020] determining second parameter information, the second parameter information comprising at least parameters of a first reinforce network and parameters of a second reinforce network, wherein the first reinforce network is configured to perform actions of inter-network slice spectrum resource allocation, and the second reinforce network is configured to perform actions of intra-network slice spectrum resource allocation.
[0021] performing iterative training on the reinforcement learning model based on the second parameter information using a reinforce algorithm.
[0022] Optionally, the process of performing iterative training on the reinforcement learning model based on the second parameter information using a reinforce algorithm comprises:
[0023] performing spectrum resource allocation between network slices based on the second parameter information using the first reinforce network;
[0024] performing spectrum resource allocation within a network slice based on the second parameter information and the spectrum resource allocated to the network slice using the second reinforce network;
[0025] updating parameters of the first reinforce network and the second reinforce network using a Monte Carlo algorithm according to the executed actions of the first action network and the second action network.
[0026] In a second aspect, the present application provides a network slice resource allocation apparatus, comprising:
[0027] an acquisition module configured to acquire first state information of each network slice, second state information of each terminal corresponding to each network slice, and third state information of spectrum resources of an access network device;
[0028] an allocation module configured to allocate spectrum resources of each network slice according to the first state information and the third state information, and allocate spectrum resources of each terminal corresponding to each network slice according to the spectrum resources allocated to each network slice and the second state information.
[0029] In a third aspect, the present application provides a computer readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method of the first aspect.
[0030] In a fourth aspect, the present application provides an electronic device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor implements the method of the first aspect when executing the computer program.
[0031] In a fifth aspect, the present application provides a computer program product having a computer program stored thereon, wherein the computer program product, when executed by a processor, implements the method of the first aspect.
[0032] By the above technical solution, the application provides a network slice resource allocation method, device, storage medium and electronic equipment, wherein the method comprises: first acquiring first state information of each network slice, second state information of each terminal corresponding to each network slice, and third state information of spectrum resources of an access network device; then allocating spectrum resources of each network slice according to the first state information and the third state information, and allocating spectrum resources of each terminal corresponding to each network slice according to the spectrum resources allocated to each network slice and the second state information. By applying the technical solution of the application, effective allocation of network slice resources can be realized, the problem that the traditional method is difficult to solve under an increasingly complex network can be effectively solved, the complex environment and variable demand can be effectively adapted, the limited spectrum resources can be reasonably and efficiently allocated according to different demands of upper network slices and lower terminal users, user satisfaction can be realized at the same time, resource reconfiguration overhead between slices is reduced, and resource utilization is improved.
[0033] The above description is only a summary of the technical solution of the application. In order to more clearly understand the technical means of the application, the application can be implemented according to the content of the specification, and in order to make the above purpose and other purposes, characteristics and advantages of the application more obvious and easy to understand, the following specific embodiments of the application are described. BRIEF DESCRIPTION OF DRAWINGS
[0034] The drawings incorporated into the specification and forming part of the specification show embodiments consistent with the application and, together with the specification, serve to explain the principles of the application.
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, for those skilled in the art, other drawings can also be obtained without creative labor.
[0036] Figure 1 A flowchart of a network slice resource allocation method provided by an embodiment of the application is shown;
[0037] Figure 2 A flowchart of another network slice resource allocation method provided by an embodiment of the application is shown;
[0038] Figure 3 A flowchart of an example provided by an embodiment of the application is shown;
[0039] Figure 4 A flowchart of an example provided by an embodiment of the application is shown;
[0040] Figure 5An example flowchart provided by the embodiment of the present application is shown.
[0041] Figure 6 A structural diagram of a network slice resource allocation device provided by the embodiment of the present application is shown.
[0042] Figure 7 An example structural diagram provided by the embodiment of the present application is shown. DETAILED DESCRIPTION
[0043] The embodiments of the present application will be described in more detail below with reference to the accompanying drawings. It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.
[0044] To solve the technical problem of how to realize effective allocation of network slice resources. The embodiment provides a network slice resource allocation method, as shown in the figure, the method comprises the following steps: Figure 1
[0045] Step 101, acquiring first state information of each network slice, second state information of each terminal corresponding to each network slice, and third state information of spectrum resources of an access network device.
[0046] A network slice is a logical network that is constructed by building multiple virtualized and isolated logical networks on a physical network to meet the different service quality requirements of specific business types or industry users. Each network slice can be customized and configured according to different business needs and service quality requirements. In the embodiment, the first state information can be used to determine the state of the network slice, such as performance indicators, availability, and load conditions. Each network slice can serve multiple terminals, and the second state information can be used to determine the state of the terminal corresponding to the network slice, such as the connection state, signal strength, and data usage. The access network device is responsible for managing and allocating wireless spectrum resources to support communication between different network slices and terminals, and the third state information can be used to determine the state of the spectrum resources of the access network device, such as available spectrum resources and spectrum utilization.
[0047] The execution subject of the embodiment can be a device or equipment for allocating network slice resources, which can be configured on the side of the access network device, or the access network device is the execution subject of the embodiment, etc.
[0048] In some embodiments, the access network device is, for example, a node or device that accesses a terminal to a wireless network, and the access network device can include at least one of an evolved node B (eNB) in a 5G communication system, a next generation eNB (ng-eNB), a next generation node B (gNB), a node B (NB), a home node B (HNB), a home evolved node B (HeNB), a wireless backhaul device, a radio network controller (RNC), a base station controller (BSC), a base transceiver station (BTS), a base band unit (BBU), a mobile switching center, a base station in a 6G communication system, an open base station (Open RAN), a cloud base station (Cloud RAN), a base station in other communication systems, an access node in a Wi-Fi system, but is not limited thereto.
[0049] It should be noted that the terms "access network device (AN device)", "radio access network device (RAN device)", "base station (BS)", "radio base station", "fixed station", "node", "access point", "transmission point (TP)", "reception point (RP)", "transmission / reception point (TRP)", "panel", "antenna panel", "antenna array", "cell", "macro cell", "small cell", "femto cell", "pico cell", "sector", "cell group", "carrier", "component carrier", "bandwidth part (BWP)", and the like can be replaced with each other.
[0050] Step 102: allocating spectrum resources of each network slice according to the first status information and the third status information, and allocating spectrum resources of each terminal corresponding to each network slice according to the spectrum resources allocated to each network slice and the second status information.
[0051] By applying the technical solutions of the present application, effective allocation of network slice resources can be achieved.
[0052] In some example scenarios, the rapid development of 5G communication provides strong technical support for various emerging applications such as robots, Internet of Things, industrial automation, vehicle-road cooperation, and smart cities. The integration of 5G technology and robot technology is considered one of the important directions of the next generation of intelligent manufacturing and service robots, which can significantly improve the perception, decision-making, and execution capabilities of robots. 5G technology has the characteristics of high speed, low latency, high reliability, high spectrum efficiency, and large-scale connection. As one of the core technologies of 5G, network slicing technology divides physical network resources into multiple logically independent networks, and each virtual network can be customized for different functional scenarios. This flexible resource management method perfectly adapts to the characteristics of "5G + robot" applications and can significantly improve production efficiency and safety. However, how to effectively allocate slice resources on the access network side still has great challenges.
[0053] In the "5G + robot" application, different robot application scenarios have different requirements for bandwidth allocation, latency guarantee, and reliability requirements. For example, control robots in industrial automation require extremely low latency and high reliability support, i.e., ultra-reliable low-latency communication (uRLLC), scanning robots may focus more on bandwidth and coverage, i.e., enhanced mobile broadband (eMBB), and robotic factories have numerous sensors that need to implement massive machine communication (mMTC). How to dynamically adjust the allocation of network resources according to the needs of different application scenarios is the key to ensuring the efficient operation of 5G + robots.
[0054] Based on the above example scenarios, research on slice resource allocation on the access network side has focused on traditional optimization algorithms and model-based allocation strategies. However, these methods have problems such as dimension disaster, high computational complexity, and poor adaptability to dynamic environments when dealing with large-scale, multi-dimensional resource allocation problems in "5G + robot" application scenarios. Therefore, the present embodiment proposes a slice resource allocation method on the access network side based on hierarchical reinforcement learning. By decomposing the resource allocation on the access network side into inter-slice and intra-slice resource allocation problems, using a hierarchical reinforcement learning framework, and dynamically optimizing resource allocation according to application communication service requirements, the utilization rate of inter-slice and intra-slice resources in the wireless access network is improved, and efficient adaptation to complex environments and changing requirements is achieved. The method proposed in the present embodiment, as shown in Figure 2 may include:
[0055] Step 201, obtaining first state information of each network slice, second state information of each terminal corresponding to each network slice, and third state information of spectrum resources of an access network device.
[0056] In step 202, the frequency spectrum resources of each network slice and the frequency spectrum resources of each terminal corresponding to each network slice are determined according to the first state information, the second state information and the third state information by using the reinforcement learning model, and the frequency spectrum resources are allocated according to the determination result.
[0057] Reinforcement learning is one of the important machine learning methods, in which an agent learns an optimal policy by constantly interacting with the environment to maximize cumulative rewards. As the network size continues to expand, the state space dimension of the problem increases, which causes a large amount of storage and computing resource consumption. Hierarchical reinforcement learning divides the task into multiple levels of subtasks, making the solution of complex problems more efficient and flexible. In hierarchical reinforcement learning, the high-level strategy is responsible for planning long-term goals, and the low-level strategy is responsible for specific operation execution. The collaborative work between the two can significantly improve the learning efficiency and effect. Among them, hierarchical reinforcement learning is a reinforcement learning method that divides complex tasks into more manageable subtasks and simplifies the optimization process through a hierarchical structure. The embodiment can use option-based reinforcement learning.
[0058] As an optional mode, the reinforcement learning model can be trained using an actor-critic algorithm. In some examples, first parameter information can be determined first, which can include at least parameters of a first action network, parameters of a second action network, and parameters of a critic network, etc., wherein the first action network is used to perform the action of inter-network slice frequency spectrum resource allocation, the second action network is used to perform the action of intra-network slice frequency spectrum resource allocation, and the critic network is used to guide the action execution of the first action network and the second action network; then the reinforcement learning model is obtained by iterative training using the actor-critic algorithm based on the first parameter information.
[0059] In some examples, the process of training the reinforcement learning model in each iteration based on the first parameter information using the actor-critic algorithm includes: performing spectrum resource allocation between network slices using a first action network based on the first parameter information, and calculating a corresponding first reward value through a first reward function, the first reward function being determined based on the satisfaction of the end users, the network slice reconfiguration overhead, the spectrum resource utilization, and the decision cost of option switching; performing spectrum resource allocation within a network slice using a second action network based on the first parameter information and the spectrum resource to which the network slice is allocated, and calculating a corresponding second reward value through a second reward function, the second reward function being determined based on the satisfaction of the end users; then calculating the action value and bias of the second action network using a critic network, and updating the parameters of the second action network according to the action value and the bias; updating the experience pool based on the actions of the first action network and the second action network, and combining the first reward value, the second reward value, the state of the network slices after spectrum resource allocation, and the state of the corresponding terminals; when the data in the experience pool reaches a preset storage threshold, updating the parameters of the critic network and the first action network using the data in the experience pool, and emptying the experience pool after the updating is completed.
[0060] The above process is exemplarily illustrated as follows:
[0061] As Figure 3 Fig. 1 shows a flowchart of a slice resource allocation method on the access network side based on an asynchronous advantage option-critic (A2OC) framework, and the main steps include:
[0062] Step S1: initializing variables such as network parameters, an experience pool, an exploration probability, a switching cost, a global time, and a process time; initializing the state of each slice in the upper layer, the state of the terminal in the lower layer, and the state of the system spectrum resource; and obtaining a current environment state;
[0063] Step S2: according to the initial upper slice state in step S1 or the new upper slice state in step S4, an upper action (Actor) network (i.e., a first action network) selects an inter-slice resource allocation action according to an option policy and performs the action, and calculates a reward value (i.e., a first reward value) through an upper reward function (i.e., a first reward function);
[0064] Each option represents a resource allocation scheme between slices.
[0065] Step S3: Within the process time, each slice updates the lower terminal state according to the initial lower terminal state obtained in step S1 or the new lower terminal state in step S3, under the constraints of the inter-slice resource allocation in step S2, the lower action network (i.e., the second action network) selects and executes the intra-slice resource allocation action according to the policy, updates the lower terminal state, and calculates the reward value (i.e., the second reward value) through the lower reward function (i.e., the second reward function);
[0066] Step S4: The critic network updates the upper state by combining the upper slice state at the previous moment and the updated terminal state within the slice at the current moment, calculates the action value and bias, and updates the lower action network parameters;
[0067] Step S5: Determine whether the option enters the termination state, if yes, update the option according to the exploration probability and the updated terminal state within the lower slice, and increase the decision cost; if not, jump to step S3, and the decision cost is 0;
[0068] Step S6: Store the current moment state of the upper inter-slice environment, action, reward, and next moment state as a set of experience data in the experience pool, when the data in the experience pool reaches the storage threshold, train the algorithm network and update the network parameters, empty the experience pool after completing the policy update;
[0069] Step S7: Determine whether the current round number reaches the preset maximum value, if yes, jump to step S8, if not, jump to step S2;
[0070] Step S8: Determine whether the slice and terminal performance within the global time is within the preset index, if yes, complete the training process of the method in the embodiment, and apply the model scheme in the industrial scene to allocate the access network side slice resource, if not, jump to step S1.
[0071] Further, in order to illustrate the specific implementation process of each step, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the drawings in the embodiments of the present application. Figure 4 A complete flowchart of an access network side slice resource allocation method based on A2OC framework hierarchical reinforcement learning.
[0072] In the industrial scene, there are video transmission, control instruction transmission, data acquisition and other types of business, and these businesses have different communication needs. For example, control instruction transmission has strict requirements on delay and reliability to ensure the safety and order of industrial production; video transmission business has high bandwidth demand to ensure high-definition and non-lagging transmission, generally requiring a transmission bandwidth of ≥2 Mbps; and the acquisition type business focuses on connection density and has lower requirements on bandwidth, delay and reliability. In view of the problem of diversification of business types and differentiation of communication needs in the industrial scene, network slicing is deployed to realize the isolation of each type of business, and the resource optimization configuration of the uplink in the access network side network slicing is focused on. It is assumed that N network slices are deployed on the access network side, and N = {1, 2, …, n} is used to represent the set of network slices. These network slices serve K robots, which are randomly distributed within the coverage of the factory base station. According to the characteristics of the business, the robots are connected to the corresponding slice, and each user will connect to the corresponding slice and only connect to one slice. It is assumed that in the factory environment, the number of available spectrum resources is M resource blocks, and the subchannel bandwidth size is B. In combination with the actual system, the base station needs to allocate the limited resource blocks (RB) to different business types of network slices with a certain resource granularity, and the network slice resource scheduler allocates resources within the slice to terminal robots.
[0073] Step S1: initialize network parameters, experience pool, exploration probability, switching cost, global time and process time, etc.; initialize the state of each slice in the upper layer, the state of the terminal in the lower layer slice and the system resource state; obtain the current environment state;
[0074] Specifically, the random weight is used to initialize the parameters of the upper action network, the lower action network and the comment network, the upper experience replay pool D, the exploration probability ε, the switching cost η, the global time variable T and the process time variable t.
[0075] Further, the contents that need to be initialized include: the number of slices N, the number of terminal robots K, the total amount of system spectrum resource blocks M, the subchannel bandwidth B and the user sending buffer queue capacity;
[0076] Further, the current environment state that needs to be obtained includes the state between the upper layer slices and the state of the terminal within the lower layer slice, which is specifically: the amount of data to be sent in the buffer of slice n at time t b e , the number of resources allocated to slice n at time t θ t , the available resources of slice n at time t and the number of resources required by user k That is:
[0077] S u ={b t ,θ t}
[0078]
[0079] where S u and S l represent the upper slice inter-state space and the lower terminal state space, respectively.
[0080] Step S2: According to the obtained initial upper slice state in step S1 or the new upper slice state in step S4, the upper action network selects a slice inter-resource allocation action according to the policy in the option and executes the action, and calculates the reward value through the upper reward function;
[0081] Specifically, the action dimension output by the upper action network is determined by the number of slices N, and the probability distribution of the action value output by the upper action network is quantized into the number of RBs. Specifically, the upper action space A u is the number of resources allocated to the slice by the base station That is:
[0082]
[0083] It should be noted that in one system, without considering resource reservation, the number of RBs allocated to all network slices needs to satisfy:
[0084]
[0085] The number of RBs allocated to all users cannot exceed the total amount of system spectrum resources, that is:
[0086]
[0087] Further, the reward value is calculated according to the upper reward function. In the upper action policy learning process, the state transition process of an individual after entering an option is related to the historical process experienced, so the establishment function of the upper learning process should be the cumulative return of the upper selected action within the duration of the option. Define the upper reward function R u as:
[0088] R u = H t
[0089] where H t is specifically defined as follows:
[0090] Assume that the user k requires a service rate of the service rate of the terminal at the current time is greater than or equal to the required service rate Therefore, when the service rate of the terminal is less than the demand service rate, the user satisfaction of the terminal increases with the increase of the service rate, but when the service rate of the terminal is greater than the demand service rate, the increase of the service rate of the terminal does not bring the improvement of the user satisfaction, but causes the waste of resources, and thus the overall satisfaction of the system decreases. Therefore, the user satisfaction is defined as:
[0091]
[0092] In industrial production, some services are more important and need to be satisfied first, and thus alpha is defined as the satisfaction coefficient of the slice n, reflecting the proportional relationship between the satisfaction and the service rate in different slices. Beta is a positive value, used to represent the penalty size when the service rate of the terminal is greater than the demand service rate.
[0093] In addition, considering the overhead of inter-slice reconfiguration, the overhead of one inter-slice reconfiguration is defined as Where gamma is the inter-slice reconfiguration overhead coefficient; I(x) is an indicator function, which is 0 when x = 0, and 1 when x ≠ 0, used to represent whether the slice is reconfigured. represents the state of the slice at the current time, represents the state of the slice before adjustment.
[0094]
[0095] Finally, considering the resource utilization, the resources in the actual system can only be allocated in a certain granularity, and there will be resource waste. In order to maximize the resource utilization, the resource utilization of the slice μ is defined as utll As follows, where delta is the resource utilization coefficient, m n represents the actual number of resources allocated to the slice n.
[0096]
[0097] Finally, considering the user satisfaction, the overhead of slice reconfiguration, the resource utilization, and the decision cost of the user option switching under the A2OC framework, the upper inter-slice reward function is defined as:
[0098]
[0099] Where d is the time scale of the option after the individual enters the option, and c is the decision cost set in the A2OC framework to avoid too frequent switching of the option, which is specifically updated in step S5.
[0100] Step S3: Within the process time, each slice updates the lower terminal state according to the initial lower terminal state in step S1 or the new lower terminal state in step S3, selects the intra-slice resource allocation action according to the policy under the constraint of the inter-slice resource allocation in step S2, and executes the action, updates the lower terminal state, and calculates the reward value through the lower reward function;
[0101] Specifically, in the learning process of the lower layer, the slice controller allocates the RBs allocated to the slice to the intra-slice terminal, so that the lower action space A l is designed as:
[0102]
[0103] wherein represents the number of RBs allocated to the user k.
[0104] It should be noted that, considering the hierarchical problem of inter-slice and intra-slice resource allocation, at time t, the number of resource blocks allocated by each network slice n to the terminal robot k cannot exceed the number of RBs allocated to the slice at this time, and this constraint condition can be expressed as:
[0105]
[0106] After executing the action, the system environment changes, the terminal state is updated, and a new lower terminal state is obtained; and the reward value is calculated through the lower reward function. Since the learning process of the lower action selection policy is independent between slices, the reward obtained by the lower layer should be the sum of the intra-slice terminal satisfaction, and the lower reward function R l is defined as:
[0107]
[0108] Step S4: Update the upper state by combining the upper slice state at the previous time and the intra-slice terminal state at the current time, calculate the state value and bias of the critic network, and update the lower action network parameters;
[0109] Further, the critic network calculates the option state value function V(s; ω) and the action value function Q(s, o; ω), and calculates the TD error ξ t to reflect the difference between the current estimate and the actual return, wherein ξ t is:
[0110]
[0111] Then, the critic network parameters ω are updated using the gradient descent method, and the critic network is continuously updated and optimized, and the update step is:
[0112]
[0113] where α is the learning rate, is the gradient of critic network parameters ω.
[0114] Further, depending on the value evaluation of the critic network, the policy gradient is calculated, and the lower action network parameters are updated to learn a better policy, and to maximize the cumulative reward.
[0115] Step S5: Determine whether the option enters a termination state, if yes, update the option according to the exploration probability and the lower slice terminal state, and the decision cost increases; if not, jump to step S3, and the decision cost is 0;
[0116] Specifically, the option framework defines the time-expanded action. Generally, a Markov option can be expressed as a triple <I ω ,π ω ,β ω >. Wherein, represents the initial state set of the option; π ω represents the option internal strategy, which determines the action inside the option; β ω :S→[0,1] represents the termination condition, that is, there is a probability of β ω to terminate and exit the option when reaching state s. Assuming that at time t, the individual enters the option and the option will last for d time scales, then at t+d, the option is exited. For each time τ in the duration, the state change and action selection strategy of the option depend on the entire learning process s h,τ ={s t ,a t ,r t ;s t+1 ,a t+1 ,r t+1 ;...;s τ ,a τ ,r τ}. It is proved in the initial option file that the existence of the option makes the Markov decision process (MDP) become a semi-Markov process, which has the optimal state value function V(s; ω) and the action value function Q(s, o; ω) corresponding to the option.
[0117] When using the hierarchical reinforcement learning method of the option-critic framework for end-to-end training, it is possible that, as the training progresses, the termination gradient tends to make the option terminate as soon as possible due to the optimization process overemphasizing the immediate reward, and the scope and duration of a single option become smaller and smaller, i.e., t→τ, eventually degenerating into a basic action. In each state, the agent can choose to terminate the current option and start another option, resulting in that, in actual operation, these options no longer exhibit the advantages of the high-level decision or behavior pattern designed for them. According to the idea of bounded rationality, A2OC proposes the idea of deliberation cost to help the agent learn and plan more quickly, reduce the number of times of switching other options as much as possible, explore more stable and efficient strategy execution, and reduce the overhead of inter-slice resource scheduling. Specifically, it is expressed as:
[0118]
[0119] c represents the decision cost, and η represents the cost increment added each time the option is switched.
[0120] Step S6: store the current time state, action, reward and next time state of the upper inter-slice environment in the experience pool as a set of experience data, train the algorithm network and update the network parameters when the data in the experience pool reaches the storage threshold, and empty the experience pool after completing the policy update;
[0121] Specifically, after the upper inter-slice state is updated, the four-tuple is stored in the experience replay pool D. If the data in the experience replay pool D is full, the next state is input to the critic network to obtain the state value The discounted reward and advantage function are calculated.
[0122] Further, at each update time, the stored state is input to the critic network to obtain the state value, the loss function is calculated for gradient update; the stored state and action are input to the upper action network to obtain the new and old policy probability ratio, and the loss function is calculated for gradient update.
[0123] Step S7: determine whether the current round number reaches the preset maximum value, if yes, jump to step S8, if no, jump to step S2;
[0124] Step S8: determine whether the slice and terminal performance in the global time reach the preset target, if yes, complete the training process of the method proposed in the patent, and apply the model scheme to the access network side slice resource allocation in the industrial scene, if no, continue to train the algorithm.
[0125] As another alternative, the reinforcement learning model can also be trained using a reinforce algorithm. As an example, second parameter information can be determined first, the second parameter information including at least parameters of a first reinforce network for performing actions of inter-network slice spectrum resource allocation and parameters of a second reinforce network for performing actions of intra-network slice spectrum resource allocation; and then the reinforcement learning model is trained using the reinforce algorithm based on the second parameter information.
[0126] In some examples, the process of training the reinforcement learning model using the reinforce algorithm based on the second parameter information comprises: performing inter-network slice spectrum resource allocation using the first reinforce network based on the second parameter information; performing intra-network slice spectrum resource allocation using the second reinforce network based on the second parameter information and the spectrum resources allocated to the network slice; and updating the parameters of the first reinforce network and the second reinforce network using the Monte Carlo algorithm according to the performed actions of the first action network and the second action network.
[0127] The implementation process of the access network side slice resource allocation method based on the reinforce framework is described below as an example:
[0128] In the access network side slice resource allocation method based on the reinforce framework, based on the A2OC framework, the initialization variables such as exploration probability, switching cost, time variable are kept unchanged, the same upper slice state space and lower slice state space are established, the same user satisfaction and upper and lower space slice reward functions are defined, and the same option switching strategy is used, the A2OC network architecture is replaced by the reinforce network architecture, the data pool process is removed, only the data generated by the current round is used for iteration each time, the Monte Carlo method is used to replace the evaluation network to evaluate the reinforcement learning effect of the reinforce neural network itself, and the parameters are updated. After that, the same logic as the A2OC framework method is used to complete the parameter iteration of the multi-layer network model, whether to continue iteration or end iteration is selected according to whether the preset iteration number is reached, and the upper and lower layer resource allocation strategies are output.
[0129] Step S1: Compared with the A2OC framework scheme, the initialization process of the experience pool is removed, and the upper and lower neural networks of the reinforcement learning network are replaced by the reinforce framework. For example, the upper and lower reinforce network parameters, exploration probability, switching cost, global time and process time, etc. are initialized; the upper slice state, the lower slice terminal state and the system resource state are initialized; the current environment state is obtained;
[0130] Step S2: According to the initial upper slice state in step S1 or the new upper slice state in step S4, the upper reinforce network (i.e., the first reinforce network) selects an inter-slice resource allocation action according to the policy in the option and performs the action, calculates the reward value;
[0131] wherein each option represents an inter-slice resource allocation scheme.
[0132] Step S3: Within the process time, each slice selects an intra-slice resource allocation action according to the policy in the option based on the initial lower terminal state in step S1 or the new lower terminal state in step S3 under the constraint of the inter-slice resource allocation in step S2, and performs the action, updates the lower terminal state, calculates the reward value, and updates the parameters of the lower reinforce network (i.e., the second reinforce network) using the gradient ascent method for continuous updating and optimization, further calculates the policy gradient, updates the parameters of the lower action network, learns a better policy, and maximizes the cumulative reward.
[0133] Step S4: The derivative of the expected function of the option state value function and the action value function is calculated using the Monte Carlo method, and then the parameters are updated using the gradient ascent method for continuous updating and optimization, further, the policy gradient is calculated, the parameters of the lower action network are updated, a better policy is learned, and the cumulative reward is maximized.
[0134] Step S5: Determine whether the option enters the termination state, if yes, update the option according to the exploration probability and the lower intra-slice terminal state, and the decision cost increases; if not, jump to step S3, and the decision cost is 0.
[0135] Step S6: The current time state, action, reward, and next time state of the upper inter-slice environment are batch generated using the Monte Carlo method. The overall state value is estimated based on the samples, and the algorithm network is trained and the network parameters are updated based on this, and the policy is updated.
[0136] Step S7: Determine whether the current round number reaches the preset maximum value, if yes, jump to step S8, if not, jump to step S2.
[0137] Step S8: Determine whether the slice and terminal performance within the global time is within the preset index, if yes, complete the training process of the method in the embodiment, and apply the model scheme in the industrial scene for access network side slice resource allocation, if not, jump to step S1.
[0138] Further, in order to illustrate the specific implementation process of each step, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the drawings in the embodiments of the present application. Figure 5A complete flowchart of an access network side slice resource allocation method based on A2OC framework hierarchical reinforcement learning.
[0139] In the access network side slice resource allocation method of the reinforce framework hierarchical reinforcement learning, based on the A2OC access network side slice resource allocation method, the initialization variables such as exploration probability, switching cost, and time variable are kept unchanged, the same upper layer slice state space and lower layer slice state space are established, the same user satisfaction and upper and lower layer space slice reward function are defined, and the same option switching strategy is used. On the basis of replacing the A2OC network architecture with the reinforce network architecture, removing the data pool establishment process, using only the data generated by the current round of operation for iteration each time, using the Monte Carlo method to replace the critic network with the reinforce neural network itself to evaluate the reinforcement learning effect and update its own parameters. After that, the same logic as in the embodiment is used to complete the parameter iteration of the multi-layer network model, and whether to continue iteration or end iteration is selected according to whether the preset iteration number is reached, and the upper and lower layer resource allocation strategies are output. Figure 4
[0140] Specifically, in steps S1-S8 of the embodiment, steps S1, S4 and S6 are changed, and the remaining steps are consistent in principle and parameter definition. Figure 4 In steps S1-S8 of the embodiment, steps S1, S4 and S6 are changed, and the remaining steps are consistent in principle and parameter definition.
[0141] Step S1: Compared with the A2OC framework scheme, the initialization process of the experience pool is removed, and the upper and lower neural networks of the reinforcement learning network are replaced with the reinforce framework.
[0142] Step S4: The derivative of the expected J function of the option state value function V(s) and the action value function Q is calculated using the Monte Carlo method.
[0143] J(θ)=E S [V π (S)]
[0144]
[0145] Then the parameters are updated using the gradient ascent method for continuous updating and optimization, and the update steps are:
[0146]
[0147] Further, the policy gradient is calculated, the lower layer action network parameters are updated, a better policy is learned, and the cumulative reward maximization is realized.
[0148] Step S6: using Monte Carlo method to batch generate the current time state, action, reward and next time state of the upper slice inter-environment. The overall state value is estimated by using the sample, and the algorithm network is trained and the network parameters are updated based on this, and the policy update is completed.
[0149] To meet the diversified demand of different service types of robots for network resources, an access network side slice resource allocation method based on hierarchical reinforcement learning is proposed. According to the characteristics and demands of the business in the industrial scene, the method proposes to deploy network slices on the access network side, and realizes data isolation and differentiated service customization of different types of businesses through flexible control of spectrum resources. With the introduction of network slices, the original wireless network side resource allocation from the base station directly to the user is decomposed into inter-slice and intra-slice double-layer resource allocation, forming a clear hierarchy. Based on the hierarchical reinforcement learning of asynchronous advantage option-critic (A2OC) framework and reinforce framework, the method divides the resource allocation process into two levels. In the upper layer, the base station allocates the spectrum resources in the cell to different network slices, and each option represents the inter-slice resource allocation scheme; in the lower layer, the slice allocates resources to the data terminals in the slice according to the allocated resources data through the internal decision mechanism. Through this hierarchical reinforcement learning-based dynamic optimization and allocation method of resources, the user satisfaction can be maximized while reducing the system overhead, improving the utilization rate of inter-slice and intra-slice resources of the wireless access network, and effectively adapting to complex environments and changing demands.
[0150] Further, as a specific implementation of the method shown in Figure 1 and Figure 2 , the embodiment provides a network slice resource allocation device, as shown in Figure 6 , the device comprises an acquisition unit 31 and an allocation unit 32.
[0151] The acquisition unit 31 is configured to acquire first state information of each network slice, second state information of each terminal corresponding to each network slice, and third state information of spectrum resources of an access network device;
[0152] The allocation unit 32 is configured to allocate spectrum resources of each network slice according to the first state information and the third state information, and allocate spectrum resources of each terminal corresponding to each network slice according to the spectrum resources allocated to each network slice and the second state information.
[0153] In some examples of the present embodiment, the allocation unit 32 is specifically configured to determine the spectrum resources of each network slice and the spectrum resources of each terminal corresponding to each network slice according to the first state information, the second state information and the third state information, and allocate the spectrum resources according to the determination results.
[0154] In some examples of the present embodiment, the allocation apparatus further comprises a training unit.
[0155] The training unit is configured to determine first parameter information, the first parameter information comprising at least parameters of a first action network, parameters of a second action network and parameters of a critic network, wherein the first action network is used to perform actions of inter-network slice spectrum resource allocation, the second action network is used to perform actions of intra-network slice spectrum resource allocation, and the critic network is used to guide the action performance of the first action network and the second action network; and perform iterative training using an actor-critic algorithm based on the first parameter information to obtain the reinforcement learning model.
[0156] In some examples of the present embodiment, the training unit is specifically configured to perform inter-network slice spectrum resource allocation using the first action network based on the first parameter information, and calculate a corresponding first reward value by a first reward function, the first reward function being determined based on the satisfaction of terminal users, network slice reconfiguration overhead, spectrum resource utilization and decision cost of option switching; perform intra-network slice spectrum resource allocation using the second action network based on the first parameter information and the spectrum resources to which the network slices are allocated, and calculate a corresponding second reward value by a second reward function, the second reward function being determined based on the satisfaction of terminal users; calculate the action value and bias of the second action network using the critic network, and update the parameters of the second action network according to the action value and bias; update the experience pool based on the actions of the first action network and the second action network, and in combination with the first reward value, the second reward value, the state of inter-network slices after spectrum resource allocation and the state of corresponding terminals; when the data in the experience pool reaches a preset storage threshold, update the parameters of the critic network and the first action network using the data in the experience pool, and empty the experience pool after the update is completed.
[0157] In some examples of the present embodiment, the training unit is further configured to determine second parameter information, the second parameter information comprising at least parameters of a first reinforce network and parameters of a second reinforce network, wherein the first reinforce network is used to perform the action of inter-network slice spectrum resource allocation, and the second reinforce network is used to perform the action of intra-network slice spectrum resource allocation; and perform iterative training using a reinforce algorithm based on the second parameter information to obtain the reinforcement learning model.
[0158] In some examples of the present embodiment, the training unit is further configured to perform inter-network slice spectrum resource allocation using the first reinforce network based on the second parameter information; perform intra-network slice spectrum resource allocation using the second reinforce network based on the second parameter information and the spectrum resource to which the network slice is allocated; and update the parameters of the first reinforce network and the second reinforce network using a Monte Carlo algorithm according to the execution action of the first action network and the second action network.
[0159] It should be noted that other corresponding descriptions of the functions of the various functional units involved in the network slice resource allocation apparatus provided by the present embodiment can be referred to the corresponding descriptions in Figure 1 and Figure 2 , which will not be repeated here.
[0160] Based on the above-mentioned method and apparatus of the embodiments, as an example, as shown in Figure 7 , the present embodiment provides a network side slice resource allocation apparatus based on hierarchical reinforcement learning, comprising:
[0161] a perception module configured to collect terminal (such as robots, etc.) network state information and perform service modeling;
[0162] an allocation and scheduling module: a processor executing a hierarchical reinforcement learning algorithm, and a memory in communication with the processor to store a large number of weights, activations, etc. generated by the hierarchical reinforcement learning; to ensure that each network slice meets the specific quality of service (QoS) requirements of the terminal, and to optimize the QoS of the slice through the reinforcement learning algorithm;
[0163] an execution module: executing operation commands through network devices (base stations, routers, switches) to implement the application of the slice resource allocation strategy.
[0164] Based on the above-mentioned method as shown in Figure 1 and Figure 2 , accordingly, the present embodiment also provides a computer readable storage medium having a computer program stored thereon, which is executed by a processor to implement the above-mentioned method as shown inFigure 1 and Figure 2 the method shown in
[0165] Based on the above method as Figure 1 and Figure 2 shown, accordingly, the embodiment also provides a computer program product, which stores a computer program, and the computer program product is executed by a processor to implement the above method as Figure 1 and Figure 2 shown.
[0166] Based on such understanding, the technical scheme of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a U disk, a mobile hard disk, etc.), and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the method of each implementation scenario of the present application.
[0167] Based on the above method as Figure 1 and Figure 2 shown, and Figure 6 the virtual device embodiment shown, in order to achieve the above purpose, the embodiment of the present application also provides an electronic device, such as an access network device, or a device configured on the side of a receiving network device, etc., which includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the above method as Figure 1 and Figure 2 shown.
[0168] Optionally, the above-mentioned entity device can also include a user interface, a network interface, a camera, a radio frequency (Radio Frequency, RF) circuit, a sensor, an audio circuit, a WI-FI module, etc. The user interface can include a display screen (Display), an input unit such as a keyboard (Keyboard), etc. The optional user interface can also include a USB interface, a card reader interface, etc. The network interface can optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), etc.
[0169] Those skilled in the art can understand that the above-mentioned entity device structure provided by the embodiment does not constitute a limitation on the entity device, and can include more or fewer components, or combine certain components, or different component arrangements.
[0170] The storage medium can also include an operating system, a network communication module. The operating system is a program that manages the hardware and software resources of the above-mentioned entity device, supports the running of information processing programs and other software and / or programs. The network communication module is used to realize the communication between the internal components of the storage medium, and the communication with other hardware and software in the information processing entity device.
[0171] Through the above description of the implementation methods, those skilled in the art can clearly understand that this application can be implemented using software plus necessary general-purpose hardware platforms, or it can be implemented through hardware. In industrial scenarios, intelligent manufacturing is constantly developing. By deploying network slices on the access network side, multiple logically independent networks are built on the same physical infrastructure to achieve data isolation for different types of services and customized differentiated services. This embodiment proposes a network slice resource allocation method based on hierarchical reinforcement learning on the access network side. This method allocates limited spectrum resources reasonably and efficiently, reducing the overhead of resource reconfiguration between slices while achieving user satisfaction and improving resource utilization. Specifically, the A2OC hierarchical reinforcement learning framework prevents agents from frequently switching options, making the algorithm more stable and efficient; the reinforcement learning framework simplifies the neural network structure and speeds up training. Furthermore, the overhead of reconfiguration between slices is much higher than the overhead of user reconfiguration within a slice, and the actual technical switching frequency may not reach this level. By using the A2OC or reinforcement algorithm, the overhead cost of reconfiguration between slices can be effectively reduced. By dynamically allocating resources between and within slices using hierarchical reinforcement learning algorithms, the dynamic performance of the system can be improved, and scheduling is more efficient than traditional mathematical methods.
[0172] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0173] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A method for allocating network slice resources, characterized in that, The method comprises the following steps: obtaining first state information of each network slice, second state information of each terminal corresponding to each network slice, and third state information of spectrum resources of an access network device; allocating spectrum resources of each network slice according to the first state information and the third state information, and allocating spectrum resources of each terminal corresponding to each network slice according to the spectrum resources allocated to each network slice and the second state information.
2. The method of claim 1, wherein, The method for allocating spectrum resources of each network slice according to the first state information and the third state information, and allocating spectrum resources of each terminal corresponding to each network slice according to the spectrum resources allocated to each network slice and the second state information, comprises: determining spectrum resources of each network slice and spectrum resources of each terminal corresponding to each network slice by using a reinforcement learning model according to the first state information, the second state information and the third state information, and performing allocation of the spectrum resources according to the determination result.
3. The method of claim 2, wherein, The training process of the reinforcement learning model comprises: determining first parameter information, wherein the first parameter information at least comprises parameters of a first action network, parameters of a second action network and parameters of a comment network, wherein the first action network is used to perform an action of inter-network slice spectrum resource allocation, the second action network is used to perform an action of intra-network slice spectrum resource allocation, and the comment network is used to guide the action execution of the first action network and the second action network; performing iterative training of the reinforcement learning model based on the first parameter information by using an actor-critic algorithm.
4. The method of claim 3, wherein, The process of training the reinforcement learning model by using the actor-critic algorithm based on the first parameter information in each iteration comprises: performing inter-network slice spectrum resource allocation by using the first action network based on the first parameter information, and calculating a corresponding first reward value by using a first reward function, wherein the first reward function is determined based on satisfaction of a terminal user, network slice reconfiguration overhead, spectrum resource utilization and decision cost of option switching; performing intra-network slice spectrum resource allocation by using the second action network based on the first parameter information and the spectrum resources allocated to the network slice, and calculating a corresponding second reward value by using a second reward function, wherein the second reward function is determined based on satisfaction of a terminal user; calculating an action value and a bias of the second action network by using the comment network, and updating the parameters of the second action network according to the action value and the bias; updating an experience pool based on the actions of the first action network and the second action network, and combining the first reward value, the second reward value, a state after spectrum resource allocation between network slices and a state of corresponding terminals; when the data in the experience pool reaches a preset storage threshold, updating the parameters of the comment network and the first action network by using the data in the experience pool, and emptying the experience pool after the updating is completed.
5. The method of claim 2, wherein, The training process of the reinforcement learning model comprises: determining second parameter information, the second parameter information comprising at least parameters of a first reinforce network and parameters of a second reinforce network, wherein the first reinforce network is used to perform an action of inter-network slice spectrum resource allocation, and the second reinforce network is used to perform an action of intra-network slice spectrum resource allocation; iteratively training the reinforce learning model based on the second parameter information using a reinforce algorithm.
6. The method of claim 5, wherein, The process of training the reinforce learning model based on the second parameter information using the reinforce algorithm in each iteration comprises: performing inter-network slice spectrum resource allocation using the first reinforce network based on the second parameter information; performing intra-network slice spectrum resource allocation using the second reinforce network based on the second parameter information and spectrum resources allocated to the network slice; updating parameters of the first reinforce network and the second reinforce network using a Monte Carlo algorithm according to the performed actions of the first action network and the second action network.
7. A network slice resource allocation device, characterized in that, comprises: an obtaining module configured to obtain first state information of each network slice, second state information of each terminal corresponding to each network slice, and third state information of spectrum resources of an access network device; an allocating module configured to allocate spectrum resources of each network slice according to the first state information and the third state information, and allocate spectrum resources of each terminal corresponding to each network slice according to spectrum resources allocated to each network slice and the second state information.
8. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by a processor to implement the method of any one of claims 1 to 6.
9. An electronic device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, The processor executes the computer program to implement the method of any one of claims 1 to 6.
10. A computer program product having stored thereon a computer program, characterized in that The computer program product is executed by a processor to implement the method of any one of claims 1 to 6.