Continuous network slicing in a 5G cellular communications network via a delayed deep deterministic policy gradient

GB2599196BActive Publication Date: 2025-07-09EBOS TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
GB2021008215
Authority / Receiving Office
GB · GB
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-06-29
Filing Date
2021-06-09
Publication Date
2025-07-09
Estimated Expiration
2041-06-09

Smart Images

  • Figure 00000001_0000
    Figure 00000001_0000
  • Figure 00000001_0001
    Figure 00000001_0001
Patent Text Reader

Abstract

Delayed deep deterministic policy gradient (DDPG) based continuous network slicing defines at least two network slices in a central unit (CU) of a 5G network, and at different time steps identifies a
Need to check novelty before this filing date? Find Prior Art

Description

5 BACKGROUND OF THE INVENTION

[0001] Field of the Invention

[0002] The present invention relates to the field of cellular data communications and more particularly to network slicing in a fifth generation (5G) cellular telecommunications network. 10

[0003] Description of the Related Art

[0004] Cellular data communications refer to the exchange of data traffic over a cellular telecommunications network. Digital cellular data communications require the presence of an underlying physical data communications infrastructure layered upon a cellular network, such as that first evidenced by second generation digital cellular 15 communications, and more recently by the substantially more robust and reliable fourth generation (4G) long term evolution (LIE) cellular data communications network. In 4G LIE, the network architecture supports the connectivity of user equipment (UE) to different base stations (eNBs) clustered in different radio access networks (RANs) with the RANs each coupled to the core network (CN). 20

[0005] The eNBs send and receive radio transmissions to each UE using the analog and digital signal processing functions of the LTE air interface through different multiple input multiple output (MIMO) antenna arrays. Each eNB also controls the low-level 1 9 09 9^ operation of each coupled UE, by sending the UE signaling messages such as handover commands. Finally, each eBN connects with the CN, also known as the "Evolved Packet Core" (EPC), by means of an SI protocol stack interface. Of note, each eBN also may be communicatively coupled to another, nearby eBN by an X2 interface, so as to support 5 signaling and packet forwarding during handover of a communication with UE from eBN to eBN (cell to cell). The EPC, in turn, is a framework for providing converged voice and data on the 4G LTE network. Whereas 2G and third generation (3G) network architectures process and switch voice and data through two separate sub-domains— circuit-switched (CS) for voice and packet-switched (PS) for data—EPC unifies voice and 10 data on an Internet Protocol (IP) service architecture and voice is treated as just another Internet Protocol (IP) application.

[0006] Whereas 4G represented a giant leap in performance over 2D and 3G networks 5G represents an enormous improvement over 4G. Capitalizing on Massive MIMO antenna arrays in each base station, the utilization of millimeter wave radio 15 communications, beamforming for direct wireless communications with individual UE, and a bifurcated centralized unit (CU) and distributed unit (DU) architecture, 5G is able to achieve a data exchange capacity of nearly thirteen terabytes—almost a twenty times improvement over 4G LTE. The CN of the 5G architecture reflects a substantial change over the EPC of 4G. In the CN of 5G, the changes have been reduced, abstractly, into 20 what has been referred to as the "Four Modernizations". The first is "information technology" or "IT", the second is the "Internet", the third is "extremely simplified", and the fourth is "service-based". The most typical change in the network architecture of the 1 9 09 9^ CN is the service-based network architecture of the CN so as to separate the control plane from the user plane. Other technologies support network slicing and edge computing.

[0007] As to the IT modernization, the essential characteristic of the 5G architecture is the notion of network function virtualization (NFV). NFV decouples software from 5 hardware by replacing various network functions such as firewalls, load balancers and routers with virtualized instances running as software. This eliminates the need to invest in many expensive hardware elements and can also accelerate installation times, thereby providing revenue generating services to the customer faster. NFV enables the 5G infrastructure by virtualizing appliances within the 5G network. This includes the 10 network slicing technology that enables multiple virtual networks to run simultaneously. NFV can address other 5G challenges through virtualized computing, storage, and network resources that are customized based on the applications and customer segments.

[0008] Some have referred to network slicing as the "key ingredient" of 5G, enabling the full potential of 5G architecture to be realized. Network slicing adds an extra 15 dimension to the NFV domain by allowing multiple logical networks to simultaneously run on top of a shared physical network infrastructure. As such, network slicing becomes integral to 5G architecture by creating end-to-end virtual networks that include both networking and storage functions. Operators of a 5G network then can effectively manage diverse 5G use cases with differing throughput, latency and availability demands 20 by partitioning network resources to multiple users or “tenants”. With strategically tuned 1 9 09 9^ network slicing and the optimized allocation of virtual network function (VNF) instances, the cost of operating a 5G architected network can be optimized.

[0009] In respect to the optimization of the configuration of different network slices, zero-touch and fully automated operations and management have become quintessential 5 to harness the potential gain of dynamic resource allocation in an NFV-enabled network slice. To that end, many have proposed autonomous management and orchestration of VNFs, where the CU "learns" to re-configure resources, deploy new VNF instances or offload jobs to a central cloud. One notable proposal refers to a deep reinforced learning (DRL)-based solution dubbed parameterized action twin (PAT) Deep Deterministic 10 Policy Gradient (DDPG) which leverages the actor-critic method to learn to provision network resources to VNFs in an online manner, given the current network state and the requirements of the deployed VNFs.

[0010] Of note, the PAT DDPG solution outperforms all benchmark DRL schemes as well as heuristic greedy allocation in a variety of network scenarios. However, although 15 DDPG is capable of providing excellent results, it has its drawbacks. Like many reinforced learning algorithms training DDPG can be unstable and heavily reliant on finding the correct hyper parameters for the current task. This is caused by the algorithm continuously over-estimating the Q-values of the critic (value) network. These estimation errors accumulate over time and can lead to the agent falling into a local optima or 20 experience catastrophic forgetting. 1 9 09 9^ BRIEF SUMMARY OF THE INVENTION

[0011] Embodiments of the present invention address deficiencies of the art in respect to network slicing in a 5G network and provide a novel and non-obvious method, system 5 and computer program product for the continuous network slicing utilizing a delayed DDPG. In an embodiment of the invention, at least two network slices are defined within a CU of a 5G network architected cellular communications network. Thereafter, at different time steps a network slicing function identifies a state of each of the network slices and determines from a reinforced learning policy assigned to one of the network 10 slices, for a contemporaneous one of the time steps, a scaling operation in allocating different computing resources to corresponding VNFs in the CU for one of the network slices based upon the identified state of the one of the network slices. The network slicing function further applies the determined scaling operation to the CU by allocating the different computing resources to the corresponding VNFs in the CU to the network 15 slice.

[0012] Of importance, the reinforced learning policy includes an actor policy and a critic model. The actor policy takes into account the state of the one of the network slices as input and delivers as output a determined scaling operation allocating different computing resources to corresponding virtual network functions (VNFs) in the CU for 20 one of the network slices based upon the identified state of the one of the network slices. The critic model, in turn, takes into account the state of the one of the network slices in combination with the determined scaling operation as input and delivers a statistical Q-value as output. Optionally, the critic model can be embodied by an amalgamation of 1 9 09 9^ twin critic models with the statistical Q-value being a minimization of the individual Q-values produced by each of the twins. Notably, both the actor policy and the critic model can be implemented according to a deep neural network that self-learns based upon feedback applied to the network. 5

[0013] Once the scaling operation has been applied, the network slicing function monitors a resource cost outcome of the determined scaling operation in the CU and compares the monitored outcome to a pre-determined optimal outcome for the determined scaling operation. The network slicing function then determines a statistical Q-value in the critic model based upon a difference between the monitored outcome and 10 the optimal outcome and computes a gradient for each of the actor policy and the critic model accounting for the determined statistical Q-value. Finally, the network slicing function applies the computed gradients to each of the actor policy and the critic network model for use in a next determined scaling operation at a subsequent one of the time steps. However, the network slicing function applies the gradient computed for the actor 15 policy at a rate which is less frequent than an application of the computed gradient for the critic model.

[0014] In one aspect of the embodiment, the determined scaling operation is determined m consideration of a state space for the one of the network slices, the state space including a number of new UE connections to the one of the network slices, 20 computing resources allocated to each of the VNFs in the CU for the corresponding one of the network slices, a delay status with respect to latency cost for each of the network 1 9 09 9^ slices, an energy status with respect to an energy cost for use of the computing resources by each of the network slices, a number of users served in each one of the network slices, and a number of VNF instantiations in each one of the network slices. In another aspect of the embodiment, the scaling operation is part of a vertical scaling action space that 5 includes scaling up to increased capacity in the one of the network slices, and scaling down to reduced capacity in the one of the network slices. In yet another aspect of the embodiment, the optimal outcome includes a maximized inverse of a total network cost of the monitored outcome at the contemporaneous one of the time steps.

[0015] In another embodiment of the invention, a C-RAN architected data processing 10 system may be adapted for continuous network slicing utilizing a delayed DDPG The system includes a host computing platform disposed within a CU of a 5G network architected cellular communications network. The system also includes a delayed DDPG based continuous network slicing module. The module includes computer program instructions enabled while executing m the host computing platform to define at least two 15 network slices in the CU, load for the network slices a reinforced learning policy that include an actor policy and a critic model. The computer program instructions further continuously at different time steps identify a state of each of the network slices, provide the identified state to the reinforced learning policy and receive from the reinforced learning policy for a contemporaneous one of the time steps, an outputted scaling 20 operation. 1 9 09 9^

[0016] The program instructions yet further apply the outputted scaling operation to the CU by allocating the different computing resources to the corresponding VNFs in the CU to the one of the network slices and monitor a resource cost outcome of the determined scaling operation in the CU while comparing the monitored outcome to a pre- 5 determined optimal outcome for the determined scaling operation. The program instructions even yet further determine a statistical Q-value in the critic model based upon a difference between the monitored outcome and the optimal outcome and compute a gradient for each of the actor policy and the critic model accounting for the determined statistical Q-value. Finally, the program instructions apply the computed gradients, 10 respectively to the actor policy and the critic model, for use in a next determined scaling operation at a subsequent one of the time steps. Notably, the application of the computed gradient for the actor policy occurs at a rate which is less frequent than an application of the computed gradient for the critic model.

[0017] Additional aspects of the invention will be set forth in part in the description 15 which follows, and in part will be obvious from the description, or may be learned by practice of the invention. The aspects of the invention will be realized and attained by means of the elements and combinations particularly pointed out in the appended claims. It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the 20 invention, as claimed. 1 9 09 9^ BRIEF DESCRIPTION OF THE SEVERAL VIEWS OF THE DRAWINGS

[0018] The accompanying drawings, which are incorporated in and constitute part of this specification, illustrate embodiments of the invention and together with the description, serve to explain the principles of the invention. The embodiments illustrated 5 herein are presently preferred, it being understood, however, that the invention is not limited to the precise arrangements and instrumentalities shown, wherein:

[0019] Figure 1 is a schematic illustration of a C-RAN adapted for twin delayed DDPG based continuous network slicing; and,

[0020] Figure 2 is a flow chart illustrating a process for C-RAN architected for twin 10 delayed DDPG based continuous network slicing. DETAILED DESCRIPTION OF THE INVENTION

[0021] Embodiments of the invention provide for C-RAN architected for delayed DDPG based continuous network slicing. The C-RAN includes a host computing platform disposed within a CU of a 5G network architected cellular communications 15 network. A DDPG based continuous network slicing module executes in the memory of the platform and during execution, defines two network slices in the CU. The module also loads for each of the network slices, a reinforced learning policy that includes a contemporaneous actor policy taking the state of the one of the network slices as input and delivering as output a determined scaling operation allocating different computing 20 resources to corresponding VNFs in the CU for a corresponding one of the network slices 1 9 09 9^ based upon the identified state of the one of the network slices. The reinforced learning policy also includes a critic model that takes the state of a corresponding one of the network slices in combination with the determined scaling operation as input and delivers a statistical Q-value as an output. 5

[0022] Once the network slices have been defined and the policies loaded, the module continuously, at different time steps, identifies a state of each of the network slices, provides the identified state to the actor policy and receives from the actor policy for a contemporaneous one of the time steps, an outputted scaling operation. The module then applies the outputted scaling operation to the CU by allocating the different computing 10 resources to the corresponding VNFs in the CU to the one of the network slices. Then, the module monitors a resource cost outcome of the determined scaling operation in the CU while comparing the monitored outcome to a pre-determined optimal outcome for the determined scaling operation.

[0023] Finally, the module determines a statistical Q-value in the critic model based 15 upon a difference between the monitored outcome and the optimal outcome and computes a gradient for each of the actor policy and the critic model accounting for the determined statistical Q-value. The module then applies the computed gradients to each respective one of the actor policy and the critic model for use in a next determined scaling operation at a subsequent one of the time steps. Importantly, however, the 20 module only updates the contemporaneous actor policy with the corresponding computed gradient at a rate which is less frequent than application of each computed gradient to the 1 9 09 9^ critic model. In this way, delayed DDPG based continuous network slicing may be achieved producing a substantial performance improvement over a traditional dynamic reinforced learning-based network slicing process. In particular, delayed DDPG addresses DDPG drawbacks by focusing on reducing the overestimation bias seen in 5 previous algorithms through the addition of 3 key features: • Clipped double Q-learning with pair of critic models • Delayed policy updates and target models • Target policy smoothing and noise regularization

[0024] In further illustration, Figure 1 is a schematic illustration of a C-RAN adapted 10 for delayed DDPG based continuous network slicing. As shown in Figure 1, a C-RAN 130 may be implemented to include a host computing platform 100 that includes one or more computers 110 each with memory 140 and one or more processors 120. Multiple different CUs 150 for respective network slices 170 are defined in the memory 140, each including one or more VNFs 160 to support processing of 5G cellular network 15 connections with different UEs 190 through DUs 180. Importantly, a delayed DDPG network slicing module 200 is included in the memory 140 and executes by at least one of the processors 120 of the host computing platform 100.

[0025] The delayed DDPG network slicing module 200 includes computer program instructions that when executing in the memory 140, receive a state for one of the 20 network slices 170 and provide the state to an actor policy 115 A. The actor policy 115 1 9 09 9^ returns a scaling operation including adding more CPUs 120 to a corresponding one of the CUs 150 for the network slice 170 or removing one or more CPUs 120 from a corresponding one of the CUs 150 for the network slice 170. The program instructions then monitor an outcome of the scaling operation and compare the outcome to an optimal 5 outcome. The comparison is provided to the pair of critic models 135A, 145A along with scaling operation so that the critic models each produce a corresponding statistical Q-value which are then amalgamated with a minimization operation.

[0026] The program instructions then provide the amalgamated Q-value in a gradient for each of the actor policy 115A and the critic models 135A, 145A and the gradients for 10 each of the actor policy 115A and the critic models 135A, 145A are applied, respectively, to a corresponding actor target 115B and pair of critic targets 135B, 145B. Finally, the program instructions update the critic models 135A, 145A with the critic targets 135B, 145B and the program instructions update the actor policy 115A with the actor target 115B, but the program instructions perform the updating of the actor policy 115A at a 15 rate which is less frequent than a rate of updating the critic models 135A, 145 A. In this way, delayed DDPG can be achieved in the C-RAN 130 of the 5G network architected cellular telecommunications network.

[0027] As it will be recognized, the program instructions of the delayed DDPG network slicing module 200 aim to adjust the parameters (p of the actor policy 115A in 20 the direction of a performance gradient VqJ^tp). The performance gradient as to be applied to the actor policy 115A can be reflected mathematically as follows: 1 9 09 9^ ^7(¾) = Jsp^(s) J4           (s, a)dads which equals to (o| S) |a= 7t(s) V©TT© ($)] 5 The actor policy 115A may be parameterized as a value function with the goal of finding the optimal policy 7i(p in which q> includes updating the weight of the actor policy 115 A. The expected return can be approximated in many ways. In one example, the gradient of expected return may be computed according to parameters of (p as V<pJ((p). As can be 10 seen, gradient ascent is preferred over gradient descent for updating the parameters, q>t+i = (pt + aVcpJC^Icpt.

[0028] In the actor-critic method of Figure 1, two models work concurrently where the actor policy 115A is a policy taking state as input and delivering actions as output, while the critic models 135A, 145A each takes states and actions concatenated together and 15 return the Q-value so that the actor policy 115A can be updated through the deterministic policy gradient, V.,.J(0) - Es~„, where 20 is the statistical Q-value also known as the value function or critic.

[0029] More particularly, initially a random experience is stored in buffer [3. In the other words, (st, at, rt, St+1) is stored in order to train a Deep Q-Network. A random batch 25 B is then selected in the buffer p and for all transitions (s® , a® , Hb , s®+l) of P, the predictions are Q(s® , a®) and the targets consider as optimal immediate return that are 1 9 09 9^ exactly first part of the temporal difference (ID) learning error as R(stB , a® ) + vmaxalQfitB-1 ,a)). Over the entire batch B, the loss between predictions and the targets in the batch B may be calculated. Preferably, another target model is used instead of using the Q-network to calculate the target in order to fulfill more stability for the learning 5 algorithm. As it will be recognized, then, the TD process is based on the actor-critic model while leveraging three additional processes in order to improve the TD algorithm:

[0030] (1) Clipped double Q-leaming with pair of critic models:

[0031] The first additional process utilizes two Deep Neural Networks (DNNs) as the two actor models 115 A, 115B and are denoted by <p as a DNN for the actor policy 115A 10 and (p1 as a DNN for the actor target 115B. In addition, two pairs of DNNs are provided for critic models 135A, 145A and critic targets 135B, 145B and denoted as 0i, 02 for the parameterization of a value network as critic models 135A, 145 A, and 0'i, 02 as critic targets 135B, 145B. Therefore, two machine learnings occur simultaneously, namely, Q-learning and policy learning, and the combination addresses approximation error, 15 reduction of bias and finding the highest statistical Q-value. For each element and transition of the batch, the actor target 115B plays a' based on s' while Gaussian noise is added to a'. The critic targets 145A, 145B take the couple (s', a') and return two Q-values Q'ti and Qe as output. Then, the (min Q'ti,Q't2) as an amalgamation of statistical Q-values is considered as an approximated value for DNNs of the critic targets 145 A, 145B. 20

[0032] The DNNs for the critic targets 145A, 145B are used to provide the value estimates through an amalgamation of produced statistical Q-values as follows: 1 9 09 9^ Qt r + 7 * Q^2) Consequently, the DNNs for the two critic models 135A, 135B return two Q-values as Qi(s,a) and Q2(s,a). The loss may then be calculated based on the two critic models 135A, 135B and with Mean Squared Error (MSE). In order to minimize the loss over 5 iterations via back-propagation technique, an efficient optimizer known as Adaptive Moment Estimation may be used: L — Imse(Q^ Qt) + Imse^Qz, Qt) :::: N ' £ [ V„Q», ( S. (,f) (7 10

[0033] (2) Delayed policy updates and target models:

[0034] The second additional process provides for a delayed updating of the actor policy 115 A. Specifically, the DNN of the actor policy 115 A is updated less frequently than the DNNs for the critic models 135A, 135B so as to estimate values with lower variance. The update rule is given by Polyak Averaging, so as to update the parameters 15 by: -- T0f + (1 — T)0^ -- T0 + (1 where t <1 is an hyperparameter to tune the speed of updating. 20

[0035] (3) Target policy smoothing and noise regularization: 1 9 09 9^

[0036] The third additional process acts to smooth the actor target 115B and critic targets 135B, 145B. In this regard, when updating the critic models 135A, 145A, a learning target 135B, 145B using a deterministic policy is highly susceptible to inaccuracies induced by function approximation error, increasing the variance of the 5 target. This induced variance is reduced through regularization to be sure for exploration of all possible continuous parameters. To that end, Gaussian noise is added to the next action a' to prevent two large actions from disturbing the state of environment: a <—+ €, e ~          cr), — c, c) 10 where the noise s is sampled from a Gaussian distribution with zero and certain standard deviation and clipped in a certain range of value between c and c to encourage exploration. To avoid error of using an impossible value of actions, the added noise is clipped to the range of possible actions (min action, max action).

[0037] The foregoing TD3-based network slicing method can be summarized as 15 follows: 1 9 09 9^ Initialize actor network $ and aide networks -¾ Initializie (copy parameters) target networks Initialize replay buffer ,S Import custom gym environment (‘smartech-vO’) white r <maxL-fwesteps' do if r <stotjthen | a = en¥.acdOT_space,s»pteQ else | a-f--'^(s) + c, « . / / ((), 0-) end nexLstate, reward, done, _ = eawstep(a) store the new transition cm, t*, ^ia-i.) into if t >then sample batch of transitions B , afcJ?, r* o . st ^.4.-4) a --^ / (V) + e; e ~ cKp(jV(O; 0 —c, e) Qt = rf7* mm(Q^ , Q^) dy --argmm^A7"1 V(L if reg == 0 then = AT * [VaQ^ 4--r^y + (1 - r)t^ 0' i-- 7^ +(1- r)^z end end if rff.w then | »bs, done = env^setf). False end i=Wl

[0038] In further illustration and summarization of the foregoing TD3-based network slicing methodology, Figure 2 is a flow chart illustrating a process for C-RAN architected for delayed DDPG based continuous network slicing. Beginning in block 210 a network 5 slice is created in the CU as an environment and in block 215, is initialized in memory for the network slice. In block 220, an actor policy and an actor target are loaded into memory for a selected network slice. Then, in block 225 two critic models and two critic targets are also loaded. Thereafter, m block 230, the actions play randomly and in block 1 9 09 9^ 235, a batch of transitions are sampled. In block 240 the actor target takes the next state and plays a next action. In block 245, gaussian noise is added to the next action and the next action is clamped. In block 250 the critic targets compute Q-values from the state and the action and in block 255, a minimum is computed for the Q-values. In block 260, 5 a final target is determined with respect to a discount factor and in block 265 the critic models receive the action and state and return Q-values. In block 270 the critic loss is computed with respect to the final target and block 275, the critic loss is backpropagated in order to update the parameters of the critic models. In decision block 280, it is determined if the updating of the actor model is to be delayed. If yes, in block 285 the 10 actor model is updated with the output of the first critic model and in block 290, the weights of each of the actor and critic targets are updated.

[0039] Thus, in accordance with the present invention, a reward-penalty mechanism is provided in order to mitigate a negative impact of destabilizing training. The rewardpenalty mechanism clips the network values to some constant and constraint values 15 relating to quality of service (QoS) and other thresholds. Thus, as it will be recognized, the proposed technique applies robotic algorithms in the domain of telecommunications. As well, experience replay is one of the main aspects of learning behaviors in biological systems. Here, to accelerate the training process and to improve learning efficiency, a score-based asynchronous actor-learner is optimized for the network slicing environment. 20

[0040] The present invention may be embodied within a system, a method, a computer program product or any combination thereof. The computer program product may 1 9 09 9^ include a computer readable storage medium or media having computer readable program instructions thereon for causing a processor to carry out aspects of the present invention. The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer 5 readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing.

[0041] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or 10 to an external computer or external storage device via a network. The computer readable program instructions may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. Aspects of the present invention are described herein with reference to flowchart illustrations and / or block 15 diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions. 20

[0042] These computer readable program instructions may be provided to a processor of a general-purpose computer, special purpose computer, or other programmable data 1 9 09 9^ processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions may also be 5 stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein includes an article of manufacture including instructions which implement aspects of the function / act specified in the flowchart and / or block diagram block or blocks. 10

[0043] The computer readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other 15 device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0044] The flowchart and block diagrams in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. 20 In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of instructions, which includes one or more executable instructions 1 9 09 9^ for implementing the specified logical function(s). In some alternative implementations, the functions noted in the block may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending 5 upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or carry out combinations of special purpose hardware and computer instructions. 10

[0045] Finally, the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "include", "includes", and / or "including," when used in this specification, specify the 15 presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0046] The corresponding structures, materials, acts, and equivalents of all means or step plus function elements in the claims below are intended to include any structure, 20 material, or act for performing the function in combination with other claimed elements as specifically claimed. The description of the present invention has been presented for 1 9 09 9^ purposes of illustration and description, but is not intended to be exhaustive or limited to the invention m the form disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the invention. The embodiment was chosen and described in order to best explain the 5 principles of the invention and the practical application, and to enable others of ordinary skill in the art to understand the invention for various embodiments with various modifications as are suited to the particular use contemplated.

[0047] Having thus described the invention of the present application in detail and by reference to embodiments thereof, it will be apparent that modifications and variations 10 are possible without departing from the scope of the invention defined in the appended claims as follows: 1 9 09 9^

Claims

We claim:

1. A method for continuous network slicing, the method comprising:defining at least two network slices in a central unit (CU) of a 5G network5 architected cellular communications network;loading for the at least two network slices a reinforced learning policy comprising an actor policy taking the state of one of the at least two network slices as input and delivering as output a determined scaling operation allocating different computing resources to corresponding virtual network functions (VNFs) in the CU for one of the at10 least two network slices based upon the identified state of the one of the network slices, the reinforced learning policy additionally comprising a critic model taking the state of the one of the at least two network slices in combination with the determined scaling operation as input and delivering a statistical Q-value as output for use in reinforcing learning by the actor policy;15 continuously at different time steps identifying a state of each of the at least twonetwork slices, providing the identified state to the reinforced learning policy and receiving from the reinforced learning policy for one of the time steps contemporaneous with the actor policy taking the state of one of the at least two network slices as input, an outputted scaling operation;20 applying the outputted scaling operation to the CU by allocating the differentcomputing resources to the corresponding VNFs in the CU to the one of the at least two network slices;1 9 09 9^monitoring a resource cost outcome of the determined scaling operation in the CU and comparing the monitored outcome to a pre-determined optimal outcome for the determined scaling operation;determining a statistical Q-value in the critic model based upon a difference5 between the monitored outcome and the optimal outcome and computing a gradient for each of the actor policy and the critic model accounting for the determined statistical Q-value; and,applying the computed gradients to each respective one of the actor policy and the critic model for use in a next determined scaling operation at a subsequent one of the time10 steps, but applying one of the computed gradients corresponding to the actor policy at arate which is less frequent than an application of the other of the computed gradients to the critic model.

2. The method of claim 1, wherein the determined scaling operation is determined in 15 consideration of a state space for the one of the network slices comprising a number ofnew user equipment (UE) connections to the one of the network slices, computing resources allocated to each of the VNFs in the CU for the one of the network slices, a delay status with respect to latency cost for each of the at least two network slices, an energy status with respect to an energy cost for use of the computing resources by each of20 the at least two network slices, a number of users served in each one of the at least two network slices and a number of VNF instantiations in each one of the at least two network slices.1 9 09 9^3. The method of claim 1, wherein the scaling operation is part of a vertical scaling action space that includes scaling up to increased capacity in the one of the network slices, and scaling down to reduced capacity in the one of the network slices.

54. The method of claim 1, wherein the optimal outcome comprises a maximized inverse of a total network cost of the monitored outcome at the contemporaneous one of the time steps.10 5. The method of claim 1, wherein the reinforced learning policy comprises twins ofthe critic model.

6. The method of claim 5, wherein the Q-value used in computing the gradient for the actor policy is a minimum of Q-values provided by each of the twins.

157. A cloud radio access network (C-RAN) architected data processing system adapted for continuous network slicing, the system comprising:a host computing platform disposed within a central unit (CU) of a 5G network architected cellular communications network, the CU comprising a communicative20 coupling to a multiplicity of different distributed units (DUs), at least one of the DUs comprising a massive multiple input, multiple output (MIMO) antenna transmitting over1 9 09 9^millimeter wave frequencies, the platform comprising one or more computers, each comprising memory and at least one processor; and,a DDPG based continuous network slicing module comprising computer program instructions enabled while executing in the host computing platform to perform:5 defining at least two network slices in the CU;loading for the network slices a reinforced learning policy comprising an actor policy taking the state of the one of the network slices as input and delivering as output a determined scaling operation allocating different computing resources to corresponding virtual network functions (VNFs) in the CU for one of10 the network slices based upon the identified state of the one of the network slices,the reinforced learning policy additionally comprising a critic model taking the state of the one of the network slices in combination with the determined scaling operation as input and delivering a statistical Q-value as output for use in reinforcing learning by the actor policy;15 continuously at different time steps identifying a state of each of thenetwork slices, providing the identified state to the reinforced learning policy and receiving from the reinforced learning policy for one of the time steps contemporaneous with the actor policy taking the state of one of the at least two network slices as input, an outputted scaling operation;20 applying the outputted scaling operation to the CU by allocating thedifferent computing resources to the corresponding VNFs in the CU to the one of the network slices;1 9 09 9^monitoring a resource cost outcome of the determined scaling operation in the CU and comparing the monitored outcome to a pre-determined optimal outcome for the determined scaling operation;determining a statistical Q-value in the critic model based upon a5 difference between the monitored outcome and the optimal outcome andcomputing a gradient for each of the actor policy and the critic model accounting for the determined statistical Q-value; and,applying the computed gradients to each respective one of the actor policy and the critic model for use in a next determined scaling operation at a subsequent10 one of the time steps, but applying one of the computed gradients correspondingto the actor policy at a rate which is less frequent than an application of the other of the computed gradients to the critic model.

8. The system of claim 7, wherein the determined scaling operation is determined in15 consideration of a state space for the one of the network slices comprising a number ofnew user equipment (UE) connections to the one of the network slices, computing resources allocated to each of the VNFs in the CU for the one of the network slices, a delay status with respect to latency cost for each of the at least two network slices, an energy status with respect to an energy cost for use of the computing resources by each of20 the at least two network slices, a number of users served in each one of the at least two network slices and a number of VNF instantiations in each one of the at least two network slices.1 9 09 9^9. The system of claim 7. wherein the scaling operation is part of a vertical scaling action space that includes scaling up to increased capacity in the one of the network slices, and scaling down to reduced capacity in the one of the network slices.

510. The system of claim 7, wherein the optimal outcome comprises a maximized inverse of a total network cost of the monitored outcome at the contemporaneous one of the time steps.10 11. The system of claim 7, wherein the reinforced learning policy comprises twins ofthe critic model.

12. The system of claim 11, wherein the Q-value used in computing the gradient for the actor policy is a minimum of Q-values provided by each of the twins.1513. A computer program product for continuous network slicing, the computer program product including a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a device to cause the device to perform a method including:20 defining at least two network slices in a central unit (CU) of a 5G networkarchitected cellular communications network;loading for the network slices a reinforced learning policy comprising an actor1 9 09 9^policy taking the state of the one of the network slices as input and delivering as output a determined scaling operation allocating different computing resources to corresponding virtual network functions (VNFs) in the CU for one of the network slices based upon the identified state of the one of the network slices, the reinforced learning policy5 additionally comprising a critic model taking the state of the one of the network slices in combination with the determined scaling operation as input and delivering a statistical Q-value as output for use in reinforcing learning by the actor policy;continuously at different time steps identifying a state of each of the network slices, providing the identified state to the reinforced learning policy and receiving from10 the reinforced learning policy for one of the time steps contemporaneous with the actorpolicy taking the state of one of the at least two network slices as input, an outputted scaling operation;applying the outputted scaling operation to the CU by allocating the different computing resources to the corresponding VNFs in the CU to the one of the network15 slices;monitoring a resource cost outcome of the determined scaling operation in the CU and comparing the monitored outcome to a pre-determined optimal outcome for the determined scaling operation;determining a statistical Q-value in the critic model based upon a difference20 between the monitored outcome and the optimal outcome and computing a gradient for each of the actor policy and the critic model accounting for the determined statistical Q-value; and,1 9 09 9^applying the computed gradients to each respective one of the actor policy and the critic model for use in a next determined scaling operation at a subsequent one of the time steps, but applying one of the computed gradients corresponding to the actor policy at a rate which is less frequent than an application of the other of the computed gradients to5 the critic model.

14. The computer program product of claim 13, wherein the determined scaling operation is determined in consideration of a state space for the one of the network slices comprising a number of new user equipment (UE) connections to the one of the network10 slices, computing resources allocated to each of the VNFs in the CU for the one of the network slices, a delay status with respect to latency cost for each of the at least two network slices, an energy status with respect to an energy cost for use of the computing resources by each of the at least two network slices, a number of users served in each one of the at least two network slices and a number of VNF instantiations in each one of the15 at least two network slices.

15. The computer program product of claim 13, wherein the scaling operation is part of a vertical scaling action space that includes scaling up to increased capacity in the one of the network slices, and scaling down to reduced capacity in the one of the network20 slices.1 9 09 9^16. The computer program product of claim 13, wherein the optimal outcome comprises a maximized inverse of a total network cost of the monitored outcome at the contemporaneous one of the time steps.5 17. The computer program product of claim 13, wherein the reinforced learningpolicy comprises twins of the critic model.

18. The computer program product of claim 17, wherein the Q-value used in computing the gradient for the actor policy is a minimum of Q-values provided by each 10 of the twins.

Citation Information

Patent Citations

  • Improvements in and relating to telecommunication networks

    GB2577055A