Network Slicing Resource Allocation Method and System Based on LSTM and A2C

By building a neural network based on LSTM and A2C, collecting and training user communication data, the low spectrum efficiency and lag problems in network slice resource allocation are solved, and the optimal allocation of spectrum resources and the improvement of system utilization are achieved.

CN115550181BActive Publication Date: 2025-07-22RESEARCH INSTITUTE OF TSINGHUA UNIVERSITY IN SHENZHEN
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211112978.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-14
Publication Date
2025-07-22
Estimated Expiration
2042-09-14

AI Technical Summary

Technical Problem

In the prior art, network slice resource allocation has low spectrum efficiency, difficulty in allocating resources on demand, and lag caused by dynamic changes in base station business requirements is difficult to achieve real-time data statistics and resource optimization.

Method used

Using the network slice resource allocation method based on LSTM and A2C, we use the neural network to collect slice service request data during user communication, and use the A2C algorithm to perform iterative training, output bandwidth allocation strategy, and realize the rational allocation of spectrum resources.

Benefits of technology

When the base station statistics and collects network slice service request data lag, the system spectrum efficiency and system utilization rate are improved, while ensuring the user experience quality of each slice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115550181B_ABST
    Figure CN115550181B_ABST
Patent Text Reader

Abstract

The present invention discloses a network slice resource allocation method and system based on LSTM and A2C. The method includes: constructing a neural network based on LSTM; collecting slice service request data during the actual communication process of users; based on the A2C algorithm, inputting the slice service request data into the neural network for iterative training, and outputting a bandwidth allocation strategy. By using the present invention, it is possible to improve the system spectrum efficiency and system utilization rate in the case where the base station has a lag in statistics and collection of network slice service request data. As a network slice resource allocation method and system based on LSTM and A2C, the present invention can be widely applied to the field of wireless communication technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of wireless communication technologies, and particularly to a method and system for network slice resource allocation based on LSTM and A2C. Background Art

[0002] With the development of wireless communication technologies, diverse application scenarios have emerged. Different application scenarios have different requirements in terms of rate, latency, etc. If an operator only constructs one network, it is difficult to meet the needs of all scenarios simultaneously. For this reason, the network slicing technology has emerged. Through network slicing, an operator can construct multiple logically independent virtual networks on a common physical network. Each virtual network has different functions and can flexibly respond to different service requirements. Therefore, in order to support the functions of different network slices, resources need to be allocated for their slices. However, there are the following problems in slice resource allocation at the present stage. For the radio access network, the spectrum is a scarce resource, so it is difficult to ensure the spectrum efficiency; the service level agreement (SLA) signed with slice tenants usually has strict regulations. Since the users served by the cellular network generally have strong mobility, this will cause the service requirements of each network slice of the base station to change dynamically, making it difficult to allocate network resources on demand; in practical applications, it is difficult to obtain and statistically analyze the service request data of each network slice in real time, and there is a certain lag. Summary of the Invention

[0003] In order to solve the above technical problems, the purpose of the present invention is to provide a method and system for network slice resource allocation based on LSTM and A2C, which can improve the system spectrum efficiency and system utilization rate when there is a lag in the base station's statistics and collection of network slice service request data.

[0004] The first technical solution adopted by the present invention is: A method for network slice resource allocation based on LSTM and A2C, including the following steps:

[0005] Construct a neural network based on LSTM;

[0006] Collect slice service request data during the actual communication process of users;

[0007] Based on the A2C algorithm, input the slice service request data into the neural network for iterative training, and output a bandwidth allocation strategy.

[0008] Further, the construction of the neural network based on LSTM includes a prediction network, a state processing network, a value network, and a policy network, where:

[0009] The prediction network consists of four layers of LSTM networks and two layers of fully connected layers. The first fully connected layer is used to process the feature of the service request data, and the second fully connected layer is used to process the output data of the first fully connected layer to obtain the final prediction result;

[0010] The state processing network serves as an intermediate layer between the observation variable and the A2C network. This layer of network consists of a single layer of LSTM and is used to extract features from the observation variable;

[0011] The value network includes two layers of fully connected layers. Among them, the first fully connected layer is used to extract value features from the output data of the state processing network, and the second fully connected layer processes the output data of the first fully connected layer to obtain an estimated value of the state value;

[0012] The policy network includes two layers of fully connected layers. Among them, the first fully connected layer is used to extract features from the output data of the state processing network, and the second fully connected layer is used to process the output data of the first fully connected layer for policy probability distribution;

[0013] Further, the step of collecting the slice service request data in the actual communication process of the user specifically includes:

[0014] Set the wireless spectrum bandwidth, randomly adopt a bandwidth allocation strategy, and determine three types of UE users with different service types;

[0015] Based on the wireless communication scenario, perform data interaction between the wireless spectrum bandwidth and the UE users to obtain the slice service request data in the actual communication process of the user.

[0016] Further, the step of inputting the slice service request data into the neural network for iterative training based on the A2C algorithm and outputting the bandwidth allocation strategy specifically includes:

[0017] Based on the A2C algorithm, input the slice service request data into the neural network;

[0018] Based on the prediction network, predict the historical data to obtain the predicted slice service request data, where the historical data is the historical interaction data between the wireless spectrum bandwidth and the UE users;

[0019] Concatenate the slice service request data and the predicted slice service request data to obtain the concatenated state vector;

[0020] Based on the state processing network, value network and policy network, train the concatenated state vector to obtain the first bandwidth allocation strategy;

[0021] According to the first bandwidth allocation strategy, perform data interaction between the wireless spectrum bandwidth and UE users, obtain the first slice service request data, and update the neural network;

[0022] Input the first slice service request data into the updated neural network for training to obtain the second bandwidth allocation strategy;

[0023] Repeat the update steps of the recurrent neural network until the preset number of iterative updates is satisfied, and output the bandwidth allocation strategy.

[0024] Furthermore, the step of training the spliced state vector based on the state processing network, value network, and policy network to obtain the first bandwidth allocation strategy specifically includes:

[0025] Transmit the spliced data to the state processing network for feature extraction processing to obtain the eigenvalue of the spliced state vector;

[0026] Calculate the eigenvalue of the state vector based on the advantage function on the value network, and output the estimated value of the state value;

[0027] Explore the eigenvalue of the spliced state vector based on the policy network, and output the first bandwidth allocation strategy.

[0028] Furthermore, the reward function for the prediction network to predict historical data is expressed as follows:

[0029]

[0030] In the above formula, R t represents the feedback reward function, q e represents the average experience quality of eMBB slice users, q v represents the average experience quality of VoLTE slice users, q u represents the average experience quality of URLLC slice users.

[0031] Furthermore, the advantage function for value network feature calculation is as follows:

[0032] δ t (S t ; θ v ) = R t +γV(S t+1 ; θ v ) - V(S t ; θ v )

[0033] In the above formula, θ v represents the value network parameter, γ represents the discount factor, R t represents the feedback reward, S tLet \(\mathbf{s}\) denote the state vector and \(V\) denote the estimated value of this state.

[0034] The loss function of the value network is expressed as follows:

[0035]

[0036] In the above formula, denotes the loss function of the value network.

[0037] Furthermore, the exploration loss function of the policy network is as follows:

[0038]

[0039] In the above formula, denotes the loss function of the value network, \(\theta_{\pi}\) p denotes the policy network parameters, \(\theta_{v}\) v denotes the value network parameters, \(\delta\) t denotes the advantage function, \(s\) t denotes the state vector, \(a\) t denotes the feedback reward, \(\eta\) denotes the entropy weight, and \(H\) denotes the entropy regularization term.

[0040] The second technical solution adopted by the present invention is: a network slice resource allocation system based on LSTM and A2C, including:

[0041] A construction module for constructing a neural network based on LSTM;

[0042] An acquisition module for acquiring slice service request data during the actual communication process of users;

[0043] An output module, based on the A2C algorithm, inputs the slice service request data into the neural network for iterative training and outputs a bandwidth allocation strategy.

[0044] The beneficial effects of the method and system of the present invention are: the present invention trains a neural network based on LSTM by acquiring slice service request data during the actual communication process of users, considering the actual application scenarios where users have mobility and the statistical and collection of slice aggregation requests of different service types have hysteresis, and predicts the slice service request data of the next time slot based on the slice service request data of the previous time slot collected, and obtains the slice bandwidth strategy through the Advantage Actor-Critic (A2C) reinforcement learning algorithm, which can realize the reasonable allocation of spectrum resources through intelligent learning in the case of hysteresis in the base station's statistical and collection of network slice service request data, and effectively improve the system spectrum efficiency (SE) and system utilization while ensuring the quality of experience (QoE) of each slice user. Description of the Drawings

[0045] Figure 1It is the flowchart of the steps of the network slice resource allocation method based on LSTM and A2C of the present invention;

[0046] Figure 2 It is the structural block diagram of the network slice resource allocation system based on LSTM and A2C of the present invention;

[0047] Figure 3 It is the network structure diagram of the deep reinforcement learning neural network of the present invention;

[0048] Figure 4 It is the schematic diagram of the simulation parameters of the wireless network environment of the present invention;

[0049] Figure 5 It is the specific parameter setting diagram of the network slice in the present invention;

[0050] Figure 6 It is the comparison diagram of the predicted value and the true value of the VoLTE slice aggregation request by the deep reinforcement learning algorithm (Double-A2C-LSTM) of the present invention;

[0051] Figure 7 It is the comparison diagram of the predicted value and the true value of the eMBB slice aggregation request by the deep reinforcement learning algorithm (Double-A2C-LSTM) of the present invention

[0052] Figure 8 It is the comparison diagram of the predicted value and the true value of the URLLC slice aggregation request by the deep reinforcement learning algorithm (Double-A2C-LSTM) of the present invention;

[0053] Figure 9 It is the comparison diagram of the system SE performance between the deep reinforcement learning algorithm (Double-A2C-LSTM) and the A2C-LSTM algorithm with a delay of 5 time slots;

[0054] Figure 10 It is the comparison diagram of the system utilization rate between the deep reinforcement learning algorithm (Double-A2C-LSTM) and the A2C-LSTM algorithm with a delay of 5 time slots;

[0055] Figure 11 It is the comparison data diagram of the user experience quality of the VoLTE slice between the deep reinforcement learning algorithm (Double-A2C-LSTM) and the A2C-LSTM algorithm with a delay of 5 time slots of the present invention;

[0056] Figure 12 It is the comparison data diagram of the user experience quality of the eMBB slice between the deep reinforcement learning algorithm (Double-A2C-LSTM) and the A2C-LSTM algorithm with a delay of 5 time slots of the present invention;

[0057] Figure 13It is a comparison data graph of the user experience quality of the URLLC slice between the deep reinforcement learning algorithm (Double-A2C-LSTM) of the present invention and the A2C-LSTM algorithm with a 5-slot delay. Detailed implementation manners

[0058] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. For the step numbers in the following embodiments, they are only set for the convenience of description and explanation, and no limitation is imposed on the order between steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0059] Refer to Figure 1 , the present invention provides a network slice resource allocation method based on LSTM and A2C, and the method includes the following steps:

[0060] S1. Construct a neural network based on LSTM;

[0061] Specifically, refer to Figure 3 , the neural network consists of four parts: a prediction network, a state processing network, a value network, and a policy network. Among them, the prediction network is composed of four layers of LSTM and two layers of fully connected layers. The number of LSTM neurons in the four layers is 512, 256, 128, and 64 respectively, and the number of neurons in the two layers of fully connected layers is 32 and 3 respectively. The activation function of each layer is tanh;

[0062] The state processing network is composed of a single layer of LSTM, and the number of LSTM neurons is 64;

[0063] The value network is composed of two layers of fully connected layers. The first layer of fully connected layer extracts features related to value, has 32 neurons, and the activation function is tanh. The second layer of fully connected layer has 1 neuron to output the state value V(S t );

[0064] The policy network (outputting the bandwidth allocation policy) is composed of two layers of fully connected layers and is responsible for generating actions. The first layer takes S′ t as the input and is responsible for extracting action features, has 32 neurons, and the activation function is tanh. The second layer outputs the probability distribution of x policies, and the sum of all policy probabilities is 1. A random policy is adopted to select the bandwidth allocation scheme;

[0065] Further, set the parameters corresponding to the network. The learning rate of the prediction network is set to 0.001, the learning rate of the policy network is set to 0.005, the learning rate of the value network is set to 0.008, the observation length T is set to 10, the delay time slot X is set to 5, the discount factor γ is set to 0.9, the entropy weight η is set to 0.001, the weight α of the optimization target is set to 0.01, β is set to [1, 1, 1], the slice bandwidth adjustment period is 1 s, the users within the slice are polled and scheduled with a scheduling time slot of 0.5 ms, and the number of neural network training times is 5000.

[0066] S2. Collect the slice service request data during the actual communication process of users;

[0067] Specifically, initialize the system parameters, and determine that there are different service type slices in a single base station. The user set served by each slice is denoted as The total number of aggregated request data packets for each slice is denoted as d = (d1, d2, …, d N ). All slices share the total spectrum resource W. The intelligent controller adopts a random bandwidth allocation strategy and interacts with the environment for preliminary exploration T times. At the t time slot, obtain the slice service request data before the (t - x - 1) time slot, form historical memories and save them in the cache, that is, historical data;

[0068] Further, referring to Figure 4 and Figure 5 , set the parameters of the wireless network environment. Consider a radio access network scenario with three types of services (i.e., VoLTE, eMBB, URLLC) and a 1200 m × 1200 m simulation area. There are 800 users and a single base station. The wireless network environment simulation parameters and explanations are as Figure 4 shown; the ratio of VoLTE, eMBB, and URLLC users is 1:2:2. Assume that the users within the same slice share the mobile mode (including the distribution of speed and moving direction). When a user reaches the boundary of the simulation area, its direction will bounce back, as Figure 5 describes the specific configuration and moving speed of each user. In addition, each user Figure 5 generates traffic according to the parameters in

[0069] S3. Based on the A2C algorithm, input the slice service request data into the neural network for iterative training, and output the bandwidth allocation strategy.

[0070] S31. Based on the prediction network, predict the historical data to obtain the predicted slice service request data. The historical data is the historical interaction data between the wireless spectrum bandwidth and UE users;

[0071] Specifically, according to the output of the policy network, a random policy is adopted to execute the bandwidth allocation scheme. After performing an action in the environment, a feedback reward R is obtained. t The slice aggregation request P obtained at this moment t-x is put into the cache. The reward function formula is as follows:

[0072]

[0073] In the above formula, R t represents the feedback reward function, q e represents the average experience quality of eMBB slice users, q v represents the average experience quality of VoLTE slice users, q u represents the average experience quality of URLLC slice users;

[0074] S32. Concatenate the slice service request data with the predicted slice service request data, and transmit the concatenated data to the state processing network for feature extraction processing to obtain the eigenvalue of the slice service request data;

[0075] Specifically, send S t+1 ={P t-T-x+1 ,…,P t-x} into the prediction network to recursively obtain X data, concatenate it with S t+1 and send it into the value network to obtain the true state value V(S t+1 ), so as to obtain the advantage function as follows:

[0076] δ t (S t ; θ v ) = R t +γV(S t+1 ; θ v ) - V(S t ; θ v )

[0077] In the above formula, θ v represents the value network parameter, γ represents the discount factor, R t represents the feedback reward, S t represents the state vector, and V represents the estimated value of the state value;

[0078] To sum up, the prediction network and the state processing network constructed based on LSTM have the following functions. The training objective of the prediction network is to minimize the error between the predicted value and the true value. Due to the lag in data statistics and collection, at time slot t, the controller makes x predictions in total to obtain {p′ t-x ,…,p′ t-1}; but at the next time slot t + 1, the controller can only obtain the true p t-x, so we only update the prediction network based on the prediction error at the t-x moment. And each time the intelligent controller obtains a real value, the loss function uses the Mean Square Error (MSE) and the Adam (Adaptive momentum) optimizer to update the neural network parameters. Among them, the loss function is as follows:

[0079]

[0080] In the above formula, y represents the real value, y' represents the predicted value, and N represents the slices of different service types shared by a single base station;

[0081] S33. Calculate the eigenvalue of the state vector based on the advantage function on the value network, and output an estimated value of the state value;

[0082] Specifically, the output of the estimated value of the state value is used to guide the policy network to learn better policies and output better policy schemes;

[0083] The loss function of the value network is as follows:

[0084]

[0085] In the above formula, represents the loss function of the value network;

[0086] The network parameters are updated using the gradient descent method. The update formula of the value network is as follows:

[0087]

[0088] S34. Explore the eigenvalue of the state vector based on the policy network and output the first bandwidth allocation policy;

[0089] Specifically, the policy network adds entropy regularization to the loss function of the policy network to encourage exploration. η is the entropy value weight, and the calculation formula of the loss function is as follows:

[0090]

[0091] In the above formula, represents the loss function of the value network, θ p represents the policy network parameters, θ v represents the value network parameters, δ t represents the advantage function, s t represents the state vector, a t represents the feedback reward, η represents the entropy value weight, and H represents the entropy regularization term;

[0092] The gradient descent method is used to update the network parameters, and the update formula is as follows:

[0093]

[0094] S35. The cyclic sliced service request data is input into the updated neural network for training until the number of iterations is satisfied.

[0095] Specifically, based on the first bandwidth allocation policy obtained in step S34, the controller allocates bandwidth resources to different network slices in the wireless communication system according to the first bandwidth allocation policy. At this time, the wireless communication system obtains a new allocation policy, resulting in a change in the network environment. The controller can obtain the data interaction between the wireless spectrum bandwidth and UE users based on this first bandwidth allocation policy again as new sliced service request data, and input it into the updated neural network for training again, and so on in a loop until the final preset number of iterations is satisfied;

[0096] As Figures 6 to 8 shown, the present invention considers the data of 10 historical moments to recursively predict the data of the next 5 time slots. The prediction errors are gradually accumulated, and the error of the final prediction result is the largest. Therefore, the real value of the last predicted data is compared with the predicted value;

[0097] As Figures 9 to 13 shown, it is a performance comparison diagram of the deep reinforcement learning algorithm (Double-A2C-LSTM) proposed in the present invention and the A2C-LSTM algorithm without considering the predicted time slots. The simulation considers that the LSTM-A2C lacks the data of 5 adjacent time slots and compares it with the Double-LSTM-A2C algorithm with a prediction function. The simulation results are the average values of the results of 5 independent experiments. Since LSTM-A2C does not use prediction, it will lack the data of 5 adjacent time slots due to hysteresis. From the analysis of the result graph, the average user QoE of different slices of both algorithms can reach more than 98%. However, the system SE and utilization rate of the algorithm proposed in the present invention are better than those of the LSTM-A2C algorithm. Because both algorithms give priority to ensuring the average user QoE of different slices, thereby improving the system SE and utilization rate. Due to the lack of data in the LSTM-A2C algorithm, LSTM cannot accurately capture the change law of sliced requests, resulting in the reinforcement learning being unable to learn the optimal bandwidth allocation policy. While the algorithm proposed in the present invention can accurately complement the missing data, so that the state processing network can accurately capture the change law of sliced requests, assist the reinforcement learning to learn the optimal bandwidth allocation policy, and improve the spectrum utilization efficiency under the condition of limited resources.

[0098] Referring to Figure 2 , the network slice resource allocation system based on LSTM and A2C includes:

[0099] A building module for building an LSTM-based neural network;

[0100] An acquisition module for acquiring slice service request data during the actual communication process of users;

[0101] An output module, based on the A2C algorithm, inputs the slice service request data into the neural network for iterative training and outputs a bandwidth allocation strategy.

[0102] The content in the above method embodiments is applicable to the system embodiments. The functions specifically implemented by the system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0103] The above has specifically described the preferred embodiments of the present invention. However, the present invention is not limited to the described embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A method for allocating network slice resources based on LSTM and A2C, characterized in that It includes the following steps: Construct an LSTM-based neural network; Collect slice service request data during the actual communication process of users; Based on the A2C algorithm, input the slice service request data into the neural network for iterative training, and output the bandwidth allocation strategy; The construction of the LSTM-based neural network includes a prediction network, a state processing network, a value network, and a policy network, where: The prediction network consists of four layers of LSTM networks and two layers of fully connected layers. The first fully connected layer is used to perform feature processing on the service request data, and the second fully connected layer is used to process the output data of the first fully connected layer to obtain the final prediction result; The state processing network serves as an intermediate layer between the observation variable and the A2C network. This layer of network consists of a single layer of LSTM and is used to extract features from the observation variable; The value network includes two layers of fully connected layers. Among them, the first fully connected layer is used to perform value feature extraction processing on the output data of the state processing network, and the second fully connected layer processes the output data of the first fully connected layer to obtain an estimated value of the state value; The policy network includes two layers of fully connected layers. Among them, the first fully connected layer is used to perform feature extraction processing on the output data of the state processing network, and the second fully connected layer is used to perform policy probability distribution processing on the output data of the first fully connected layer.

2. The network slice resource allocation method based on LSTM and A2C according to claim 1, characterized in that The step of collecting slice service request data during the actual communication process of users specifically includes: Set the wireless spectrum bandwidth, randomly adopt the bandwidth allocation strategy, and determine three types of UE users with different service types; Based on the wireless communication scenario, perform data interaction between the wireless spectrum bandwidth and the UE users to obtain the slice service request data during the actual communication process of users.

3. The method for allocating network slice resources based on LSTM and A2C according to claim 2, wherein The step of, based on the A2C algorithm, inputting the slice service request data into the neural network for iterative training and outputting the bandwidth allocation strategy specifically includes: Based on the A2C algorithm, input the slice service request data into the neural network; Based on the prediction network, predict the historical data to obtain the predicted slice service request data, where the historical data is the historical interaction data between the wireless spectrum bandwidth and the UE users; Concatenate the slice service request data and the predicted slice service request data to obtain the concatenated state vector; Based on the state processing network, the value network, and the policy network, train the concatenated state vector to obtain the first bandwidth allocation strategy; According to the first bandwidth allocation strategy, perform data interaction between the wireless spectrum bandwidth and the UE users to obtain the first slice service request data and update the neural network; Input the first slice service request data into the updated neural network for training to obtain the second bandwidth allocation strategy; Repeat the update step of the neural network until the maximum preset number of iterations is reached, terminate the loop, and output the bandwidth allocation strategy.

4. The network slice resource allocation method based on LSTM and A2C according to claim 3, wherein The step of, based on the state processing network, the value network, and the policy network, training the concatenated state vector to obtain the first bandwidth allocation strategy specifically includes: Transmit the concatenated state vector to the state processing network for feature extraction processing to obtain the eigenvalue of the state vector; Calculate the eigenvalues of the state vector based on the advantage function on the value network and output the estimated value of the state value; Explore the eigenvalues of the state vector based on the policy network and output the first bandwidth allocation policy.

5. The network slice resource allocation method based on LSTM and A2C according to claim 4, characterized in that, The reward function for which the prediction network predicts historical data is expressed as follows: In the above formula, R t represents the feedback reward function, and q e represents the average experience quality of eMBB slice users, q v represents the average experience quality of VoLTE slice users, q u represents the average experience quality of URLLC slice users.

6. The network slice resource allocation method based on LSTM and A2C according to claim 4, wherein The advantage function for value network feature calculation is as follows: δ t (s t ; θ v ) = R t + γV(s t+1 ; θ v ) - V(s t ; θ v ) In the above formula, θ v represents the value network parameter, γ represents the discount factor, and R t represents the feedback reward, s t represents the state vector, and V represents the estimated value of the state value; The loss function of the value network is expressed as follows: In the above formula, represents the loss function of the value network.

7. The network slice resource allocation method based on LSTM and A2C according to claim 4, characterized in that The exploration loss function of the policy network is as follows: In the above formula, represents the loss function of the value network, θ p represents the policy network parameters, θ v represents the value network parameters, δ t represents the advantage function, s t represents the state vector, a t represents the feedback reward, η represents the entropy value weight, and H represents the entropy regularization term.

8. A network slice resource allocation system based on LSTM and A2C, characterized in that, A network slicing resource allocation method based on LSTM and A2C as claimed in claim 1, comprising the following modules: A construction module for constructing a neural network based on LSTM; An acquisition module for acquiring slice service request data during the actual communication process of users; An output module, based on the A2C algorithm, inputs the slice service request data into the neural network for iterative training and outputs a bandwidth allocation policy.

Citation Information

Patent Citations

  • Method for training action planning model and target searching method

    CN110059646A

  • Multi-base-station cooperative wireless network resource allocation method based on spatial-temporal feature extraction reinforcement learning

    CN113811009A