Heterogeneous network bandwidth adaptive allocation method based on reinforcement learning

By applying reinforcement learning-based methods in heterogeneous networks and dynamically adjusting bandwidth allocation strategies, the problem of dynamic changes in bandwidth requirements in the existing technology is solved, and efficient utilization of network resources and improvement of user experience is achieved.

CN119945992APending Publication Date: 2025-05-06NANJING NORTH OPTICAL ELECTRONICS
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411968652.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art cannot effectively respond to the dynamic changes in bandwidth requirements in heterogeneous networks, resulting in poor flexibility and the inability to achieve the optimal matching of service requirements and bandwidth resources.

Method used

Using reinforcement learning-based methods, the bandwidth allocation strategy is dynamically adjusted through data acquisition, demand analysis and modeling, optimization matching and dynamic adjustment steps to achieve the optimal matching of business requirements and bandwidth resources.

Benefits of technology

It realizes dynamic adjustment and optimized allocation of heterogeneous network bandwidth, improves network resource utilization efficiency and user experience, and meets the flexibility of changes in user demand and fluctuations in network traffic.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119945992A_ABST
    Figure CN119945992A_ABST
Patent Text Reader

Abstract

The invention discloses a heterogeneous network bandwidth adaptive allocation method based on reinforcement learning. The method comprises the following steps: collecting bandwidth use information of various communication devices and bandwidth requirements of different types of services on a terminal through a user terminal; classifying and analyzing the collected user demand data to obtain the relationship between the service flow evaluation degree and the communication bandwidth; a reinforcement learning algorithm is adopted to match service requirements and bandwidth resources, decision making is carried out on bandwidth requirements at future moments to obtain optimal service quality of services, and dynamic adjustment and distribution of heterogeneous network bandwidths are achieved. Compared with the prior art, the adaptive bandwidth allocation method for the heterogeneous network can dynamically adjust the allocation strategy of bandwidth resources in real time according to user requirements and network conditions by adopting an optimization algorithm based on reinforcement learning. Through accurate modeling and optimization of bandwidth requirements of different service types, optimal configuration of bandwidth resources is realized, and then the utilization rate of network resources and the service quality of services are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of communication networks, and in particular to a heterogeneous network bandwidth adaptive allocation method based on reinforcement learning. Background Art

[0002] In natural disaster scenarios, disaster areas are isolated network information islands and urgently need external communications. Faced with the increasing demand for network bandwidth due to a large number of business needs (such as video, images, information, etc.) in a short period of time, it is urgent to solve the adaptive allocation of communication bandwidth under limited spectrum resources through heterogeneous communication networks.

[0003] Traditional bandwidth allocation methods are usually based on static priority settings, such as manually assigning priorities to different types of services in advance to ensure bandwidth guarantees for various types of services. For example, voice and video communications with high real-time requirements are given the highest priority to ensure low latency and high-quality transmission; while services with lower real-time requirements, such as file downloads and email synchronization, are set to lower priorities to allow appropriate delays when bandwidth resources are tight. This method is simple and easy to understand, but it cannot cope with the dynamic changes in bandwidth requirements in the network environment and has poor flexibility. Therefore, how to achieve the optimal match between service requirements and bandwidth resources in heterogeneous networks has become an important issue. Summary of the invention

[0004] In view of the above problems, the present invention provides a heterogeneous network bandwidth adaptive allocation method based on reinforcement learning, which uses reinforcement learning algorithm to match business demand with bandwidth resources, and makes decisions on bandwidth demand at future moments to obtain the optimal business service quality, thereby realizing dynamic adjustment and allocation of heterogeneous network bandwidth, and improving the utilization efficiency of network resources and user experience.

[0005] The method of the present invention mainly comprises the following steps: A heterogeneous network bandwidth adaptive allocation method based on reinforcement learning comprises the following steps:

[0006] Step 1: Data collection: Collect bandwidth usage information of various communication devices and bandwidth requirements of different types of services on the terminal through user terminals;

[0007] Step 2: Demand analysis and modeling: Classify and analyze the collected user demand data, identify the bandwidth demand characteristics of different service types, perform mathematical modeling for different service types, and obtain the relationship between service flow evaluation and communication bandwidth;

[0008] Step 3: Optimize matching: Evaluate the bandwidth resource status of the current heterogeneous network, use reinforcement learning algorithms to match business needs with bandwidth resources, and make decisions on bandwidth needs in the future to obtain the best business service quality;

[0009] Step 4: Dynamic adjustment: Dynamically adjust the bandwidth allocation strategy based on the reinforcement learning model.

[0010] Compared with the prior art, the present invention has the following beneficial effects:

[0011] (1) Adaptive bandwidth resource allocation: By adopting an optimization algorithm based on reinforcement learning, the bandwidth resource allocation strategy can be dynamically adjusted in real time according to user needs and network conditions. This approach can flexibly respond to network traffic fluctuations, meet uncertainties such as changes in user needs, and improve network resource utilization and service quality.

[0012] (2) Improving service quality: By accurately modeling and optimizing the bandwidth requirements of different service types (such as emergency voice, data files, video streaming, etc.), the optimal allocation of bandwidth resources can be achieved, thereby improving the service quality of various services, ensuring that users' bandwidth requirements are met, and avoiding excessive or insufficient bandwidth allocation. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 Allocate specific processes for network resources based on reinforcement learning. DETAILED DESCRIPTION

[0014] In order to more clearly illustrate the technical solutions of the embodiments of this specification, the symbols and formulas required for describing the embodiments are briefly introduced below.

[0015] Step 1: Data Collection:

[0016] In a heterogeneous network that includes multiple communication methods, user terminals are connected to various communication devices through wired connections and can obtain network parameters of communication devices at all times. In addition, various business applications run in user terminals and can report network bandwidth demand parameters required by the business to the user terminals.

[0017] The communication network parameters and service bandwidth requirement parameters collected mainly include: various service types and network parameter types; n, p, t are used to represent service index, network type index and time index respectively, and the definitions are: Among them, N is the total number of services, P is the total number of network communication means (network type), T is the total time; B represents the user communication bandwidth, B n Indicates the total bandwidth required for the transmission of the nth type of service. represents the maximum bandwidth required by the nth service at time t, It represents the maximum network bandwidth that can be provided by the p-th communication means at time t.

[0018] Step 2: Demand analysis and modeling: Classify and analyze the collected user demand data to identify the bandwidth demand characteristics of different service types; establish a service demand evaluation algorithm model for user demand data, and extract three types of service models from N types of services. The service types are mainly divided into three types of service models: emergency voice service, emergency data file service, and emergency video streaming service.

[0019] Mathematical modeling is mainly carried out for different business types to obtain the relationship between the business flow evaluation degree U(B) and the communication bandwidth B.

[0020] (1) For emergency voice services, generally speaking, as long as the bandwidth obtained exceeds the threshold B required by the service, th , the service evaluation degree will remain unchanged as the bandwidth increases. Its service flow evaluation function can be expressed as:

[0021]

[0022] Where B is the user communication bandwidth; B th is the threshold value and does not change with time.

[0023] (2) For emergency data file services, the service evaluation will also increase with the increase of allocated bandwidth, but the rate of increase will gradually slow down with the increase of bandwidth. Its service flow evaluation function can be expressed as:

[0024]

[0025] Where B is the communication bandwidth; max It is the bandwidth obtained by reaching the maximum service evaluation value, which does not change with time. e is the base of the natural logarithm.

[0026] (3) For emergency video streaming services, the relationship between output quality and service evaluation is different: as bandwidth resources increase, service evaluation increases. When the bandwidth is less than the critical value of the service quality required by the video streaming service, the function decreases monotonically; when it is greater than the critical value of the service quality required by the video streaming service, the function increases monotonically. The modeling function can be expressed as:

[0027]

[0028] In the formula, B is the communication bandwidth; C is the average value of the basic bandwidth requirement and the maximum bandwidth requirement of the video streaming service, which does not change with time.

[0029] Step 3: Optimization matching stage: The total transmission bandwidth required by N types of services needs to obtain the total maximized service evaluation flow within time T, and the network bandwidth that P types of communication means can provide at different times It is randomly allocated to N services. Combining equations (1), (2) and (3), different service bandwidths Bth , B max Different from C, to evaluate the bandwidth resource status of heterogeneous networks at different times, it is necessary to calculate the available bandwidth of heterogeneous networks at different times t. Assigned to different types of services n, that is, In the following, we use an optimization algorithm based on reinforcement learning to match business needs with bandwidth resources.

[0030] The optimal service quality is obtained by setting optimization goals and making decisions on bandwidth requirements at future times. The specific steps are as follows:

[0031] Step 3.1: Establish optimization objective function

[0032] According to the actual situation of the business flow in step 2, the bandwidth resource allocation problem can be transformed into a simple addition of the evaluation utilities of each business flow, specifically:

[0033]

[0034] In the formula, It is expressed as the sum of the communication bandwidth B allocated to the nth type of user by P communication means at time t, and the obtained service flow evaluation degree.

[0035] Step 3.2: Build a reinforcement learning optimization model

[0036] Based on the above problems, combined with the Markov model of reinforcement learning, the current state is modeled: That is, the state space s is constructed by the bandwidth that each communication means service can provide at time t, the total bandwidth demand of each service type, and the current bandwidth demand of each service type. t ;

[0037] It is the action space, which indicates the size of bandwidth allocated to each service by each communication means in the heterogeneous network at the current moment;

[0038] That is, in the current spatial state s t Next, select action state a t The return you can get after t (i.e. business value evaluation value);

[0039] Then establish the final optimization goal:

[0040] That is, the maximum value R of all business value evaluations within a period of time T total .

[0041] Step 3.3: Model training based on reinforcement learning algorithm;

[0042] (1) Based on Actor-Critic architecture:

[0043] The resource allocation of the present invention adopts a reinforcement learning algorithm based on the Actor-Critic architecture and constructs two types of neural networks, namely:

[0044] Actor: responsible for generating strategies (i.e. from state s t To action a t The neural network of the mapping).

[0045] That is, the input state is taken and a continuous action (i.e., strategy) is output. The training goal is to maximize the expected cumulative reward by using deterministic policy gradients to update the neural network parameters.

[0046] Critic: Evaluates the quality of the action selected by the Actor (i.e., calculates the action a t Q value) of the neural network.

[0047] That is, the input state s t and action a t , output the current Q value estimate, i.e. Q(s t , a t ), indicating that in state s t Take action a t The long-term return of the training is to minimize the TD (time difference) error and update the Q value estimate.

[0048] (2) Neural network construction

[0049] Construct the initial neural network θ of Actor and Critic μ and θ Q , according to the neural network input state dimension, that is, s t and a t The dimension size determines the number of layers and the number of neural networks. At the same time, the target neural network of the Actor network and the Critic is constructed. The number of layers and the number of target neural networks are the same as the initial neural network model of the Actor network and the Critic. In order to distinguish, the parameter is θ μ′ and θ Q′ express.

[0050] In order to fully explore the early stage of neural network architecture training, noise is added to the output of the Actor's initial neural network to generate a noise signal with time correlation, thereby improving the exploration in the continuous action space.

[0051] In order to break the correlation between data, an experience replay pool is used. At each time step, the initial user state will store the experienced state transition in the replay pool, that is, the four-tuple sample (s t ,a t ,s t+1 ,r t ). At each update, a batch is randomly drawn from the replay pool for training.

[0052] (3) Neural network training process

[0053] In each round, the user sets the state space s of the initial state t Input into the Actor initial neural network to get the current action state a t , after executing the action, we get the next state space s t+1 , and the current reward r t The user stores the state transitions he experiences in the replay pool, i.e., the four-tuple sample (s t ,a t ,s t+1 ,r t ), and training can be started when there is enough data in the replay pool.

[0054] During training, sample batches from the replay pool: randomly draw a small batch of data from the replay pool.

[0055] Update the initial critic network: Use batch data to calculate the target Q value and update the parameters of the critic network. The target Q value is calculated by the Bellman equation, that is, minimizing the TD (time difference) error.

[0056] Loss(θ Q )=[r t +λQ(s t+1 ,a t+1 |θ Q′ )-Q(s t ,a t |θ Q )] 2 (5)

[0057] In the formula, Loss is the loss function, which represents the difference between the value output by the Critic initial network and the expected target value. This value is used to update the Critic initial network θ Q parameter, λ is a constant between 0 and 1, Q(s t ,a t |θ Q ) is the value output by the initial Critic network; Q(s t+1 ,a t+1 |θ Q′) is the value output by the Critic target network.

[0058] Update the Actor initial network: Use the policy gradient method to update the parameters of the Actor network.

[0059]

[0060] In formula (6), Represents the Actor network θ μ The gradient of Q(s t ,μ(s t )|θ Q ) means in state s t Next, select action μ(s t ) (i.e., the action obtained according to the Actor's initial network strategy μ), the value of the Critic's initial network output, i.e., the expected cumulative reward; Indicates the value of the obtained Critic initial network output relative to the action a(μ(s t )) represents the gradient of the expected cumulative return at the current state s t Next, for the selected action μ(s t ). The gradient can be used to adjust the initial network parameters of the Actor to increase the expected return; μ(s t |θ μ ), indicating that in state s t Next, based on the current Actor initial network parameters θ μ The probability distribution of the chosen action. Indicates that in state s t Under this condition, the initial network strategy μ of Actor has an effect on the parameter θ μ The gradient of , computes the action that adjusts the policy parameters to improve the policy's choice.

[0061] Update target network: Update the target network of Actor and Critic. The parameters of the target network are the parameters of the initial network of Actor and Critic a long time ago. They are fixed for a period of time and then replaced by the parameters of the estimated network. The advantage of this is that it can speed up the convergence of the estimated network parameters.

[0062] Repeat the above steps until the model converges.

[0063] Step 4: Dynamic adjustment: Based on network conditions and user feedback, the model trained by reinforcement learning is used to dynamically adjust the bandwidth allocation strategy. An adaptive adjustment mechanism is implemented to cope with sudden traffic and network changes.

[0064] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention without departing from the principles and intent of the present invention.

Claims

1. A heterogeneous network bandwidth adaptive allocation method based on reinforcement learning, characterized in that: The following steps are involved: Step 1: Data collection: Collect bandwidth usage information of various communication devices and bandwidth requirements of different types of services on the terminal through user terminals; Step 2: Demand analysis and modeling: Classify and analyze the collected user demand data, identify the bandwidth demand characteristics of different service types, perform mathematical modeling for different service types, and obtain the relationship between service flow evaluation and communication bandwidth; Step 3: Optimize matching: Evaluate the bandwidth resource status of the current heterogeneous network, use reinforcement learning algorithms to match business needs with bandwidth resources, and make decisions on bandwidth needs in the future to obtain the best business service quality; Step 4: Dynamic adjustment: Dynamically adjust the bandwidth allocation strategy based on the reinforcement learning model.

2. The method for adaptively allocating heterogeneous network bandwidth based on reinforcement learning according to claim 1, characterized in that: In step 2, mainly for different service types, the relationship between the service flow evaluation degree U(B) and the communication bandwidth B is: For emergency voice services, the service flow evaluation function is expressed as: Where, U(B) is the service flow evaluation degree and B is the user communication bandwidth; th is the threshold value and does not change with time; For emergency data file services, the service flow evaluation function is expressed as: In the formula, B max It is the bandwidth obtained by reaching the maximum service evaluation value, which does not change with time, and e is the base of the natural logarithm; (3) For the emergency video stream service type, the service flow evaluation function is expressed as: Where C is the average value of the basic bandwidth requirement and the maximum bandwidth requirement of the video streaming service, which does not change over time.

3. The method for adaptively allocating heterogeneous network bandwidth based on reinforcement learning according to claim 1, characterized in that: In step 3, the bandwidth resource allocation problem is transformed into a simple addition of the evaluation utilities of each service flow, and the optimization objective function is: In the formula, It is expressed as the sum of the communication bandwidth B allocated to the nth service by P types of network communication means at time t, and the service flow evaluation degree is obtained; n is the service, p is the type of communication means and t is the time; N is the total number of services, P is the total number of network communication means types, and T is the total time.

4. The method for adaptively allocating heterogeneous network bandwidth based on reinforcement learning according to claim 3 is characterized in that: In step 3, the final optimization goal is established as: That is, the maximum value R of all business value evaluations within a period of time T total , returns r t is the business value evaluation value at time t.

5. The method for adaptively allocating heterogeneous network bandwidth based on reinforcement learning according to claim 4 is characterized in that: In step 3, resource allocation uses a reinforcement learning algorithm based on the Actor-Critic architecture to construct the initial neural networks of the Actor network and the Critic network; based on the input state dimension of the neural network, the number of layers and the number of neural networks are determined, and at the same time, the target neural networks of the Actor network and the Critic are constructed.

6. The method for adaptively allocating heterogeneous network bandwidth based on reinforcement learning according to claim 5, characterized in that: In step 3, the neural network training process is: In each round, the user sets the state space s of the initial state t Input into the Actor initial neural network to get the current action state a t , after executing the action, we get the next state space s t+1 , and the current reward r t ; The user stores the state transitions he experiences in the replay pool, i.e., the four-tuple sample (s t ,a t ,s t+1 ,r t ), and start training when there is enough data in the replay pool; During training, sample batches from the replay pool: randomly extract a batch of data from the replay pool; Update the initial Critic network: use batch data to calculate the target Q value and update the parameters of the Critic network; Update the Actor initial network: Use the policy gradient method to update the parameters of the Actor network; Update target network: Update the target network of Actor and Critic. The parameters of the target network are the parameters of the Actor initial network and the Critic initial network. They are fixed for a period of time and then replaced by the parameters of the estimated network. Repeat the above steps until the model converges.

7. The method for adaptively allocating heterogeneous network bandwidth based on reinforcement learning according to claim 6, characterized in that: Update the initial Critic network: The target Q value is calculated by the Bellman equation, that is, minimizing the TD error; Loss(θ Q )=[r t +λQ(s t+1 ,a t+1 |θ Q′ )-Q(s t ,a t |θ Q )] 2 (5) In the formula, Loss is the loss function, which represents the difference between the value output by the Critic initial network and the expected target value. This value is used to update the Critic initial network θ Q , λ is the value of the Critic initial network output; Q(s t+1 ,a t+1 |θ Q′ ) is the value output by the Critic target network.

Citation Information

Cited By

  • Heterogeneous network bandwidth allocation method based on path optimization

    CN122226619A