On-off-Policy deep reinforcement learning algorithm based on optimal communication resource scheduling strategy

By adopting the on-off-policy deep reinforcement learning algorithm in communication resource scheduling, combining the advantages of on-policy and off-policy, using the monotonic characteristics of the value function and the dynamic priority management of the experience pool, the problems of waste of resources, inefficiency and poor adaptability in the existing scheduling methods are solved, and efficient resource scheduling in dynamic and complex environments are achieved.

CN120091447APending Publication Date: 2025-06-03陈嘉铮
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510131469.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing communication resource scheduling methods face the problems of resource waste, inefficient scheduling efficiency and poor adaptability to environmental changes, and it is difficult to fully consider the dynamic changes in communication needs, resulting in poor results in complex environments.

Method used

The on-off-policy deep reinforcement learning algorithm based on the optimal communication resource scheduling strategy is adopted. By combining the advantages of on-policy and off-policy deep reinforcement learning, the monotonic characteristics of the value function and the dynamic priority management mechanism of the experience pool are used to achieve rapid convergence of the strategy and global optimal performance.

Benefits of technology

This algorithm can achieve efficient adaptation of resource scheduling in dynamic and complex environments, reduce computational resource waste, improve data utilization, ensure the real-time nature of scheduling strategies, and maintain efficient scheduling performance in large-scale systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120091447A_ABST
    Figure CN120091447A_ABST
Patent Text Reader

Abstract

The invention provides an on-off-policy deep reinforcement learning algorithm based on an optimal communication resource scheduling strategy. The on-off-policiy deep reinforcement learning algorithm based on the optimal communication resource scheduling strategy comprises the steps of S1, collecting state data at a moment by the algorithm for a wireless networked control system with a sensor and a channel, and predicting and updating an equipment state by using Kalman filtering, and S2, calculating a resource allocation action vector based on the collected state data. According to the on-off-policy deep reinforcement learning algorithm based on the optimal communication resource scheduling strategy, the advantages of on-policy and off-policy deep reinforcement learning are combined, and meanwhile, the monotonous characteristic of a value function and the dynamic priority management mechanism of an experience pool are utilized, so that the rapid convergence and global optimal performance of the strategy are realized; the method shows excellent application value and wide applicability in a dynamic complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of on-off-policy deep reinforcement learning algorithms, and specifically to an on-off-policy deep reinforcement learning algorithm based on an optimal communication resource scheduling strategy. Background Art

[0002] The on-off-policy deep reinforcement learning algorithm based on an optimal communication resource scheduling strategy mainly optimizes the scheduling of communication resources through deep reinforcement learning methods. The core idea of this algorithm is to model the communication resource scheduling problem as a reinforcement learning task, where the agent, during the interaction with the environment, learns the optimal resource scheduling strategy through continuous trial and error and feedback. Structurally, the algorithm adopts a two-stage deep reinforcement learning architecture: On-Policy stage: In this stage, the agent interacts with the environment based on the current policy, generating a set of data for training. This process is an update based on the current policy, ensuring that the policy minimizes losses at each decision step. Off-Policy stage: In this stage, the agent no longer fully relies on the current policy but uses historical data for training to learn a more optimized policy. This enables the algorithm to learn from past experiences and improve the flexibility and efficiency of resource scheduling.

[0003] Traditional communication resource scheduling methods face multiple challenges, the most common of which include problems such as resource waste, low scheduling efficiency, and poor adaptability to environmental changes. Existing scheduling strategies usually rely on rules or heuristic algorithms, and these methods often fail to fully consider the dynamic changes in communication requirements, resulting in poor performance in complex environments. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the present invention provides an on-off-policy deep reinforcement learning algorithm based on an optimal communication resource scheduling strategy, which solves problems such as resource waste, low scheduling efficiency, and poor adaptability to environmental changes; and the problem that it cannot fully consider the dynamic changes in communication requirements, resulting in poor performance in complex environments.

[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: An on-off-policy deep reinforcement learning algorithm based on an optimal communication resource scheduling strategy, comprising:

[0006] S1: For a wireless networked control system with N sensors and M channels, the algorithm collects state data at time t including the following:

[0007] Data of the edge device: Information freshness

[0008]

[0009] where τ n,t is defined as:

[0010]

[0011] Communication resource data: channel transmission status

[0012]

[0013] where h n,m,t represents the communication status of device n and channel m at time t, and combines with the transmission success rate p n,m,t to calculate the communication quality;

[0014] Use Kalman filter to predict and update the device status. The formula is as follows:

[0015] State prediction:

[0016]

[0017] where A n is the device state transition matrix, and W n is the noise matrix;

[0018] Kalman gain:

[0019]

[0020] where C n is the observation state transition matrix, and V n is the observation noise covariance matrix;

[0021] State update:

[0022]

[0023] where y n,t is the observed value;

[0024] S2: Based on the collected state data, calculate the resource allocation action vector is defined as follows:

[0025]

[0026] Action needs to satisfy the following constraint conditions:

[0027]

[0028] Resource allocation is achieved by optimizing the following objective function:

[0029]

[0030] where π is the scheduling policy and γ is the discount factor; the immediate cost function is defined as:

[0031]

[0032] P n,t is the error covariance matrix;

[0033] To adapt to the energy consumption sensitive scenario, the cost function is extended to:

[0034]

[0035] where e n,t is the device energy consumption;

[0036] S3: Generate a new data trajectory based on the current policy and design a loss function with monotonicity:

[0037]

[0038] S4: Store the data that meets the following conditions in the experience pool:

[0039]

[0040] and screen the samples with higher value gain through the dynamic priority mechanism;

[0041] Select data from the experience pool and calculate the off-policy loss function:

[0042]

[0043] S5: Combine the on-policy and off-policy loss functions and use the Adam optimizer to update the model parameters:

[0044]

[0045] where g t is the gradient and α is the learning rate.

[0046] Preferably, the λ 1 and λ 2 are weight parameters respectively, which can be dynamically adjusted according to the task priority. Tr(P n,t ) represents the trace of the error covariance matrix, which is used to measure the transmission accuracy, and e n,t is the device energy consumption, which is used to evaluate the energy consumption overhead of the system.

[0047] Preferably, the experience pool samples are defined as:

[0048]

[0049] Only retain samples where ΔV > ΔV threshold For the samples, dynamically adjust the priority queue of the experience pool so that samples with greater value gain have a higher probability of being selected. The number of samples in the experience pool changes dynamically, and the sample storage upper limit N is adjusted according to the complexity of the current task max This dependent claim realizes a dynamic management mechanism for the data in the experience pool by introducing a value gain criterion. Compared with the traditional random sampling or fixed storage method, it greatly improves the effectiveness of the training data and the training efficiency. Especially in large-scale dynamic scenarios, this mechanism can automatically preferentially store and use important samples, reduce data redundancy, and at the same time improve the model convergence speed and policy performance

[0050] Preferably, the load of the communication adjusts the sampling frequency, which is defined as:

[0051]

[0052] where f base is the reference sampling frequency, and the current load value is the load adjustment factor. In the case of low load, the sampling frequency is reduced to reduce the consumption of computing resources; in the case of high load, the sampling frequency is increased to ensure the real-time performance of the scheduling strategy. By dynamically adjusting the sampling frequency, this dependent claim significantly reduces the overhead of computing resources while ensuring the scheduling performance, and is particularly suitable for high-dynamic communication environments. Compared with the traditional method with a fixed sampling frequency, the present invention achieves a dynamic balance between performance and efficiency, providing an efficient and feasible solution for resource scheduling in complex scenarios

[0053] Preferably, the framework of the distributed collaborative training distributes the training tasks of the large-scale scheduling problem to multiple edge nodes, and each node independently calculates the local gradient:

[0054]

[0055] where k is the edge node number

[0056] Preferably, the local gradient of the node is uploaded to the global parameter server to calculate the global gradient:

[0057]

[0058] where K is the total number of edge nodes

[0059] Preferably, the global gradient updates the model parameters:

[0060] θ t = θ t-1 - αg t

[0061] This dependent claim first applies distributed collaborative training to the problem of wireless network resource scheduling. Through the collaborative cooperation of edge computing nodes, the computational load of a single node is significantly reduced, and at the same time, the training efficiency of the model is accelerated. This framework can make full use of the computing power of edge nodes while maintaining global consistency, providing an efficient distributed optimization solution for large-scale complex systems.

[0062] Preferably, for the hybrid training method of on-policy and off-policy advantages, where: design an on-policy loss function with a monotonically increasing property:

[0063]

[0064] Ensure that the value function of the current policy has good monotonicity. In the off-policy part, select samples from the experience pool and calculate the off-policy loss function:

[0065]

[0066] The present invention provides an on-off-policy deep reinforcement learning algorithm based on an optimal communication resource scheduling strategy. It has the following beneficial effects:

[0067] This on-off-policy deep reinforcement learning algorithm based on the optimal communication resource scheduling strategy combines the advantages of on-policy and off-policy deep reinforcement learning, and at the same time uses the monotonic property of the value function and the dynamic priority management mechanism of the experience pool to achieve fast convergence of the policy and global optimal performance, showing excellent application value and wide applicability in dynamic complex environments. By combining the state estimation of Kalman filtering and the cost function of multi-objective optimization, this solution can efficiently adapt to changing communication requirements and meet the dual optimization goals of accuracy and energy consumption. The introduction of intelligent experience pool management and dynamic sampling strategies enables this solution to maintain efficient scheduling performance in large-scale systems, reducing waste of computing resources and improving data utilization. Especially in high-dynamic environments, the dynamic sampling mechanism of this solution can quickly respond to load changes to ensure the real-time nature of the scheduling strategy.

[0068] The design of the distributed collaborative training framework enables this technical solution to exhibit significant advantages in the resource scheduling of large-scale systems. Through the collaborative cooperation of edge computing nodes, the training efficiency is greatly improved and the computing load of the central node is reduced, which is applicable to scenarios such as the Internet of Things, smart cities, and 5G networks that require large-scale distributed optimization. Finally, the hybrid training method of on-policy and off-policy combined with the design of a monotonically increasing loss function provides a faster convergence speed and higher adaptability for policy optimization, and can generate efficient and accurate scheduling policies in complex task environments, further improving the performance and robustness of wireless networks. Brief Description of the Drawings

[0069] Figure 1 It is a flowchart of the feature-driven hybrid policy deep reinforcement learning algorithm based on the present invention.

[0070] Figure 2 It is a comparison chart of the average state prediction errors between the traditional policy and the scheduling policy of the present invention during the training of the present invention (N = 20, M = 10). Detailed Description of the Invention

[0071] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0072] As Figure 1-2 shown, the embodiment of the present invention provides an on-off-policy deep reinforcement learning algorithm based on an optimal communication resource scheduling policy, including: S1: For a wireless networked control system with N sensors and M channels, the algorithm collects state data at time t including the following: data of edge devices: information freshness

[0073]

[0074] where τ n,t is defined as:

[0075]

[0076] communication resource data: channel transmission state

[0077]

[0078] where h n,m,t represents the communication state of device n and channel m at time t, combined with the transmission success rate p n,m,tCalculate communication quality;

[0079] Use Kalman filtering to predict and update the device state. The formula is as follows:

[0080] State prediction:

[0081]

[0082] where A n is the device state transition matrix, and W n is the noise matrix;

[0083] Kalman gain:

[0084]

[0085] where C n is the observation state transition matrix, and V n is the observation noise covariance matrix;

[0086] State update:

[0087]

[0088] where y n,t is the observed value;

[0089] S2: Based on the collected state data, calculate the resource allocation action vector which is defined as follows:

[0090]

[0091] The action needs to satisfy the following constraint conditions:

[0092]

[0093] Resource allocation is achieved by optimizing the following objective function:

[0094]

[0095] where π is the scheduling policy and γ is the discount factor; the immediate cost function is defined as:

[0096]

[0097] P n,t is the error covariance matrix;

[0098] To adapt to the energy consumption sensitive scenario, the cost function is extended to:

[0099]

[0100] where e n,t is the device energy consumption, and λ 1 and λ 2 are weight parameters respectively, which can be dynamically adjusted according to the task priority. Tr(P n,t ) represents the trace of the error covariance matrix, which is used to measure the transmission accuracy. e n,t is the device energy consumption, which is used to evaluate the energy consumption overhead of the system;

[0101] S3: Generate a new data trajectory based on the current policy and design a loss function with monotonicity:

[0102]

[0103] S4: Store the data that meets the following conditions in the experience pool:

[0104]

[0105] And screen out samples with higher value gains through a dynamic priority mechanism;

[0106] Select data from the experience pool and calculate the loss function of off-policy:

[0107]

[0108] Samples in the experience pool are defined as:

[0109]

[0110] Only retain samples with ΔV > ΔV threshold . Dynamically adjust the priority queue of the experience pool so that samples with greater value gains have a higher probability of being selected. The number of samples in the experience pool changes dynamically, and the upper limit N of sample storage is adjusted according to the complexity of the current task max . This dependent claim realizes a dynamic management mechanism for the data in the experience pool by introducing a value gain criterion. Compared with traditional random sampling or fixed storage methods, it greatly improves the effectiveness and training efficiency of training data. Especially in large-scale dynamic scenarios, this mechanism can automatically preferentially store and use important samples, reduce data redundancy, and at the same time improve the model convergence speed and policy performance. Adjust the sampling frequency of the communication load, which is defined as:

[0111]

[0112] where f baseTaking the reference sampling frequency, and the previous value of the load as the load adjustment factor. In the case of low load, the sampling frequency is reduced to reduce the consumption of computing resources; in the case of high load, the sampling frequency is increased to ensure the real-time performance of the scheduling strategy. By dynamically adjusting the sampling frequency, this dependent claim significantly reduces the overhead of computing resources while ensuring the scheduling performance, and is particularly suitable for highly dynamic communication environments. Compared with the traditional method with a fixed sampling frequency, the present invention achieves a dynamic balance between performance and efficiency, providing an efficient and feasible solution for resource scheduling in complex scenarios.

[0113] S5: Combining the loss functions of on-policy and off-policy, use the Adam optimizer to update the model parameters:

[0114]

[0115] where g t is the gradient, α is the learning rate. The framework of distributed collaborative training distributes the training tasks of large-scale scheduling problems to multiple edge nodes, and each node independently calculates the local gradient:

[0116]

[0117] where k is the edge node number. The local gradient of the node is uploaded to the global parameter server to calculate the global gradient:

[0118]

[0119] where K is the total number of edge nodes. Update the model parameters of the global gradient:

[0120] θ t = θ t-1 - αg t

[0121] Through the collaborative cooperation of edge computing nodes, the computing load of a single node is significantly reduced, and at the same time, the training efficiency of the model is accelerated. This framework can fully utilize the computing power of edge nodes on the basis of maintaining global consistency, providing an efficient distributed optimization solution for large-scale complex systems. A hybrid training method with the advantages of on-policy and off-policy, where: design an on-policy loss function with a monotonically increasing property:

[0122]

[0123] Ensure that the value function of the current policy has good monotonicity. In the off-policy part, samples are selected from the experience pool to calculate the off-policy loss function:

[0124]

[0125] Specific cases and simulation results:

[0126] This algorithm can be used in scenarios such as industrial automation, network robots, and autonomous driving, where sensors are installed on corresponding devices. In this patent, we used Python to conduct simulation experiments on industrial automation. In the experimental scenario, we set that the data of 20 devices needs to be transmitted, and there are a total of 10 channels available for transmission. Among them, the state transition matrix A of each device n and the observation state transition matrix C n are randomly generated, and the noise matrices V n and W n are identity matrices. The channel is set as a Rayleigh fading channel, following the Rayleigh distribution. The experimental results are shown in Figure 2 .

[0127] Figure 2 shows the performance during the training process of the traditional scheduling strategy and this strategy. We can see that the scheduling scheme generated by this on-off-policy deep reinforcement learning converges about 40% faster than the traditional method, and the average error of the state prediction after convergence is about 35% lower.

[0128] Partial code:

[0129]

[0130]

[0131]

[0132]

[0133]

[0134]

[0135]

[0136]

[0137]

[0138] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An on-off-policy deep reinforcement learning algorithm based on optimal communication resource scheduling strategy, characterized in that: include: S1: For a wireless networked control system with N sensors and M channels, the algorithm collects state data at time t Includes the following: Data from edge devices: Information freshness where τ n,t Defined as: Communication resource data: channel transmission status where h n,m,t represents the communication status of device n and channel m at time t, combined with the transmission success rate p n,m,t Calculate communication quality; Use Kalman filtering to predict and update the device status. The formula is as follows: Status prediction: Among them A n is the device state transfer matrix, W n is the noise matrix; Kalman gain: Among them C n is the observation state transfer matrix, V n is the observation noise covariance matrix; Status Update: where y n,t is the observed value; S2: Based on the collected state data, calculate the resource allocation action vector The definition is as follows: action The following constraints need to be met: Resource allocation is achieved by optimizing the following objective function: Where π is the scheduling strategy, γ is the discount factor; the instant cost function Defined as: P n,t is the error covariance matrix; To adapt to energy-sensitive scenarios, the cost function is expanded to: where e n,t is the energy consumption of the equipment; S3: Generate new data trajectories based on the current strategy and design a monotonic loss function: S4: Store the data that meets the following conditions into the experience pool: And screen samples with higher value gains through a dynamic priority mechanism; Select data from the experience pool and calculate the off-policy loss function: S5: Combine the on-policy and off-policy loss functions and use the Adam optimizer to update the model parameters: where g t is the gradient and α is the learning rate.

2. According to claim 1, an on-off-policy deep reinforcement learning algorithm based on an optimal communication resource scheduling strategy is characterized in that: The λ1 and λ2 are weight parameters, which can be adjusted dynamically according to the task priority. n,t ) represents the trace of the error covariance matrix, which is used to measure the transmission accuracy, e n,t It is the energy consumption of the device, which is used to evaluate the energy consumption of the system.

3. The on-off-policy deep reinforcement learning algorithm based on the optimal communication resource scheduling strategy according to claim 1, characterized in that: The experience pool sample is defined as: Only keep ΔV>ΔV threshold of the sample.

4. The on-off-policy deep reinforcement learning algorithm based on the optimal communication resource scheduling strategy according to claim 1, characterized in that: The load-adjusted sampling frequency of the communication is defined as: where f base is the reference sampling frequency, and the current value of the load is the load adjustment factor.

5. The on-off-policy deep reinforcement learning algorithm based on the optimal communication resource scheduling strategy according to claim 1, characterized in that: The distributed collaborative training framework distributes the training tasks of large-scale scheduling problems to multiple edge nodes, and each node independently calculates the local gradient: Where k is the edge node number.

6. The on-off-policy deep reinforcement learning algorithm based on the optimal communication resource scheduling strategy according to claim 5, characterized in that: The local gradient of the node is uploaded to the global parameter server, and the global gradient is calculated: Where K is the total number of edge nodes.

7. The on-off-policy deep reinforcement learning algorithm based on the optimal communication resource scheduling strategy according to claim 6, characterized in that: The global gradient updates the model parameters: i t =θ t-1 -ag t。 8. The on-off-policy deep reinforcement learning algorithm based on the optimal communication resource scheduling strategy according to claim 1, characterized in that: The hybrid training method of on-policy and off-policy advantages, wherein: an on-policy loss function with a monotonically increasing characteristic is designed: To ensure that the value function of the current strategy has good monotonicity, in the off-policy part, samples are selected from the experience pool and the off-policy loss function is calculated: