Deep reinforcement learning multi-task resource allocation method based on strategy gradient

Through a deep reinforcement learning method based on strategy gradients, a resource allocation intelligent model is constructed, which solves the problem that existing technology is difficult to effectively allocate resources when the environment and resource situations change, and achieves dynamic adjustment of resource allocation decisions in new environments and states to achieve the effect of maximizing returns.

CN120073768APending Publication Date: 2025-05-30THE FIFTH RES INST OF TELECOMM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510152665.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing technology is difficult to achieve the expected results by directly using historical experience training in the environment where the target is different and the current resource conditions.

Method used

The deep reinforcement learning method based on policy gradient is adopted to build a resource allocation agent model, and the agent's policies are updated through policy gradient training, so that they can dynamically adjust resource allocation decisions in a new environment and state to maximize returns.

Benefits of technology

In the new environment and state, resource allocation decisions can still be dynamically adjusted to maximize returns, solving the problem that existing technology is difficult to effectively allocate resources when environmental and resource situations change.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120073768A_ABST
    Figure CN120073768A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-task resource allocation method based on deep reinforcement learning of a strategy gradient, and the method comprises the steps: carrying out the intelligent agent training and perception of a resource allocation decision under the environment and state of a target history based on the deep reinforcement learning of the strategy gradient; according to the strategy gradient algorithm, an intelligent agent directly interacts with the environment according to a current strategy, the gradient of strategy parameters is directly calculated through track data obtained through sampling, and then the current strategy is updated, so that the current strategy is close to a target expected to be reported by the maximum strategy; in a strategy gradient algorithm, a Monte Carlo method is adopted to sample a track to estimate an action value, and an unbiased gradient can be obtained, so that the strategy can still dynamically adjust a resource allocation decision in a new environment and state, and the benefit maximization is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of task resource allocation, and in particular, to a multi-task resource allocation method for deep reinforcement learning based on policy gradient. Background Art

[0002] With the continuous development of information technology, task resource allocation methods are widely used in many fields such as communication, cloud distributed storage, edge computing, network traffic control, and energy management. In order to make full use of each resource and meet the continuous monitoring of the target, it is necessary to timely adjust the resource allocation situation of the target under different states.

[0003] Target task resource allocation includes resource allocation methods based on dynamic programming, which are divided into policy iteration and value iteration. Policy iteration consists of two parts: policy evaluation and policy improvement. The policy evaluation in policy iteration obtains the state value function of a policy through the Bellman expectation equation; while value iteration directly performs dynamic programming through the Bellman optimal equation to obtain the final optimal state value. These two resource allocation algorithms based on dynamic programming need to know the state transition function and reward function of the environment, that is, the entire Markov decision process needs to be known. In a white-box environment, the state value function can be directly solved by dynamic programming without the need to learn through a large number of interactions between the agent and the environment. However, there are few white-box environments in reality, which is also the limitation of the dynamic programming algorithm, and we cannot apply it to many actual scenarios; in addition, policy iteration and value iteration are only applicable to finite Markov decision processes, that is, the state space and action space are discrete and finite; due to the different environments where the target is located and the current resource situation, it is difficult to achieve the expected effect by directly using historical experience for training. Summary of the Invention

[0004] The main object of the present invention is to provide a multi-task resource allocation method for deep reinforcement learning based on policy gradient, aiming to solve the problem that it is difficult to achieve the expected effect by directly using historical experience for training due to the different environments where the target is located and the current resource situation.

[0005] To achieve the above object, the present invention proposes a multi-task resource allocation method for deep reinforcement learning based on policy gradient, and the multi-task resource allocation method for deep reinforcement learning based on policy gradient includes:

[0006] Construct a database of the motion trajectory data of a moving target and the monitoring results of sensor device resources, and preprocess the database of the motion trajectory data of the moving target and the monitoring results of sensor device resources to obtain a target trajectory dataset;

[0007] Train an action set according to the preprocessed monitoring result data of sensor device resources;

[0008] Perform policy gradient training on the target trajectory dataset through a deep reinforcement learning model, construct a resource allocation agent model, and obtain a training weight file;

[0009] Normalize the multi-target data state, and predict the target state data based on the resource allocation agent model and the training weight file to obtain a resource allocation result that maximizes the expectation.

[0010] In one embodiment, the specific steps of the policy gradient training are as follows: The agent randomly selects an action from the action set, gives different rewards, and uses a policy gradient training update method to approach the goal of maximizing the policy expected reward. When the reward stabilizes within a certain range, the training stops and a training weight file is obtained.

[0011] In one embodiment, the specific steps of constructing a database of the motion trajectory data of a moving target and the resource monitoring results of sensor devices are as follows: According to the motion trajectory of the moving target and the resource monitoring situation of the sensor devices, form a database of the motion trajectory data of the moving target and the resource monitoring results of the sensor devices.

[0012] In one embodiment, the specific steps of preprocessing the database of the motion trajectory data of the moving target and the resource monitoring results of sensor devices are as follows: Analyze and clean the database of the motion trajectory data of the moving target and the resource monitoring results of sensor devices.

[0013] The technical solution of the present invention conducts intelligent agent training and perception on resource allocation decisions in the environment and state where the target is located based on deep reinforcement learning with policy gradients. In the policy gradient algorithm, the agent directly interacts with the environment according to the current policy, directly calculates the gradient of the policy parameters through the trajectory data obtained by sampling, and then updates the current policy to approach the goal of maximizing the policy expected reward; the Monte Carlo method is used in the policy gradient algorithm to sample trajectories to estimate the action value, and an unbiased gradient can be obtained, enabling it to dynamically adjust resource allocation decisions in new environments and states to achieve maximum benefits. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 is a schematic flowchart of the multi-task resource allocation method based on deep reinforcement learning with policy gradients of the present invention;

[0015] Figure 2 is a schematic overall framework diagram of the multi-task resource allocation method based on deep reinforcement learning with policy gradients of the present invention;

[0016] Figure 3 is a schematic diagram of the interaction between a single agent and the environment of the multi-task resource allocation method based on deep reinforcement learning with policy gradients of the present invention;

[0017] Figure 4 This is a schematic diagram of the overall application and network model of the multi-task resource allocation method based on policy gradient deep reinforcement learning of the present invention. Specific Embodiments

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Generally, the components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.

[0019] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0020] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0021] In the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "upper", "lower", "inner", "outer", "left", "right", etc. are based on the orientation or positional relationships shown in the accompanying drawings, or the orientation or positional relationships in which the product of the present invention is customarily placed during use, or the orientation or positional relationships commonly understood by those skilled in the art. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention.

[0022] In addition, the terms "first", "second", etc. are only used for distinguishing descriptions and should not be construed as indicating or implying relative importance.

[0023] In the description of the present invention, it should also be noted that unless otherwise clearly defined and limited, terms such as "set", "connected" should be understood in a broad sense. For example, "connected" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the internal communication of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0024] Reinforcement learning is a computational method by which a machine achieves a goal by interacting with the environment; the machine makes an action decision in a state of the environment, applies this action to the environment, and the environment undergoes corresponding changes and transmits the corresponding reward feedback and the next state back to the machine; the goal of the machine is to maximize the expected cumulative reward obtained during multiple rounds of interaction; multi-agent reinforcement learning is when multiple agents interact in the same environment to maximize the overall benefit. Due to the differences in the existing environment where the goal is located and the current resource situation, such as high state dimensions, huge action sets, high computational complexity, difficult reward function design, and insufficient real-time performance and robustness, directly using historical experience for training often fails to achieve the expected results.

[0025] To solve the above problems, the present invention proposes a deep reinforcement learning multi-task resource allocation method based on policy gradient. The following will combine the accompanying drawings to elaborate on the specific implementation manners of the present invention in detail.

[0026] As Figures 1-4 shown, the deep reinforcement learning multi-task resource allocation method based on policy gradient includes the following steps:

[0027] Construct a database of the motion trajectory data of the moving target and the monitoring results of the sensor device resources, and preprocess the database of the motion trajectory data of the moving target and the monitoring results of the sensor device resources to obtain a target trajectory dataset;

[0028] Train an action set according to the preprocessed monitoring result data of the sensor device resources;

[0029] Perform policy gradient training on the target trajectory dataset through a deep reinforcement learning model, construct a resource allocation agent model, and obtain a training weight file;

[0030] Normalize the multi-target data state, and predict the target state data based on the resource allocation agent model and the training weight file to obtain a resource allocation result that maximizes the expectation.

[0031] In one embodiment, the specific steps of the policy gradient training are as follows: the agent randomly selects an action from the action set, gives different rewards, and uses a policy gradient training update method to approach the goal of maximizing the policy expected reward. When the reward stabilizes within a certain range, the training stops, and a training weight file is obtained.

[0032] In one embodiment, the specific steps of constructing the database of the motion trajectory data of the moving target and the monitoring results of the sensor device resources are as follows: according to the motion trajectory of the moving target and the resource monitoring situation of the sensor device, form a database of the motion trajectory data of the moving target and the monitoring results of the sensor device resources.

[0033] In one embodiment, the specific steps for preprocessing the motion trajectory data of the moving target and the sensor device resource monitoring result database are as follows: analyze and clean the motion trajectory data of the moving target and the sensor device resource monitoring result database.

[0034] In this embodiment, Figure 2 On the left side is model training, that is, first collect, analyze and clean historical data to make a data set that meets application requirements; then train an agent for monitoring data analysis and decision-making through a designed reinforcement learning algorithm to build a resource allocation agent model library; Figure 2 On the right side is intelligent decision-making and resource planning, that is, through the trained agent, make decisions based on historical data and the current environment, and perform intelligent planning of sensor device resources, improve the utilization rate of sensor device resources, and maximize the benefits of monitoring resource allocation for each target under multiple tasks.

[0035] As Figure 3 shown, in this embodiment, a task of one target type is designed as an agent. The resource allocation result of each task is used as the action for interacting with the environment, the receiving accuracy of the target after resource allocation is used as the reward of the action, and the transfer of time and position states is used as the next state to construct the basic structure of agent interaction.

[0036] As Figure 4 shown, in this embodiment, first clean the historical target state data, then perform structured modeling on the data, and use the utilization result of the resources as the output action of the training to obtain training data; secondly, perform model training, the training input size is 1*2, then perform a fully connected layer of 1*16384 to obtain the output of the hidden layer and perform a relu activation operation, and then map the result to the final action dimension size, and output through the softmax layer to obtain the probability distribution of each action. Finally, adopt a stochastic strategy to maximize the sampling of this distribution to obtain the final action.

[0037] The multi-task resource allocation method based on deep reinforcement learning with policy gradient of the present invention performs agent training and perception on resource allocation decisions in the environment and state where the target is located based on deep reinforcement learning with policy gradient. In the policy gradient algorithm, the agent directly interacts with the environment according to the current policy, and directly calculates the gradient of the policy parameters through the trajectory data obtained by sampling, and then updates the current policy to make it approach the goal of maximizing the expected return of the policy; the Monte Carlo method is used to sample trajectories in the policy gradient algorithm to estimate the action value, and an unbiased gradient can be obtained, so that it can still dynamically adjust the resource allocation decision in the new environment and state to achieve the maximum benefit.

[0038] Taking a certain task as an example, the representation of the target agent modeling is as follows:

[0039] : Current state [Current time , current position ;

[0040] : Action for interacting with the environment in the current state;

[0041] : Reward under the current action;

[0042] : Current target policy adopted;

[0043] : Cumulative return under the current policy;

[0044] Among them, is to obtain the range of sensor device resources [A, B, C,... N] according to the cleaned data, and obtain its resource combinations as the action set of the current task, such as [A, B], [A, C], [A, B, C], etc. One combination is used as an action, so several actions can be obtained, that is, the optional actions of the agent for the environment; represents the return of the current agent, ; represents the distance between the device and the target; represents the angle between the device and the target; The policy of the agent is , The policy is a function, indicating the probability of taking an action in the case of the input state. The present invention adopts a stochastic policy, and outputs a probability distribution of actions for each state of the target. Then, an action can be obtained by sampling according to this distribution. The ultimate optimization goal of the reinforcement learning task is to maximize the return of the agent's optimal resource allocation strategy in the process of interacting with the dynamic environment for each target state.

[0045] The policy-based method first needs to parameterize the policy: Assume the target policy is a stochastic policy and is differentiable everywhere, where is the corresponding parameter; Use a neural network model to model such a policy function, input a certain state, and then output a probability distribution of an action; The goal of the present invention is to find an optimal policy and maximize the expected return of this policy in the environment, that is, to define the objective function of policy learning as: ; Among them, represents the initial state; After taking the derivative of the objective function with respect to the policy and obtaining the derivative, the gradient ascent method can be used to maximize this objective function, thereby obtaining the optimal policy.

[0046] Taking the gradient of the objective function, we get: , and this gradient is used to update the policy. At each state, the modification of the gradient is to make the policy sample more actions that bring higher values and fewer actions that bring lower values. In the formula for calculating the policy gradient, is estimated using the Monte Carlo method. For an environment with a finite number of steps, its policy gradient is , where T is the maximum number of steps of interacting with the environment, that is, the farthest moment of the time evolution of the target task data. The specific algorithm solving steps are as follows:

[0047] Initialize the policy parameters ;

[0048] for episode e = 1 → E do:

[0049] Use the current policy to sample a trajectory ;

[0050] Calculate the return from each moment of the current trajectory onwards , denoted as ;

[0051] Update , .

[0052] In summary, deep reinforcement learning lies in combining the perception ability of deep learning with the decision-making ability of reinforcement learning, which can directly control according to the task input and is an artificial intelligence method closer to the human thinking mode. On the basis of the combination of the two, the present invention proposes a deep reinforcement learning multi-task resource allocation method based on policy gradient. This method first trains an agent in an infinite state space and action space according to the resource decision scheme of the target in history, and then perceives according to the environment where the current target is located to complete the dynamic adjustment of resources. Finally, the resource allocation in the current state is selected to maximize the benefit of the target.

[0053] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A multi-task resource allocation method for deep reinforcement learning based on policy gradient, characterized in that: The multi-task resource allocation method for deep reinforcement learning based on policy gradient includes the following steps: Constructing a database of motion trajectory data of a motion target and a sensor device resource monitoring result database, and preprocessing the motion trajectory data of the motion target and the sensor device resource monitoring result database to obtain a target trajectory data set; Training the action set based on the preprocessed sensor device resource monitoring result data; Performing policy gradient training on the target trajectory dataset through a deep reinforcement learning model, building a resource allocation agent model, and obtaining a training weight file; The multi-objective data states are normalized, and the target state data are predicted based on the resource allocation agent model and the training weight file to obtain a resource allocation result that maximizes expectations.

2. The multi-task resource allocation method for deep reinforcement learning based on policy gradient according to claim 1, characterized in that: The specific steps of the policy gradient training are: the agent randomly selects actions in the action set, gives different rewards, adopts the policy gradient training update method to approach the goal of maximizing the expected return of the strategy, and stops the training when the reward stabilizes within a certain range to obtain the training weight file.

3. The multi-task resource allocation method for deep reinforcement learning based on policy gradient according to claim 1, characterized in that: The specific steps of constructing the motion trajectory data of the motion target and the sensor equipment resource monitoring result database are: forming the motion trajectory data of the motion target and the sensor equipment resource monitoring result database according to the motion trajectory of the motion target and the resource monitoring situation of the sensor equipment.

4. The multi-task resource allocation method for deep reinforcement learning based on policy gradient according to claim 1, characterized in that: The specific steps of preprocessing the motion trajectory data of the motion target and the sensor device resource monitoring result database are: analyzing and cleaning the motion trajectory data of the motion target and the sensor device resource monitoring result database.