A dynamic resource allocation method with an unknown objective function
The distributed reinforcement learning algorithm solves the problem of unknown objective function in dynamic resource allocation, achieving low-cost and efficient resource allocation, providing privacy protection and information security, and improving the scalability and convergence of the algorithm.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-15
- Publication Date
- 2026-04-03
AI Technical Summary
In dynamic resource allocation problems, especially in applications such as smart grids, existing technologies struggle to effectively handle situations where the objective function is unknown, resulting in high computational costs, low efficiency, and poor robustness.
Employing a distributed reinforcement learning algorithm, this paper designs a dynamic resource allocation method with an unknown objective function through a multi-agent system and network communication graph. Resource allocation is achieved by utilizing local information and exchanging with neighboring agents, and by combining policy update rules and exploration-utilization balance to solve the problem of an unknown objective function.
It achieves resource allocation with low computational complexity, privacy protection, and information security, improves the scalability and fast convergence of the algorithm, reduces computational costs, and ensures the feasibility and efficiency of resource allocation.
Smart Images

Figure CN116260775B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of network resource allocation technology, and in particular to a dynamic resource allocation method with an unknown objective function. Background Technology
[0002] Dynamic resource allocation problems have attracted increasing attention due to their wide applicability in communication networks, sensor networks, power grids, and many other fields. Centralized algorithms, as an early method for solving resource allocation, require a central controller to collect information from each agent. Their drawbacks include high computational cost, low computational efficiency, low robustness, and susceptibility to interference. Unlike centralized algorithms, distributed algorithms achieve collective intelligence by agents exchanging information with neighboring agents through communication networks. Therefore, distributed algorithms for dynamic resource allocation have received increasing attention. However, in many cases, the objective function of dynamic resource allocation is complex and difficult to express explicitly. For example, in smart grids, an application scenario for dynamic resource allocation problems, the explicit mathematical expression of the generator cost function is generally difficult to accurately characterize due to the influence of generator operating conditions, rendering distributed algorithms requiring an objective function expression ineffective. Therefore, designing distributed algorithms to solve dynamic resource allocation problems with unknown objective functions has become a significant challenge in the practical application of algorithms. Reinforcement learning mainly describes how agents achieve their learning objectives through interaction with an unknown environment. This patent, based on the learning method of reinforcement learning, designs a dynamic resource allocation method with an unknown objective function for dynamic resource allocation problems. The main contributions of this patent are as follows: (1) For the dynamic resource allocation problem with an unknown objective function and discrete feasible resource constraints, a dynamic resource allocation method with an unknown objective function is proposed. Compared with the traditional centralized algorithm, the dynamic resource allocation method with an unknown objective function proposed in this patent has the characteristics of reduced computational complexity and good scalability. (2) For the dynamic resource allocation problem with an unknown objective function and discrete feasible resource constraints, this patent proposes a dynamic resource allocation method with an unknown objective function, which effectively ensures that the joint strategy of the agent can generate feasible resource allocation at each time step, thereby ensuring the execution of the entire training process. Strategy. Summary of the Invention
[0003] This invention provides a dynamic resource allocation method for problems with unknown objective functions, which incorporates a novel distributed reinforcement learning algorithm. Based on a multi-agent system and a reinforcement learning model, this proposed method addresses the dynamic resource allocation problem in a distributed manner, enabling network resource allocation among agents even when the objective function is unknown. Furthermore, this method not only provides privacy and information security but also improves the algorithm's scalability. Simulation results demonstrate the good performance and effectiveness of this method in numerical examples of dynamic resource allocation problems with unknown objective functions.
[0004] A dynamic resource allocation method with an unknown objective function includes the following steps:
[0005] Step 1: Construct a model for the dynamic resource allocation problem with an unknown objective function;
[0006] Step 2: Design a dynamic resource allocation method with an unknown objective function, obtain a distributed iterative formula, and solve iteratively.
[0007] The dynamic resource allocation problem model in step 1) is specifically referred to as model (1).
[0008] (1)
[0009] in Represents intelligent agents The cost function, Represents intelligent agents In time Local resource allocation, express Total network resources at any given time For resource transfer functions, For intelligent agents Local resource allocation constraints. Additionally, agents can communicate via network graphs. To exchange information.
[0010] The dynamic resource allocation method with an unknown objective function in step 2 is designed as follows:
[0011] The objective function in the dynamic resource allocation model (1) has an unknown functional expression. In the dynamic resource allocation method where the objective function is unknown, assume that the first... Second trial Moment The value of , feasible resource allocation is The overall objective function value is Each intelligent agent Only local objective function values can be used. and local resource data It also exchanges information with neighboring intelligent agents.
[0012] Update the intelligent agent of Function: for any time intelligent agent Local q-function Depend on and local actions Definition. For time... and all potential ,set up Initial value of a function ,in It is a sufficiently large constant. Assume the agent is in the... Second trial 'Adopt feasible resource allocation within the time limit' ,in intelligent agent of The update rules for the function are as follows:
[0013] (5)
[0014] in , In particular, in At any moment, for all potential , ,make .
[0015] Update the intelligent agent Local strategy For all potential [times] in time t ,definition Allocate feasible resources. Intelligent agent. In the Second trial The local policy update rules within a given time are as follows:
[0016] (6)
[0017] Balancing exploration and exploitation: the use of intelligent agents Strategy balance exploration and utilization, in the first Second trial At that moment, with probability of use Or with The probability of using other feasible resource allocations.
[0018] Compared to existing technologies, the advantages of this invention are as follows: 1) The structure of the dynamic resource allocation method with an unknown objective function provides favorable characteristics in terms of privacy protection, information security, and scalability. The use of reinforcement learning effectively addresses the difficulty of an unknown objective function in dynamic resource allocation problems; 2) The dynamic resource allocation method with an unknown objective function proposed in this invention can solve the dilemma of obtaining the objective function in resource allocation problems. The dynamic resource allocation method with an unknown objective function proposed in this invention can efficiently and cost-effectively solve dynamic resource allocation problems with unknown objective functions. Its objective function optimization and fast convergence characteristics can bring economic and time benefits to network resource allocation. Using the dynamic resource allocation method with an unknown objective function proposed in this invention can achieve rapid modeling, ease of computation, rapid iteration, and rapid convergence, while reducing the computational resource consumption of the algorithm. Attached Figure Description
[0019] Figure 1 This is a network communication diagram;
[0020] Figure 2 A schematic diagram illustrating the learning process of a dynamic resource allocation method with an unknown objective function;
[0021] Figure 3 A schematic diagram of the strategy learned by a dynamic resource allocation method with an unknown objective function. Detailed Implementation
[0022] This invention provides a dynamic resource allocation method for problems with unknown objective functions, which incorporates a novel distributed reinforcement learning algorithm. Based on a multi-agent system and a reinforcement learning model, this proposed method addresses the dynamic resource allocation problem in a distributed manner, enabling network resource allocation among agents even when the objective function is unknown. Furthermore, this method not only provides privacy and information security but also improves the algorithm's scalability. Simulation results demonstrate the good performance and effectiveness of this method in numerical examples of dynamic resource allocation problems with unknown objective functions.
[0023] Example 1: A dynamic resource allocation method with an unknown objective function, comprising the following steps:
[0024] Step 1: Construct a model for the dynamic resource allocation problem with an unknown objective function;
[0025] Step 2: Design a dynamic resource allocation method with an unknown objective function, obtain a distributed iterative formula, and solve iteratively.
[0026] The dynamic resource allocation problem model in step 1) is specifically referred to as model (1).
[0027] (1)
[0028] in Represents intelligent agents The cost function, Represents intelligent agents In time Local resource allocation, express Total network resources at any given time For resource transfer functions, For intelligent agents Local resource allocation constraints. Additionally, agents can communicate via network graphs. To exchange information.
[0029] The dynamic resource allocation method with an unknown objective function in step 2 is designed as follows:
[0030] The objective function in the dynamic resource allocation model (1) has an unknown functional expression. In the dynamic resource allocation method where the objective function is unknown, assume that the first... Second trial Moment The value of , feasible resource allocation is The overall objective function value is Each intelligent agent Only local objective function values can be used. and local resource data It also exchanges information with neighboring intelligent agents.
[0031] Update the intelligent agent of Function: for any time intelligent agent Local q-function Depend on and local actions Definition. For time... and all potential ,set up Initial value of a function ,in It is a sufficiently large constant. Assume the agent is in the... Second trial 'Adopt feasible resource allocation within the time limit' ,in intelligent agent of The update rules for the function are as follows:
[0032] (5)
[0033] in , In particular, in At any moment, for all potential , ,make .
[0034] Update the intelligent agent Local strategy For all potential [times] in time t ,definition Allocate feasible resources. Intelligent agent. In the Second trial The local policy update rules within a given time are as follows:
[0035] (6)
[0036] Balancing exploration and exploitation: the use of intelligent agents Strategy balance exploration and utilization, in the first Second trial At that moment, with probability of use Or with The probability of using other feasible resource allocations.
[0037] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific numerical simulation implementations. Consider a dynamic resource allocation problem with four agents. The communication topology of the four agents is shown in Figure 1. Assume that the agents... The objective function is ,in , It is an intelligent agent The objective function coefficients. Consider the dynamic resource allocation problem (1) where the objective function is unknown, its (MV), (MV). Assume the initial resources are 200, and the agent's resource transfer function is , , ,in .Depend on Figure 2 As shown, the objective function generated by the algorithm proposed in this patent converges to the optimal value. The feasible resource configurations of the agents generated by the algorithm proposed in this patent over seven time periods are shown in Figure 3.
[0038] Simulation results Figure 2-3 As shown. Figure 2 The evolution of the objective function value generated by the algorithm during the learning process is shown.
[0039] from Figure 2 It can be seen that after 900,000 iterations of learning, the algorithm converges to the optimal objective function value.
[0040] Figure 3 It shows the resource allocation size of the agent at each time step, and the resource allocation of the agent learned by the distributed reinforcement learning algorithm converges to the optimal resource allocation.
[0041] It should be noted that the above embodiments are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Equivalent substitutions or alternatives made based on the above technical solutions shall all fall within the scope of protection of the present invention.
Claims
1. A dynamic resource allocation method with an unknown objective function, characterized in that, The method includes the following steps: Step 1: Construct a model for the dynamic resource allocation problem with an unknown objective function; Step 2: Design a dynamic resource allocation method with an unknown objective function, obtain a distributed iterative formula, and solve iteratively; The dynamic resource allocation problem model in step 1 is specifically referred to as model (1). (1) in Represents intelligent agents The cost function, Represents intelligent agents In time Local resource allocation, express Total network resources at any given time For resource transfer functions, For intelligent agents The local resource allocation is discretely constrained; in addition, agents interact with each other through network communication graphs. The dynamic resource allocation method with an unknown objective function in step 2 is designed as follows. In the dynamic resource allocation problem model, the expression of the objective function is unknown. In dynamic resource allocation methods where the objective function is unknown, assume the first... Second trial Moment The value of , feasible resource allocation is The overall objective function value is Each intelligent agent Able to use local objective function values The generated values, the agent interacts with its neighboring agents on local resources information; Update the intelligent agent of Function: for any time intelligent agent Local q-function Depend on and local actions Definition, for time and all potential ,set up Initial value of a function ,in It is a constant, assuming the agent is in the th... Second trial 'Adopt feasible resource allocation within the time limit' ,in intelligent agent of The update rules for the function are as follows: (5) in , , exist At any moment, for all potential , ,make , Update the intelligent agent Local strategy For all potential [times] in time t ,definition For feasible resource allocation, intelligent agents In the Second trial The local policy update rules within a given time are as follows: (6) Balancing exploration and exploitation: the use of intelligent agents Strategy balance exploration and utilization, in the first Second trial At that moment, with probability of use Or with The probability of using other feasible resource allocations.
Citation Information
Patent Citations
Path planning method based on reinforcement learning algorithm in dynamic environment
CN111649758A
System and Method for Control Constrained Operation of Machine with Partially Unmodeled Dynamics Using Lipschitz Constant
US20210003973A1