Network resource allocation method and device, equipment and medium

Through the reinforcement learning agent method, dynamically adjusting the network resource allocation strategy, the problem of resource allocation in the existing technology cannot respond quickly to traffic changes, and efficient utilization and optimization of network resources are achieved.

CN120281652APending Publication Date: 2025-07-08SHENZHEN JIUNIU YIMAO INTELLIGENT IOT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510432142.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the prior art, network resource allocation methods have limitations in a highly dynamically changing network environment, and cannot respond quickly to traffic changes, resulting in waste or overuse of resources, making it difficult to achieve optimal resource utilization.

Method used

The reinforcement learning agent method is adopted, and the network resource allocation strategy is dynamically adjusted by setting state space, action space and reward functions, and the reinforcement learning agent interacts in the network environment to learn how to take actions to maximize cumulative rewards and achieve adaptive optimization of network resources.

Benefits of technology

It improves network performance and service quality, maximizes resource utilization, reduces resource waste and delays, and realizes efficient utilization of network resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281652A_ABST
    Figure CN120281652A_ABST
Patent Text Reader

Abstract

The invention discloses a network resource allocation method, device, equipment and medium, the allocation of network resources is realized through a reinforcement learning agent, a state space, an action space and a reward function are set in the agent, available network resources are described through the state space, and the action space is set in the reward function. A network resource allocation strategy which may be adopted is described through an action space, and the effect of network resource allocation is determined through a reward function. In a dynamic network resource allocation process, a reward value is determined according to a network resource allocation condition of a previous state, and a network resource allocation strategy of a next state is determined according to the reward value, so that a complex strategy for automatically adjusting resource allocation is realized through a reinforcement learning agent, the network performance and the service quality are improved, and the network resource allocation efficiency is improved. The self-adaptive optimization of dynamic network resource allocation is realized, meanwhile, the resource utilization rate is maximized, the waste is reduced, the delay is reduced, and the efficient utilization of network resources is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer network technologies, and particularly to a method, apparatus, device, and medium for network resource allocation. Background Art

[0002] With the development of digital transformation and Internet technologies, the network has become an indispensable part of modern society. Whether it is individual users or enterprise organizations, they all need to access various online services through the network. With the rapid growth of data traffic, the demand for network resources is also continuously increasing. However, the contradiction between limited network resources and the growing demand has become increasingly prominent, resulting in problems such as increased latency and insufficient bandwidth, which directly affect the user experience and service quality.

[0003] In the prior art, network resource allocation is generally based on static configuration, or a simple dynamic adjustment strategy is adopted. Among them, static configuration refers to a pre-set resource allocation method that remains unchanged during operation; the simple dynamic adjustment strategy is a method of adjusting resource allocation in real time according to the resource allocation situation to adapt to changing demands.

[0004] However, based on static configuration or simple dynamic adjustment strategies. Although these methods can alleviate the resource allocation problem to a certain extent, they have certain limitations in a highly dynamic network environment. Among them, the static configuration method cannot respond quickly when the traffic pattern changes, which may lead to resource waste or overuse. Although the simple dynamic strategy can make some adjustments according to real-time traffic, it lacks the ability of global view and long-term planning, and it is difficult to achieve the optimal resource utilization state.

[0005] Therefore, the above problems existing in the prior art remain to be improved. Summary of the Invention

[0006] The main purpose of the present invention is to provide a method, apparatus, device, and medium for network resource allocation to solve the above technical problems.

[0007] In a first aspect, the present invention provides a method for network resource allocation, the method comprising:

[0008] Obtain a target intelligent agent, the target intelligent agent including a state space, an action space, and a reward function, wherein the state space includes a set of network resource states at different times; the action space includes a plurality of action strategies, and the action strategies are used for allocating network resources; the reward function is used to evaluate the immediate feedback of taking a preset action strategy under a preset network resource state;

[0009] Input the network resource status information monitored at each moment from the state space into the reinforcement learning environment, so that the target agent generates the first action policy from the action space according to the preset rules;

[0010] The reward function obtains the first reward value according to the network resource status allocated by the first action policy;

[0011] The target agent obtains the second action policy according to the first reward value until the network resource allocation is completed.

[0012] Preferably, obtaining the target agent includes:

[0013] Process at least one of the CPU utilization rate, memory utilization rate, storage space utilization rate, or network bandwidth utilization rate in the network resources into a target matrix, and all possibilities of changes in the target matrix are the state space.

[0014] Preferably, processing at least one of the CPU utilization rate, memory utilization rate, storage space utilization rate, or network bandwidth utilization rate in the network resources into a target matrix includes:

[0015] Process the CPU utilization rate M C , memory utilization rate M M , storage space utilization rate M S and network bandwidth utilization rate M B into the first matrix;

[0016] Perform Max-Min normalization processing on the CPU utilization rate M C , memory utilization rate M M , storage space utilization rate M S and network bandwidth utilization rate M B in the first matrix to obtain the second matrix;

[0017] Stack the second matrix to obtain the target matrix:

[0018] X = [M C , M M , M S , M B

[0019] All possibilities of the target matrix are the state space:

[0020]

[0021] Where S represents the state space and X represents different states of the target matrix.

[0022] Preferably, obtaining the target agent includes:

[0023] Obtain the action space: ​

[0024]

[0025] a1, a2, …, an represent different action strategies, defined as:

[0026] a t =[A C,i (t), A M,i (t), A S,i (t), A B,ij (t)]

[0027] Among them, a t represents the action strategy of network resource allocation, A C,i (t) represents the CPU resources allocated to the i-th node at time t, A M,i (t) represents the memory resources allocated to the i-th node at time t, A S,i (t) represents the storage space resources allocated to the i-th node at time t, A B,ij (t) represents the bandwidth resources allocated to the link from node i to node j at time t.

[0028] Preferably, obtaining the target agent includes:

[0029] Design the reward function as:

[0030]

[0031] Among them, represents the reward value, represents the objective function, U(t) represents the resource utilization rate, D(t) represents the average delay of data packet transmission in network resources, C(t) represents the cost, and α, β, and γ represent weight factors.

[0032] Preferably, the resource utilization rate:

[0033]

[0034] Among them, C max,i , M max,i , S max,i and B max,ij are respectively the maximum CPU, memory, storage space utilization rates of the i-th node, and the maximum bandwidth utilization rate from node i to node j;

[0035] The average delay:

[0036]

[0037] Among them, F(t) represents the number of data packets at time t, d f (t) represents the transmission time of the f-th data packet;

[0038] Cost:

[0039]

[0040] Among them, A C,i (t), A M,i (t) and A S,i (t) respectively represent the CPU resources, memory resources, and storage space resources allocated to the i-th node at time t, and A B,ij (t) represents the bandwidth resources allocated to the link from node i to j at time t.

[0041] Preferably, the target agent obtains a second action policy according to the first reward value until the network resource allocation is completed, including:

[0042] When all available resources have been allocated or the maximum allocation times limit is reached, or when the quality of service QoS metric reaches a preset threshold, it is determined that the network resource allocation is completed.

[0043] In a second aspect, the present invention also provides a network resource allocation device, including:

[0044] An acquisition unit, configured to acquire a target agent, where the target agent includes a state space, an action space, and a reward function. Among them, the state space includes a set of network resource states at different times; the action space includes multiple action policies, and the action policies are used to allocate network resources; the reward function is used to evaluate the immediate feedback of taking a preset action policy under a preset network resource state;

[0045] An execution unit, configured to generate a first action policy through the target agent according to the real-time network resource state, and the real-time network resource state is the network resource state monitored currently;

[0046] A feedback unit, configured to obtain a first reward value according to the network resource state allocated by the reward function according to the first action policy;

[0047] The execution unit is further configured to: the target agent obtains a second action policy according to the first reward value until the network resource allocation is completed.

[0048] In a third aspect, the present invention also provides a computer device, including: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the network resource allocation method as described in the first aspect.

[0049] In a fourth aspect, the present invention also provides a storage medium, where the storage medium includes a stored computer program. When the computer program runs, it controls the device where the storage medium is located to execute the network resource allocation method as described in the first aspect.

[0050] Advantageous technical effects of the present invention: The present invention realizes the allocation of network resources through an agent of reinforcement learning. A state space, an action space, and a reward function are set in the agent. Among them, the available network resources are described through the state space, the possible network resource allocation strategies are described through the action space, and the effect of network resource allocation is determined through the reward function. In the process of dynamic network resource allocation, the reward value is determined according to the network resource allocation situation of the previous state, and the network resource allocation strategy of the next state is determined according to the reward value. In this way, the complex strategy of automatically adjusting resource allocation is realized through the agent of reinforcement learning, the network performance and service quality are improved, the adaptive optimization of dynamic network resource allocation is realized, and at the same time, the resource utilization rate is maximized, waste is reduced, delay is reduced, and the efficient utilization of network resources is realized. Brief Description of the Drawings

[0051] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0052] Figure 1 Schematic diagram of an embodiment of the network resource allocation method provided by an embodiment of the present invention;

[0053] Figure 2 Schematic diagram of the state space in the network resource allocation method provided by an embodiment of the present invention;

[0054] Figure 3 Schematic diagram of another embodiment of the network resource allocation method provided by an embodiment of the present invention;

[0055] Figure 4 Schematic diagram of the network resource allocation device provided by an embodiment of the present invention;

[0056] Figure 5 Schematic diagram of the computer device provided by an embodiment of the present invention. Detailed Embodiments

[0057] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0058] It should be understood that when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.

[0059] It should also be understood that the terms used in this specification of the present invention are merely for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.

[0060] It should be further understood that the term "and / or" used in this specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0061] The embodiments of the present application provide a network resource allocation method, which is used for the dynamic allocation of network resources. Through the guidance of reward values, it can automatically adjust the resource allocation strategy according to the changes in network conditions, improve network performance and service quality, realize the adaptive optimization of dynamic network resource allocation. At the same time, by intelligently allocating resources, it maximizes resource utilization, reduces waste, reduces latency, and realizes the efficient utilization of network resources.

[0062] First, the relevant technical concepts involved in the embodiments of the present application are described. The embodiments of the present application need to allocate resources such as CPU, memory, storage space or network bandwidth to different virtual machines or containers. The goal is to maximize the resource utilization rate U(t), minimize the time delay D(t) and cost C(t), and ensure the quality of service (QoS), etc. The relevant parameters and their meanings are defined as follows:

[0063] CPU utilization rate: C i (t), which represents the CPU utilization rate of the i-th node at time t.

[0064] Memory utilization rate: M i (t), which represents the memory utilization rate of the i-th node at time t.

[0065] Storage space utilization rate: S i (t), which represents the storage space utilization rate of the i-th node at time t.

[0066] Network bandwidth utilization rate: B ij (t), which represents the bandwidth utilization rate of the link from node i to node j at time t.

[0067] The resource utilization rate U(t) is calculated as shown in formula (1), where Cmax,i , M max,i , S max,i and B max,ij are respectively the maximum CPU, memory, storage space utilization rate of the i-th node, and the maximum bandwidth utilization rate from node i to node j.

[0068]

[0069] The delay D(t) is defined as the average transmission time of all data packets, as shown in formula (2), where F(t) represents the number of data packets at time t, and d f (t) represents the transmission time of the f-th data packet.

[0070]

[0071] The cost C(t) is defined as the total cost required for resource allocation, as shown in formula (3), where A C,i (t), A M,i (t), and A S,i (t) respectively represent the CPU resource, memory resource, and storage space resource allocated to the i-th node at time t, and A B,ij (t) represents the bandwidth resource allocated to the link from node i to j at time t.

[0072]

[0073] The optimization objective is to maximize the resource utilization rate U(t) and minimize the delay D(t) and cost C(t). Define the objective function as shown in formula (4), where β l ∈[0, 1], l = 1, 2, 3 are weight factors used to balance the relationship between resource utilization rate, delay, and cost:

[0074]

[0075] Based on this, as Figure 1 shown, the network resource allocation method provided by the embodiments of this application includes the following steps:

[0076] S1. Obtain the target agent, which includes a state space, an action space, and a reward function. Among them, the state space includes a set of network resource states at different times; the action space includes multiple action strategies, and the action strategies are used to allocate network resources; the reward function is used to evaluate the immediate feedback of taking a preset action strategy under a preset network resource state.

[0077] In this embodiment, the target agent, as an implementation of artificial intelligence, is used to achieve the dynamic allocation of network resources. In recent years, significant progress has been made in the field of artificial intelligence, and reinforcement learning (RL), as a machine learning method, has received extensive attention. Reinforcement learning enables an agent to interact with the environment and learn how to take actions to maximize a certain cumulative reward, providing a new approach to solving complex decision-making problems. It is particularly suitable for environments with uncertainty and dynamic change characteristics, which are exactly the characteristics of the network resource allocation problem.

[0078] The key to the algorithm design of the target agent in the embodiments of this application lies in the design of the agent's state space, action space, and reward function. Among them, the state space includes the set of network resource states at different times, that is, all available network resources; the action space includes multiple action strategies, that is, the set of action strategies, and the action strategies are used to allocate network resources; the reward function is used to evaluate the immediate feedback of taking a preset action strategy under a preset network resource state, and is used to evaluate the benefits of the current network resource allocation.

[0079] The target agent in the embodiments of this application masters the relevant situations of network resource allocation through the state space, action space, and reward function. For the specific implementation methods of the state space, action space, and reward function, the embodiments of this application do not limit them. For ease of understanding, the preferred implementation methods are provided as follows, but they do not constitute a limitation to the embodiments of this application.

[0080] (1) State space.

[0081] In the embodiments of this application, the state space refers to the set of all possible states of the environment where the target agent is located. These states can be discrete or continuous. It defines all possible situations that the agent can be in at any time, usually including a complete or partial description of the environment.

[0082] In the solution of the embodiments of this application, at least one of the CPU utilization rate, memory utilization rate, storage space utilization rate, or network bandwidth utilization rate in the network resources is processed as a target matrix, and all possibilities of change of the target matrix are the state space.

[0083] Preferably, in the embodiments of this application, the CPU utilization rate M C , memory utilization rate M M , storage space utilization rate M S , and network bandwidth utilization rate M B are processed into a first matrix. Since their numerical sizes are different, we perform Max-Min normalization on these four parameters to obtain the following second matrix:

[0084]

[0085] Among them, m ij represents the element in the i-th row and j-th column of the matrix.

[0086] Stack the second matrix to obtain the above target matrix:

[0087] X = [M C , M M , M S , M B

[0088] This target matrix is the set of all possible states of the network resources in the environment where the target agent is located. Its schematic diagram is as Figure 2 shown. All possibilities of the multi-channel matrix change are the state space of the agent, denoted as For example, allocating resources to a certain node or releasing resources of a certain node, etc., will cause the current state s t to change to s t+1 . Through such a design, the target agent can decide the best resource allocation strategy according to the current resource usage and resource request situations.

[0089] (2) Action space.

[0090] In the embodiments of the present application, the action space refers to the set of all possible actions that the target agent can take. These actions can be discrete or continuous, depending on the requirements of the application scenario. The discrete action space usually consists of a set of finite options, while the continuous action space allows the agent to freely select actions within a certain range. The action in the embodiments of the present application is to allocate network resources, which is a discrete action, and its meaning is the resource allocation strategy taken by the agent for the node at different times.

[0091] It consists of the following parameters:

[0092] 1) CPU allocation: A C,i (t), which represents the CPU resources allocated to the i-th node at time t.

[0093] 2) Memory allocation: A M,i (t), which represents the memory resources allocated to the i-th node at time t.

[0094] 3) Storage space allocation: A S,i (t), which represents the storage space resources allocated to the i-th node at time t.

[0095] 4) Network bandwidth allocation: A B,ij (t), which represents the bandwidth resources allocated to the link from node i to node j at time t. ​

[0096] Action a at the current moment t =[A C,i (t), A M,i (t), A S,i (t), A B,ij (t)], different allocation situations represent different actions, so the action space is expressed as wherein, the above a1, a2,..., an represent different action strategies.

[0097] In this way, through the allocation situations of CPU allocation, memory allocation, storage space allocation and network bandwidth allocation at different times, all possible sets of allocation strategies are combined to form the action space defined by the embodiments of the present application.

[0098] (3) Reward function.

[0099] In the embodiments of the present application, the reward function defines the immediate feedback obtained by the agent after taking a certain action in a given state. The reward function is a key component in the agent's learning process, which guides the agent to identify which behaviors are beneficial and which are harmful. The reward function is usually a scalar value that the agent attempts to maximize, which reflects the quality of the agent's behavior.

[0100] The embodiments of the present application aim to maximize the resource utilization rate U(t) and minimize the delay D(t) and cost C(t), so the reward function is designed as shown in formula (6) to guide the agent to achieve the optimization goal.

[0101]

[0102] Through the above steps, the construction of the target agent is realized, and through this target agent, the dynamic allocation of network resources can be realized.

[0103] S2. Input the network resource status information monitored at each moment from the state space into the reinforcement learning environment, so that the target agent generates the first action strategy from the action space according to the preset rules.

[0104] In this embodiment, the state space contains a set of network resource states at different times. Input the resource status information monitored at each moment into the reinforcement learning environment and process it into the state matrix of the target agent, so that the target agent starts the allocation work of network resources. The above first action strategy is generated from the action space according to the preset rules, and the specific method is not limited in the embodiments of the present application.

[0105] S3. The reward function obtains the first reward value according to the network resource status allocated by the first action strategy.

[0106] In this embodiment, the reward function obtains the next state and the reward value obtained after the execution of the current action according to the action strategy adopted by the current agent. In this way, the agent can judge the benefit of the first action strategy through the reward value, so as to determine whether the allocation strategy is reasonable, thereby establishing a decision-making basis for the next network resource allocation.

[0107] S4. The target agent obtains a second action strategy according to the first reward value until the network resource allocation is completed.

[0108] In this embodiment, after the target agent outputs the second action strategy through the action space, the above work is still repeated to determine the reward value obtained by the second action strategy. In this way, through the target agent, the dynamic adjustment of network resources is realized.

[0109] It should be particularly noted that after each execution of the action strategy, the target agent will update the current state and judge whether the current state is a termination state. If it is a non-termination state, an action strategy will be generated according to the current state, and the action will be input into the reinforcement learning environment for execution, and the current process will continue until the termination state is reached, indicating that the current resource allocation is completed.

[0110] The above "termination state" refers to:

[0111] (1) Resource allocation completed: All available resources have been allocated or the maximum allocation times limit has been reached.

[0112] (2) Performance index up to standard: The quality of service (QoS) index has reached the preset threshold. QoS (Quality of Service) is the quality of service. Under limited bandwidth resources, QoS allocates bandwidth for various services and provides end-to-end service quality assurance for services. For example, voice, video, and important data applications can be given priority services through QoS configuration in network devices. In this embodiment, when, for example, the average load is lower than the preset threshold, it can be determined that the network resource allocation is in a termination state.

[0113] In this embodiment, the concept of an agent is defined for network resource allocation. Further, a state space, an action space, and a reward function are set in the agent. Among them, the available network resources are described by the state space, the possible network resource allocation strategies are described by the action space, and the effect of network resource allocation is determined by the reward function. In the dynamic network resource allocation process, the reward value is determined according to the network resource allocation situation of the previous state, and the network resource allocation strategy of the next state is determined according to the reward value. In this way, guided by the reward value, the resource allocation strategy can be automatically adjusted according to the changes in the network conditions, improving the network performance and service quality. Through the technical advantages of reinforcement learning, the target agent interacts with the environment to learn how to take actions to maximize a certain cumulative reward, providing a new way to solve the network resource allocation problem of complex decision-making. The adaptive optimization of dynamic network resource allocation is realized. At the same time, by intelligently allocating resources, the resource utilization rate is maximized, waste is reduced, and latency is decreased, achieving the efficient utilization of network resources.

[0114] The above steps S1 to S4 introduce an implementation method of the network resource allocation method provided by the embodiments of the present application. The specific implementation manner of this method is not limited by the embodiments of the present application. For ease of understanding, as Figure 3 shown, a preferred work process is provided as follows.

[0115] S10. Monitor the current network resource allocation information.

[0116] In this embodiment, the system monitors the network resource allocation information in real time, thus providing a judgment basis for the decision-making of the target agent.

[0117] S20. Send the resource allocation information to the target agent.

[0118] In this embodiment, the information after each resource allocation is sent to the target agent so that the target agent can make a decision.

[0119] S30. The target agent generates a reward value according to the resource allocation information.

[0120] In this embodiment, the target agent generates a reward value for the previous round of network resource allocation through the reward function, thereby evaluating the benefit of this round of network resource allocation.

[0121] S40. The target agent determines an action strategy according to the reward value.

[0122] In this embodiment, the target agent determines the action strategy for the next round of network resource allocation according to the reward value of the previous round of network resource allocation. In this way, through the computational advantages of artificial intelligence, the dynamic allocation efficiency of network resources is ensured to be maximized.

[0123] S50. Determine whether the network resource allocation has reached the termination state.

[0124] In this embodiment, when the network resources have not reached the termination state, the above steps S20 to S50 are repeated until the network resource allocation reaches the termination state. After that, the algorithm completes the resource allocation and ends the network resource allocation process.

[0125] The above steps S10 to S50, as a specific implementation manner of the algorithm flow in this embodiment, introduce a preferred embodiment of the network resource allocation method provided by the embodiments of the present application.

[0126] For ease of understanding, the embodiments of the present application further provide a more detailed specific embodiment as follows.

[0127] S100. Initialize the network parameters θ of the estimated value Q(s, a; θ), the network parameters θ' of the target value Q'(s, a; θ'), and the experience replay pool D.

[0128] S200. Reset the environment to obtain the initial state s t , and reset the total reward value.

[0129] In this embodiment, the above steps S100 and S200, as initialization steps, perform an initialization reset on information such as the reward value of the reward function.

[0130] S300. Conduct exploration and select the action of the target agent. If the random probability value is less than the exploration strategy parameter ∈, a random action a is generated t . If the random probability value is greater than the exploration strategy parameter ∈, the current state s t is input into the estimated value network to obtain the action a t .

[0131] In this embodiment, the above step S300 gives a preferred embodiment of the rule for the target agent to initially select the action strategy from the action space. Taking the strategy parameter ∈ as the standard, the initial action strategy at is determined by the random probability value.

[0132] S400. Interact with the environment. The target agent executes the action a in the environment t , that is, perform a resource allocation once, change the state, and obtain the next state s t+1 , the reward value r generated by executing this action t and whether it is the termination state flag bit done, and accumulate the current reward value. Then, store the current experience (s t , a t , r t , s t+1 ) into the experience replay pool D.

[0133] In this embodiment, during the network resource allocation process, the target agent updates the action policy according to the feedback of the reward function. And the obtained experience is put back into the experience pool. In this way, during the working process, the target agent always learns the action policy of network resource allocation, continuously improving the maturity of the target agent.

[0134] S500. Conduct training. If the number of experiences in the current experience replay pool is greater than the minimum training data batch size, start training. Sample batch_size experience data (s t , a t , r t , s t+1 ) from the experience replay pool D, calculate the loss function L(θ) through the experience data, and then update the network parameters through gradient descent.

[0135] In this embodiment, the target agent is trained through the experience values put back into the experience pool.

[0136] S600. Determine whether the current episode number ep is a multiple of the update frequency τ of the target network. If so, update the target network parameter θ′ by estimating the value network parameter θ.

[0137] S700. Determine whether the current state is a termination state. If it is a termination state, output the allocation policy, end the loop, and enter step S800. If not, continue to explore and return to step S300.

[0138] Step 8: Update the next state s t+1 to the current state s t , and reduce the exploration policy parameter ∈. Enter the next episode ep. If ep is maxEpisode, the algorithm ends. Otherwise, return to step S200.

[0139] The above steps S100 to S800, as a more specific embodiment, detail the working process of the target agent and the process of update and iteration through the experience pool.

[0140] It should be noted that for the above method as a working method of computer software, the specific implementation manner of its code is not limited in the embodiments of the present application. For ease of understanding, a piece of pseudocode for implementing the network resource allocation method of the present application is provided as follows.

[0141] Input: CPU utilization rate C i (t), memory utilization rate M i (t), storage space utilization rate S i (t), network bandwidth utilization rate B ij (t), weight factor β l∈ [0, 1], l = 1, 2, 3, learning rate α, reward discount factor γ, target network update frequency τ, exploration strategy parameter ∈, and exploration decay factor ∈_decay, training batch size batch_size, as well as experience replay pool D, maximum number of training episodes maxEpisode

[0142] Output: Resource allocation strategy.

[0143] 1: Initialize the network parameters θ of the estimated value Q(s, a; θ).

[0144] 2: Initialize the network parameters θ′ of the target value Q′(s, a; θ′).

[0145] 3: Initialize the experience replay pool D

[0146] 4: For ep = 1 ← maxEpisode do

[0147] 5: Reset the environment to obtain the initial state s t ;

[0148] 6: Reset the total reward value total_reward = 0;

[0149] 7: While True

[0150] 8: # The agent selects an action based on the current state

[0151] 9: If random() < ∈ do

[0152] 10: Randomly select an action a t ;

[0153] 11: Else

[0154] 12: Input the current state s t into the estimated value network Q(s, a; θ) to obtain the action;

[0155] 13: # Interact with the environment

[0156] 14: Obtain the next state s t+1 , reward value r t and state done ← interact with the environment env.step(a t )

[0157] 15: Accumulate the reward value total_reward += r t

[0158] 16: Store the experience (s t , a t , r t , st+1 ) Store in the experience replay pool D

[0159] 17: # If the number in the experience pool is greater than the batch size of the training data, start training

[0160] 18: If len(D)>batch_size do

[0161] 19: Sample batch_size experience data (s t ,a t ,r t ,s t+1 )

[0162] 20: Calculate the loss function L(θ)←(r+γMaX a′ Q(s′,a′;θ′)-Q(s,a;θ)) 2

[0163] 21: Update network parameters

[0164] 22: If ep%τ==0do

[0165] 23: θ′←θ # Update target network parameters

[0166] 24: If done

[0167] 25: Output the allocation strategy;

[0168] 26: break;

[0169] 27: s t ←s t+1 # Update the next state to the current state

[0170] 28: ∈←∈*∈_decay # Decrease the exploration rate

[0171] 29: End While

[0172] 30: End For

[0173] In summary, the network resource allocation method provided by the embodiments of the present application defines the concept of an agent, and further sets a state space, an action space, and a reward function in the agent. Among them, the available network resources are described by the state space, the possible network resource allocation strategies are described by the action space, and the effect of network resource allocation is determined by the reward function. In the dynamic network resource allocation process, the reward value is determined according to the network resource allocation situation of the previous state, and the network resource allocation strategy of the next state is determined according to the reward value. In this way, guided by the reward value, the resource allocation strategy can be automatically adjusted according to the changes in the network conditions, improving the network performance and service quality, realizing the adaptive optimization of dynamic network resource allocation. At the same time, by intelligently allocating resources, the resource utilization rate is maximized, waste is reduced, and the delay is reduced, realizing the efficient utilization of network resources.

[0174] Referring to Figure 4 , Figure 4 FIG. is a schematic diagram of a network resource allocation device provided by an embodiment of the present invention. Corresponding to the above network resource allocation method, an embodiment of the present invention further provides a network resource allocation device, including:

[0175] An acquisition unit 10, configured to acquire a target agent, where the target agent includes a state space, an action space, and a reward function. Among them, the state space includes a set of network resource states at different times; the action space includes multiple action strategies, and the action strategies are used to allocate network resources; the reward function is used to evaluate the immediate feedback of taking a preset action strategy under a preset network resource state;

[0176] An execution unit 20, configured to input the network resource state information monitored at each moment from the state space into a reinforcement learning environment, so that the target agent generates a first action strategy from the action space according to a preset rule;

[0177] A feedback unit 30, configured to obtain a first reward value according to the network resource state allocated by the first action strategy by the reward function;

[0178] The execution unit 10 is further configured to: the target agent obtains a second action strategy according to the first reward value until the network resource allocation is completed.

[0179] In this embodiment, the target agent obtained by the acquisition unit is provided with a state space, an action space, and a reward function. Among them, the available network resources are described by the state space, the possible network resource allocation strategies are described by the action space, and the effect of network resource allocation is determined by the reward function. In the dynamic network resource allocation process, the reward value is determined according to the network resource allocation situation of the previous state, and the network resource allocation strategy of the next state is determined according to the reward value. In this way, guided by the reward value, the resource allocation strategy can be automatically adjusted according to the changes in the network conditions, improving the network performance and service quality, realizing the adaptive optimization of dynamic network resource allocation. At the same time, by intelligently allocating resources, the resource utilization rate is maximized, waste is reduced, and latency is decreased, achieving the efficient utilization of network resources.

[0180] Referring to Figure 5 , Figure 5 FIG. is a schematic diagram of a computer device provided by an embodiment of the present invention. Corresponding to the above network resource allocation method, an embodiment of the present invention also provides a computer device. The device includes: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the network resource allocation method of any of the foregoing embodiments is implemented.

[0181] In this embodiment, the memory of the computer device may be any one or a combination of several of various media that can store computer programs, such as a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a register, a hard disk, a mobile hard disk, a solid state drive (SSD), etc. Among them, various program codes and data required to implement the network resource allocation function are stored in the memory, and these codes and data are organized and stored according to functional modules.

[0182] The processor of the computer device may be at least one of a central processing unit (CPU), a graphics processing unit (GPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), etc. The processor is connected to the memory through a bus and is used to execute the computer program stored in the memory. When the computer device runs, the processor loads and executes the corresponding program codes from the memory, thereby implementing the following functions:

[0183] Obtain a target agent, where the target agent includes a state space, an action space, and a reward function. Among them, the state space includes a set of network resource states at different times; the action space includes multiple action strategies, and the action strategies are used to allocate network resources; the reward function is used to evaluate the immediate feedback of taking a preset action strategy under a preset network resource state;

[0184] For inputting the network resource status information monitored at each moment from the state space into the reinforcement learning environment, so that the target intelligent agent generates a first action policy from the action space according to a preset rule;

[0185] The reward function obtains a first reward value according to the network resource status allocated by the first action policy;

[0186] The target intelligent agent obtains a second action policy according to the first reward value until the network resource allocation is completed.

[0187] In this computer device, the processor and the memory are connected and communicate through a standard system bus. The system bus may include a data bus, an address bus, and a control bus. In addition, this computer device may further include a network interface for data communication with other devices. All these hardware components cooperate together to implement the various functions of the network resource allocation method when the processor executes the program code in the memory.

[0188] Through this implementation manner based on general computing hardware, the embodiments of the present invention can be flexibly deployed on various server platforms, with good adaptability and portability. At the same time, due to the programmed implementation manner, the system functions are also easy to upgrade and expand, and can timely adapt to the changes in business requirements.

[0189] Corresponding to the above network resource allocation method, an embodiment of the present invention further provides a storage medium, where the storage medium includes a stored computer program, and when the computer program runs, it controls the device where the storage medium is located to execute the network resource allocation method of any of the foregoing embodiments.

[0190] In this embodiment, the storage medium is used to store a computer program for implementing the network resource allocation method. The storage medium may be various computer-readable storage media such as ROM, RAM, disk, optical disc, USB flash drive, SD card, etc. that can store program codes.

[0191] Specifically, the computer program stored in the storage medium includes multiple program modules, and when these program modules are loaded and executed by the computer device, the computer device can be enabled to execute the foregoing network resource allocation method.

[0192] When the storage medium is installed in the computer device and loaded and run by the processor, the stored program code will enable the computer device to execute each step of the network resource allocation method and realize the dynamic allocation of network resources. In this way, the embodiments of the present invention can be conveniently deployed and used on various computer devices, with good portability and practicality.

[0193] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0194] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of each unit is only a logical function division for the network resource allocation method. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed.

[0195] The steps in the method embodiments of the present invention can be adjusted, combined, and deleted according to actual needs. The units in the device embodiments of the present invention can be combined, divided, and deleted according to actual needs. In addition, the functional units in each embodiment of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0196] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention.

[0197] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A network resource allocation method, characterized in that, The method includes: Obtaining a target agent, where the target agent includes a state space, an action space, and a reward function. Among them, the state space includes a set of network resource states at different times; the action space includes multiple action policies for allocating network resources; the reward function is used to evaluate the immediate feedback of taking a preset action policy under a preset network resource state. Inputting the network resource state information monitored at each moment from the state space into a reinforcement learning environment, so that the target agent generates a first action policy from the action space according to a preset rule. The reward function obtains a first reward value according to the network resource state allocated by the first action policy. The target agent obtains a second action policy according to the first reward value until the network resource allocation is completed.

2. The method according to claim 1, wherein The obtaining of the target agent includes: Processing at least one of the CPU utilization rate, memory utilization rate, storage space utilization rate, or network bandwidth utilization rate in the network resources into a target matrix, and all possibilities of the change of the target matrix are the state space.

3. The method according to claim 2, wherein The processing of at least one of the CPU utilization rate, memory utilization rate, storage space utilization rate, or network bandwidth utilization rate in the network resources into a target matrix includes: Process the CPU utilization rate M C , the memory utilization rate M M , the storage space utilization rate M S and the network bandwidth utilization rate M B into the first matrix; For the CPU utilization rate M in the first matrix C , the memory utilization rate M M , the storage space utilization rate M S and the network bandwidth utilization rate M B perform Max-Min normalization to obtain a second matrix; Stacking the second matrix to obtain the target matrix: X = [M C , M M , M S , M B ​ All possibilities of the target matrix are the state space: S = {X1, X2, …, X n} = {s1, s2, …, s n} Among them, S represents the state space, and X represents different states of the target matrix.

4. The method according to claim 1, characterized in that The obtaining of the target agent includes: Obtaining an action space: The a1, a2,..., an represent different action policies, which are defined as: a t = [A C,i (t), A M,i (t), A S,i (t), A B,ij (t)] Among them, a t represents the action policy for network resource allocation, A C,i (t) represents the CPU resources allocated to the i-th node at time t, A M,i (t) represents the memory resources allocated to the i-th node at time t, A S,i (t) represents the storage space resources allocated to the i-th node at time t, A B,ij (t) represents the bandwidth resources allocated to the link from node i to node j at time t.

5. The method according to claim 1, wherein The obtaining of the target agent includes: Designing the reward function as: Among them, represents the reward value, represents the objective function, U(t) represents the resource utilization rate, D(t) represents the average delay of data packet transmission in the network resources, C(t) represents the cost, and the α, β, and γ represent weight factors.

6. The method according to claim 5, wherein The resource utilization rate: Among them, C max,i , M max,i , S max,i and B max,ij are respectively the maximum CPU, memory, storage space utilization rate of the i-th node, and the maximum bandwidth utilization rate from node i to node j; The average delay: where F(t) represents the number of data packets at time t, and d f (t) represents the transmission time of the f-th data packet; The cost: Among which A C,i (t), A M,i (t) and A S,i (t) respectively represent the CPU resources, memory resources and storage space resources allocated to the i-th node at time t, and A B,ij (t) represents the bandwidth resources allocated to the link from node i to j at time t.

7. The method according to any one of claims 1 to 6, characterized in that The target agent obtains a second action policy according to the first reward value until the network resource allocation is completed, including: When all available resources are allocated or the maximum allocation times limit is reached, or when the quality of service QoS index reaches a preset threshold, it is determined that the network resource allocation is completed.

8. A network resource allocation device, characterized in that, Includes: An obtaining unit for obtaining a target agent, where the target agent includes a state space, an action space, and a reward function. Among them, the state space includes a set of network resource states at different times; the action space includes multiple action policies for allocating network resources; the reward function is used to evaluate the immediate feedback of taking a preset action policy under a preset network resource state. An execution unit for inputting the network resource state information monitored at each moment from the state space into a reinforcement learning environment, so that the target agent generates a first action policy from the action space according to a preset rule. A feedback unit for the reward function to obtain a first reward value according to the network resource state allocated by the first action policy. The execution unit is further used for: the target agent obtains a second action policy according to the first reward value until the network resource allocation is completed.

9. A computer device, characterized in that, Includes: A processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the network resource allocation method according to any one of claims 1 to 7 is implemented.

10. A storage medium, the storage medium comprising a stored computer program, characterized in that, When the computer program is running, control the device where the storage medium is located to execute the network resource allocation method according to any one of claims 1 to 7.