An Intelligent Joint Optimization Method for Resources Based on Deep Reinforcement Learning

By applying the resource intelligent joint optimization method of deep reinforcement learning in complex communication networks, the accuracy and real-time problems of resource allocation under dynamic situations are solved, and more efficient resource utilization and lower computing time consumption are achieved.

CN119854204BActive Publication Date: 2025-06-24NANJING UNIV OF INFORMATION SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510315710.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-24
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

The prior art is difficult to effectively manage and allocate communication and computing resources in complex communication networks, especially in dynamic situations and unpredictable environments, resulting in insufficient accuracy and real-time decision-making.

Method used

Using the resource intelligent joint optimization method based on deep reinforcement learning, by building a communication network topology, detecting data traffic and resource usage, randomly generating task flows, finding the optimal path, and allocating computing and communication resources on the basis of the D3QN algorithm to minimize task delay.

Benefits of technology

On the premise of ensuring the task delay requirements, more concurrent tasks can be effectively accommodated, resource utilization rate can be maximized, calculation time consumption can be reduced, and resource allocation accuracy and real-time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119854204B_ABST
    Figure CN119854204B_ABST
Patent Text Reader

Abstract

The present invention discloses a resource intelligent joint optimization method based on deep reinforcement learning, and the method comprises the following steps: Step 1: Build a communication network topology and detect data traffic in the network topology; Step 2: Randomly generate task flows and simulate the parsing and sending of tasks in the network topology; Step 3: The switch receives the tasks and finds n simple paths of the tasks; Step 4: Set constraint conditions and screen candidate paths from the n simple paths; Step 5: With the goal of minimizing the total delay, select the optimal path from the candidate paths and allocate the optimal computer resources and communication resources for the optimal path; Step 6: Send the communication resources allocated to each link of the optimal path to the switch to implement communication resource allocation. The method proposed by the present invention can effectively accommodate more concurrent tasks on the premise of ensuring the task delay requirements, thereby maximizing the resource utilization rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of communication technologies, and particularly relates to an intelligent joint optimization method for resources based on deep reinforcement learning. Background Art

[0002] In today's complex and ever-changing communication environment, communication networks not only need to efficiently manage and allocate communication resources, but also must consider the joint allocation of computing resources. The dynamic situation in the communication environment is often unpredictable, which poses higher requirements for decision-makers. They need to make quick and accurate judgments in the rapidly changing situation and formulate corresponding strategies. However, relying solely on the wisdom and experience of decision-makers is difficult to ensure the accuracy of decisions. Therefore, introducing a computer-aided decision-making system has become an urgent problem to be solved.

[0003] In recent years, with the rapid development of artificial intelligence technologies, researchers have begun to explore how to assist decision-makers in making decisions through intelligent algorithms. This can not only improve the scientificity and accuracy of decisions, but also effectively reduce the consumption of human resources. In this context, the joint allocation of communication and computing resources is particularly important. At present, there are mainly the following three common methods: game theory methods, heuristic algorithms, and artificial intelligence methods.

[0004] Game theory, as a mathematical tool for studying the interaction and competition relationships among decision-makers, has been widely applied to the resource allocation problem in the communication field. In recent years, with the development of information technologies, game theory has been gradually introduced into resource allocation. Especially in the context of information-based communication, game theory is used to analyze the resource allocation and strategy selection of both parties. The prior art constructs a new multi-leader single-follower Stackelberg game model for the shortage problems of communication and computing resources, aiming to minimize the task processing delay and cost. At the same time, a resource allocation algorithm based on selection optimization is proposed, and the Karush-Kuhn-Tucker algorithm is used to jointly optimize the task allocation and bandwidth allocation ratio to minimize the task processing delay. In addition, the binary search method is used to refine the optimization of computing resources, aiming to minimize the total cost on the premise of ensuring that the task delay limit is satisfied. However, the envisioned scenario is relatively simple, and its effectiveness in complex communication networks cannot be fully verified. There is also the use of game theory and Lagrangian function to solve the optimization problem of minimizing the energy consumption of all users, and an optimal resource allocation scheme is obtained. However, the application scenario is also relatively ideal, without considering the impact caused by factors such as unexpected events that may be encountered, and the results lack practicality.

[0005] Heuristic algorithms have shown good performance in solving complex optimization problems due to their flexibility and efficiency, especially in the dynamic allocation of communication network resources. In recent years, scholars at home and abroad have proposed a variety of heuristic methods for problems such as task scheduling and spectrum allocation, such as genetic algorithms, particle swarm optimization, etc. These algorithms can not only quickly solve approximate optimal solutions, but also maintain good robustness under uncertain conditions. Aiming at the problem that in an uncertain network, due to the randomness of node movement and the time-varying nature of wireless channel states, uncertain network environment characteristics such as uncertain queuing delays and network device connection times appear, which greatly affect the utilization rate of computing and communication resources. With the goal of minimizing the total system energy consumption, a multi-stage stochastic programming optimization algorithm for joint resource allocation based on stochastic simulation (SS-MSSP) is proposed. The genetic algorithm is used to obtain the optimal allocation strategy of transmission power and computing resources. However, this method only considers modeling random dynamic factors from the time dimension and has not considered other dimensions, making it difficult to adapt to more application scenarios. There is also a new graph-based workflow strategy proposed for the problem that most of the existing strategies do not consider the dependencies between computing tasks, or are usually based on heuristic search algorithms, which are unacceptable for delay-sensitive services. This strategy can handle complex workflow applications with non-linear structures. At the same time, using graph-based partitioning techniques, it finds the decision-making plan with the lowest energy consumption under delay constraints, without considering the allocation problem of multi-hop bandwidth.

[0006] The development of artificial intelligence technology has provided new ideas for communication network decision-making, especially showing strong potential in resource allocation and management. In recent years, AI technologies such as deep learning and reinforcement learning have received increasing attention in the application of communication networks. Some research institutions at home and abroad have begun to explore how to integrate AI systems into communication networks to improve response speed and allocation quality. Research on allocation strategies based on artificial intelligence has been carried out, attempting to achieve intelligent allocation and dynamic adjustment of resources through machine learning algorithms, aiming to improve overall communication efficiency. For example, for the problems of insufficient node computing power and uneven bandwidth allocation in the computing power network, a deep reinforcement learning algorithm based on an incentive mechanism combined with proximal policy optimization (PPO) is proposed. Under the constraints of computing resources and transmission power, a distribution plan that maximizes the system revenue is calculated by jointly optimizing computing resources and transmission power. However, this plan only considers the overall revenue and does not consider the needs of delay-sensitive tasks. There are also solutions that use deep Q networks to solve bandwidth allocation problems and deep deterministic policy gradient algorithms to solve continuous power allocation problems. A meta-based DRL algorithm is also proposed to enhance the rapid adaptability of resource allocation strategies in dynamic environments. However, the problem of complex and variable link bandwidth caused by different transmission methods is not considered. In addition, existing technologies construct an optimization problem of delay and resource consumption through mathematical modeling and describe it as a Markov decision process. By prioritized experience replay, the traditional soft behavior critic deep reinforcement learning framework is improved, and a resource allocation algorithm based on SACPER is proposed to minimize the total cost of the system. However, this method does not consider the scenario of a large number of user devices accessing simultaneously, and the designed solution does not estimate the constraints of dependencies between tasks. Summary of the Invention

[0007] Object of the Invention: To solve the problems existing in the above-mentioned prior art, the present invention provides an intelligent joint optimization method for resources based on deep reinforcement learning.

[0008] Technical Solution: The present invention discloses an intelligent joint optimization method for resources based on deep reinforcement learning, and the method includes the following steps:

[0009] Step 1: Build a communication network topology, detect the data traffic in the network topology, and record the bandwidth, remaining bandwidth of the link, and the usage of computing resources of each communication node;

[0010] Step 2: Randomly generate a task flow T, and simulate the parsing and sending of tasks in the network topology;

[0011] Step 3: The switch receives the tasks in the task flow After parsing the data packet header, query whether there is a corresponding flow rule for the data packet in its own flow table. If there is a corresponding flow rule, directly forward the task to the corresponding port set in the flow rule; otherwise, find n simple paths for the task.

[0012] Step 4: Based on the task The minimum required transmission rate and the nodes and The link between them is the transmission rate allocated for the task , set the constraint conditions to filter candidate paths from the n simple paths; nodes i and j are the nodes at both ends of the link, i = 1, 2,..., n; j = 1, 2,..., n, and n represents the total number of nodes;

[0013] Step 5: With the goal of minimizing the total delay, select the optimal path from the candidate paths and allocate the optimal computing resources and communication resources for the optimal path;

[0014] Step 6: Send the communication resources allocated to each link of the optimal path to the switch to implement communication resource allocation.

[0015] Furthermore, the calculation formula for the remaining bandwidth free_bw in Step 1 is as follows:

[0016] free_bw = capability – speed

[0017] where capability represents the total bandwidth of the link and speed is the flow rate.

[0018] Furthermore, Step 3 uses the breadth-first algorithm to find n simple paths.

[0019] Furthermore, the constraint conditions in Step 4 are:

[0020] ;

[0021] where is the available transmission rate of the link between nodes and ; is the computing resource allocated for the task by node , is the available computing resource of node , The expression of

[0022] ;

[0023] where is the node and The link between them is the bandwidth allocated for the task , is the node and The link between them is the transmission power allocated for the task ; represents the background noise of the channel

[0024] Furthermore, in step 5, the D3QN algorithm is used to select the optimal path from the candidate paths and allocate the optimal computing resources and communication resources for the optimal path. Specifically:

[0025] Set the objective function: Set a leaky bucket controller with a capacity of b, and the fluid leaks out at a rate of units per second; the peak rate of the fluid is , and use the leaky bucket controller to calculate the end-to-end delay bound of the task :

[0026] ;

[0027] Among them, is the size of the largest packet in the task ; represents that if , , otherwise , represents the length of the largest data packet among all task flows in the router, represents the length of the largest data packet in the leaky bucket controller; represents the rate at which the node sends to the node ;

[0028] Based on the end-to-end delay bound of the task calculate the total delay of the task transmission, and take the minimum total delay as the objective function:

[0029] ;

[0030] Among them, is the total delay of the task transmission, is the delay generated by parsing the task , The expression of

[0031] ;

[0032] Among them, is the node for the task Allocated computing resources for the node of the number of cycles for the task is the total amount of data to be computed;

[0033] Set the state space as:

[0034] ;

[0035] where h represents the current moment, M represents the sequence of computing data volumes of the task, M , where represents the computing data volume of the task , X represents the total number of tasks, represents the sequence of minimum required transmission rates of the tasks, , represents the sequence of maximum allowable end - to - end delays of the tasks, , represents the maximum allowable end - to - end delay of the task ; represents the sequence of available computing resources of the node at the current moment, , represents the available computing resource of the i - th node at the current moment; represents the sequence of available transmission rates, , for the node and is the available transmission rate of the link between them;

[0036] Set the action space as:

[0037] ;

[0038] The agent selects a value from the discrete numerical set and assigns it to , Y represents the total number of discrete values, represents the preset parameter, represents the allocation of communication resources. When represents solving for communication resources using a non - uniform bandwidth allocation algorithm, represents solving for communication resources using a uniform allocation algorithm;

[0039] Set the reward function:

[0040] .

[0041] Further, the agent adopts a greedy learning strategy to select a value from a discrete numerical set and assign it to .

[0042] A computer device includes a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the processor executes the computer program, the resource intelligent joint optimization method is implemented.

[0043] A computer-readable storage medium is used to store a program, and the program is executed to implement the resource intelligent joint optimization method.

[0044] Beneficial effects: From the perspective of intelligent resource allocation for communication networks, the present invention proposes a resource intelligent joint optimization method based on deep reinforcement learning. On the premise of ensuring the task delay requirements, it can effectively accommodate more concurrent tasks, thereby maximizing resource utilization and reducing computational time consumption, and effectively improving the accuracy and real-time performance of resource allocation. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 is a flowchart of a resource intelligent joint optimization algorithm based on deep reinforcement learning;

[0046] Figure 2 is a schematic diagram of the architecture of a communication network resource joint allocation system;

[0047] Figure 3 is a flowchart of randomly generating a task flow;

[0048] Figure 4 is a diagram of issuing a meter table using the OpenFlow protocol;

[0049] Figure 5 is a comparison chart of the computational time overhead required for each method;

[0050] Figure 6 is a comparison chart of the success rates of each method. DETAILED DESCRIPTION OF THE INVENTION

[0051] The accompanying drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.

[0052] As Figure 1 shown, the present invention provides a resource intelligent joint optimization method based on deep reinforcement learning, including the following steps:

[0053] Step1: Build the communication network resource allocation system architecture, construct the communication network topology, use SDN technology to detect the data traffic in the network topology, and record the bandwidth of the links and the usage of computing resources of each communication node for the Double Dueling Deep Q-Network (D3QN) algorithm;

[0054] Step2: Randomly generate a task flow T in the network topology to simulate the parsing and sending of tasks;

[0055] Step3: Use the routing algorithm to find n shortest paths from node to node;

[0056] Step4: Based on the n shortest paths found, judge whether the available computing resources in the path are not empty and whether the available communication resources can meet the minimum transmission rate requirement of the task. If not, discard the path and judge the paths in sequence. Find the candidate paths. If not found, it is considered that the task transmission fails;

[0057] Step5: The agent selects the optimal computing resource allocation plan and the most suitable communication resource allocation plan in the current network resource environment;

[0058] Step6: Use the meter table in OpenFlow to issue the communication resources allocated to each link of the optimal path to the switch to achieve communication resource allocation.

[0059] In step 1, the present invention first constructs the communication network resource joint allocation system architecture.

[0060] The communication network resource joint allocation system architecture is as Figure 2 shown. The core function of the system is to obtain the data traffic in the network topology in real time through the data acquisition module and calculate the available bandwidth of each link. This available bandwidth information is stored in the controller in the form of a graph (graph). At the same time, the controller also stores the computing resource data of all nodes internally and can dynamically adjust the resource allocation according to task changes. The data acquisition module regularly issues commands to the switch through the SDN controller to obtain the statistical information of the switch port traffic, so as to ensure that the system continuously monitors the network status and provides an accurate traffic basis. After obtaining the traffic statistical information, the computing and storage module is responsible for calculating the remaining available bandwidth of the link, comprehensively considering the current traffic and past trends to ensure the effectiveness of resource allocation. Two formulas are needed for calculation:

[0061] Flow velocity formula: speed = (s(t1)-s(t0)) / (t1-t0).

[0062] In the formula: s(t1) represents the traffic at time t1, and s(t0) represents the traffic at time t0.

[0063] Remaining bandwidth formula: free_bw = capability – speed.

[0064] Where: capability represents the total link bandwidth.

[0065] In step 2, two different hosts are randomly selected, and one of the three set task types is selected as the generated task and sent through Iperf. Iperf is a network performance testing tool that can test the maximum TCP and UDP bandwidth performance. Iperf is implemented based on the Server-Client mode. When measuring network parameters, Iperf differentiates between two roles: the listener and the speaker. The speaker sends a certain amount of data to the listener, and the listener counts and records parameters such as the speaker's bandwidth usage and delay jitter. After all the speaker's data is sent, the listener sends a data packet to the speaker to inform the speaker of the measured data.

[0066] The flowchart of step 2 is as Figure 3 shown, specifically:

[0067] (1) Randomly generate a host IP for the Client side as the speaker in Iperf, responsible for sending data to the listener;

[0068] (2) Randomly generate a host IP for the Server side as the listener in Iperf, responsible for counting the data sent by the listener and recording parameters such as bandwidth and delay jitter;

[0069] (3) Determine whether the host IPs generated in steps (1) and (2) are the same. If they are the same, go to step (2); otherwise, go to step (4);

[0070] (4) Randomly select one of multiple types of tasks as the data parameter to be sent;

[0071] (5) Determine whether the total amount of generated tasks reaches the condition. If it reaches, end this process; otherwise, go to step (1).

[0072] Examples of the generated task types are shown in Table 1 below:

[0073] Table 1

[0074] Task type Total data volume of computing tasks Minimum required transmission rate Maximum end-to-end delay 1-1 4Mbits 4Mbits 400ms 1-2 2Mbits 4Mbits 400ms 1-3 1Mbits 4Mbits 400ms 2-1 4Mbits 2Mbits 25ms 2-2 2Mbits 2Mbits 25ms 2-3 1Mbits 2Mbits 25ms 3-1 4Mbits 1Mbits 15ms 3-2 2Mbits 1Mbits 15ms 3-3 1Mbits 1Mbits 15ms

[0075] In step 3, after the switch receives the data packet and parses the packet header, it queries whether there is a corresponding flow rule for the data packet in its own flow table. By matching the flow rule corresponding to the data packet, it mainly compares the IP addresses of the source host and the destination host, as well as the computing resources, transmitted bandwidth, and delay parameters. If the match is successful, the data packet is directly forwarded to the corresponding port set in the flow rule; if the match fails, a Packet-in event is generated according to the corresponding packet, and the Packet-in data packet is sent to the controller. After receiving the Packet-in data packet, the controller uses the breadth-first algorithm to find n simple paths for the task end-to-end.

[0076] In step 4, in the communication network resource joint optimization environment, the network topology can be stored as a graph set , where represents the edge set, O represents the total number of edges, represents the Oth edge, represents the vertex set, n represents the total number of nodes, represents the nth node, and the task set is represented as , X represents the total number of tasks; since resource joint optimization needs to consider both the communication resources and computing resources required by the task, each task can be represented by a triple as =( , , ), where is the total amount of data that task needs to calculate, is the minimum available transmission rate required by task , is the maximum end-to-end delay that task can tolerate.

[0077] Using the Shannon formula to simplify the bandwidth and power resources in the communication network, the mathematical problem is modeled as the following formula:

[0078] ;

[0079] where, is the transmission power allocated to task and for the link between node , is the bandwidth allocated to task and for the link between node ; is the transmission rate allocated to task and for the link between node ; To represent the background noise of the channel, the bandwidth resources and power resources in the communication network can be simplified from a two-dimensional problem to a one-dimensional transmission rate problem through the above formula.

[0080] In a communication network, all tasks issued need to go through a series of computational analysis processes before being sent to the designated node. This process is not only an important link to ensure the correct execution of tasks, but also leads to the consumption of computing resources and the generation of time delays.

[0081] To achieve efficient task processing, it is necessary to reasonably allocate computing resources. In the case of multi-task parallel processing, it is necessary to ensure that as many tasks as possible can be accepted and processed. In addition, the time delay generated by each task needs to be controlled within an acceptable range, so as to minimize the impact on the time delay of the entire task link. This is crucial for the real-time performance and reliability of the communication network. The present invention models the time delay generated by the parsing of computing tasks, as shown in the formula:

[0082] ;

[0083] where is the computing resource allocated to node for task , is the number of cycles of node , is the total amount of data that task needs to calculate; is the time delay generated by the parsing of computing task .

[0084] The constraint conditions for selecting a path can be obtained through the above two formulas:

[0085] ;

[0086] where is the available transmission rate of the link between node and , and nodes i and j are the nodes at both ends of the link; is the available computing resource of node ; i = 1, 2,..., n; j = 1, 2,..., n. and ensure that the transmission rate of the link between nodes and is neither less than the minimum transmission rate required by the task nor exceeds the available transmission rate of the link. and ensure that node The available computing resources allocated for a task can neither be zero nor exceed the maximum available computing resources of the current node.

[0087] In step 5, it is assumed that the variable-bit data stream in packet switching passes through a shaper composed of a leaky bucket, and the shaped data stream conforms to an arrival curve of , also known as the traffic specification - traffic flow T-SPEC, where h represents time, is the length of the largest packet in the vulnerability, is the peak rate, is the sustainable rate, is the burst tolerance. In the terminology of the integrated service Internet, this quadruple is the traffic characteristic specification adopted by the Internet Engineering Task Force, and it is widely used to describe the traffic characteristics of services in the integrated service network. According to network calculus, the definition of the leaky bucket controller is as follows:

[0088] Definition 1 The leaky bucket controller is used to define a standard model that conforms to the traffic in the integrated service network, and it is a device that analyzes the flow data in the following manner . Suppose there is a fluid bucket (pool) with a capacity of , and the initial state of the bucket is empty. There is a small hole at the bottom of the bucket, and when the bucket is not empty, the fluid leaks out at a rate of units per second. The data from the flow must pour into the bucket a fluid equal in volume to the data volume. The data that causes the leaky bucket to overflow is called non-conforming data, otherwise it is called conforming data. The fluid in the leaky bucket model does not represent data, however, it uses the same measurement unit as the data. A flow constrained by a leaky bucket, satisfying T-SPEC , when passing through a node that provides a rate-latency service curve, the delay bound of the flow is shown as follows:

[0089] ;

[0090] where, represents the service rate provided by the node to the data stream, represents the latency existing in the system. According to the integrated service model of Internet routers, the service curve provided by a router that guarantees the rate to a flow is a rate-latency function, and the rate and the latency have the following relationship as shown below:

[0091] ;

[0092] where, is the length of the largest packet of the flow, is the length of the largest packet of all flows in the router, is the total rate of the scheduler.

[0093] Obtain the end-to-end delay bound as follows:

[0094] ;

[0095] wherein, denotes that if , , otherwise .

[0096] Therefore, the end-to-end delay bound of task is obtained through the above formula as follows:

[0097] ;

[0098] wherein, denotes the rate at which node sends to node , denotes the minimum transmission rate allocated for task among all links;

[0099] In a communication network environment, a task usually includes data parsing and information transmission. The parsing task consumes a certain amount of computing resources, while the sending task has specific requirements for communication resources. The two are interdependent and have a profound impact on each other. Therefore, when optimizing resources, it is necessary to comprehensively evaluate the requirements of these two aspects to achieve the optimization of the overall performance. Since the delays in the parsing and sending processes will be superimposed on each other, resulting in an increase in the response time of the final task, the total delay becomes the core objective of the optimization of the present invention, as shown in the following formula:

[0100] .

[0101] The SDN controller is deployed as an Agent to globally control the resources in the entire communication network and dynamically formulate resource allocation schemes for each task. The command node generates some tasks based on the current environmental situation, and the information required for these tasks is sent to the Agent as the state information s. The Agent calculates and allocates computing and communication resources according to the state information. The allocation scheme is a continuous process with consistency. First, Y action strategies are adopted for the allocation of computing resources; in this embodiment, Y = 10, and these Y action strategies are respectively 10 discrete computing resource values. Then, 2 action strategies are adopted for the allocation of communication resources, namely the non-uniform bandwidth reservation algorithm and the uniform bandwidth reservation algorithm, and a continuous communication resource allocation scheme is obtained through the selected algorithm. The Agent returns the scheme generated by the action strategy to the node; the node will return a reward r to the Agent. In order to evaluate the quality of the current action strategy, the Agent and the environmental state interact continuously, and the optimal resource allocation strategy is obtained by comparing the cumulative rewards between different strategies and returned to the node.

[0102] The settings of the state space, action space, and reward function in the D3QN algorithm are as follows:

[0103] (1) State space: In the definition of the state space, the Agent needs to select corresponding action strategies according to the obtained environmental state information. Therefore, two parts of the environmental state need to be considered comprehensively. The first part is the characteristics of the tasks, that is, the Agent needs to obtain the task set , the total amount of data calculated by each task , the minimum required transmission rate sequence and the maximum end-to-end delay allowed by the task , where represents the amount of computing data of task . The second part is the current environmental information of the communication network, that is, at moment, all available computing resources of all nodes in the entire network and the available transmission rate in the link.

[0104] The settings of the state space are as follows:

[0105] ;

[0106] Action space: In the definition of the action space, the Agent will have its own action set. Whenever a task is generated, the Agent will judge according to the current environmental state and select a corresponding action strategy. After executing the corresponding action strategy, it will enter a new environmental state and obtain the reward feedback of the corresponding action strategy. In order to improve the convergence speed and obtain better action strategies faster, the present invention introduces a greedy learning strategy and adopts a greedy factor When the random number value between 0 and 1 is less than , a random value will be taken from the action space. When it is greater than or equal to , the action with the highest Q value of the current environmental state will be adopted. The present invention aims to jointly optimize the allocation of computing resources and communication resources. Therefore, there are two parts of value-taking schemes in the action space. The first part is Y discrete values of the number of CPU cycles of computing resources. The agent selects a value from for the current task, is a preset parameter and serves as the action strategy for the computing resources of the current task; the second part is two value-taking schemes for communication resources , where represents solving by using a non-uniform bandwidth allocation algorithm, represents solving by using a uniform allocation algorithm.

[0107] The setting of the action space is as follows:

[0108]

[0109] Reward function: The reward function is the feedback signal obtained by the agent from the environment after selecting and executing a certain action strategy. This signal is used to evaluate the effectiveness of the strategy and the correctness of the execution. Therefore, the design of the reward function must be closely related to the optimization goal. The proposed optimization goal needs to be able to ensure that all relevant tasks can be completed within the tolerable delay constraint under limited constraints, and at the same time, as many concurrent tasks as possible can be accommodated. To achieve this optimization goal, the reward function should reflect several key factors. First of all, the timeliness of the task is very important. Therefore, if the task can be completed within its maximum tolerable delay , a positive reward should be provided to the agent. On the contrary, if the task fails to be completed on time, a negative penalty should be imposed to prompt the agent to seek more efficient decisions.

[0110] The setting of the reward function is as follows. This design concept not only enables the agent to make effective decisions in a dynamic environment but also can continuously optimize the resource allocation strategy to meet the changing task requirements and network states. Through a reasonable reward mechanism, the agent will be more motivated to explore the optimal action strategy, thereby improving the overall system performance.

[0111] .

[0112] In step 6, as Figure 4 shown, the communication resources allocated to each link of the optimal path use the meter table in OpenFlow, and the Drop attribute is selected in the action and sent to the switch to implement the allocation of communication resources. The specific flow table is shown in Table 2 below, and the meter table is shown in Table 3 below:

[0113] Table 2

[0114] Source address 10.0.0.1 Destination address 10.0.0.4 Input port 1 Instruction Meter serial number: 1, Output port: 4

[0115] Table 3

[0116] Meter serial number 1 Type Discard Rate 2000 bits per second

[0117] Figure 5 For the bar chart comparing the time consumption calculated by multiple methods, in this embodiment, the deep deterministic policy gradient algorithm (DDPG), the double deep Q-network algorithm (DDQN), and the multi-agent deep reinforcement learning algorithm (MADRL) are used as the comparison methods in this embodiment; the abscissa represents the number of nodes of different topologies, and the ordinate represents the time spent on the calculation of different algorithms, with the unit of ms. The same number of task flows are sent in different network topologies. Each time the controller uses an agent to solve the optimal calculation resource and communication resource allocation scheme, the time spent is recorded. Finally, the average value of the time spent by different methods is obtained to get this graph. It can be seen from the graph that the algorithm proposed in the present invention always has a lower calculation time overhead because the present invention can select the optimal calculation resource and communication resource allocation scheme faster.

[0118] Figure 6 For the bar chart comparing the number of successfully allocated tasks in multiple network topologies using different methods. The abscissa represents the number of nodes of different topologies, and the ordinate represents the number of tasks successfully allocated by the algorithm. The same number of task flows are sent in different network topologies. Each time the controller successfully allocates the required resources for a task, it is recorded as 1. Finally, the average value of the total number of successfully allocated tasks in different network topologies is obtained to get this graph. It can be seen from the graph that the number of successfully completed tasks of all algorithms is proportional to the scale of the topology nodes. This is because as the scale of the topology nodes increases, the number of nodes inside increases, and the corresponding number of links also increases. The result is that the number of available calculation resource and communication resource allocation schemes also increases. From Figure 6 it can be seen that in different topology networks, the D3QN algorithm proposed in this embodiment is always optimal compared with the other three algorithms, so it has better resource utilization and better communication service quality. Compared with the DDQN algorithm, which is also a deep reinforcement learning algorithm, D3QN introduces the value function and the action value function, making the learning of the agent more stable, enabling the Q value of the evaluation network to be closer to the Q value calculated by the target network, and thus being able to solve a better strategy.

[0119] In summary, from the perspective of intelligent resource allocation in a communication network, this embodiment analyzes the advantages and disadvantages of several resource joint allocation algorithms that are commonly used today. On this basis, a resource intelligent joint optimization method based on deep reinforcement learning is proposed, which solves some of the disadvantages of the current resource joint allocation algorithms to a certain extent. On the premise of ensuring the task delay requirements, it can effectively accommodate more concurrent tasks, thereby maximizing the resource utilization rate and having a lower computational time overhead for the algorithm, effectively improving the accuracy and real-time performance of resource allocation. Therefore, the resource intelligent joint optimization method based on deep reinforcement learning proposed by the present invention is feasible and effective.

[0120] In addition, it should be noted that, in the above specific embodiments, the various specific technical features described can be combined in any suitable manner without conflict. To avoid unnecessary repetition, the present invention does not separately describe various possible combination methods.

Claims

1. A resource intelligent joint optimization method based on deep reinforcement learning, characterized in that: The method comprises the following steps: Step 1: Build a communication network topology, detect the data traffic in the network topology, record the link bandwidth, remaining bandwidth, and usage of computing resources of each communication node; Step 2: Randomly generate task flow T and simulate task parsing and sending in the network topology; Step 3: The switch receives the task in the task flow After parsing the data packet header, check whether there is a corresponding flow rule for the data packet in its own flow table. If there is a corresponding flow rule, forward the task directly to the corresponding port set in the flow rule, otherwise find n simple paths for the task; Step 4: Task-based Minimum required transfer rate and nodes and The link between the tasks Assigned transfer rate , set constraints to select candidate paths from n simple paths; nodes i and j are nodes at both ends of the link, i=1,2,...,n; j=1,2,...,n, n represents the total number of nodes; Step 5: With the goal of minimizing the total delay, select the optimal path from the candidate paths and allocate optimal computing resources and communication resources to the optimal path; Step 6: Send the communication resources allocated to each link of the optimal path to the switch to implement communication resource allocation; The constraints in step 4 are: ; in, For Node and The available transmission rate of the link between For Node For the task Allocated computing resources, For Node Available computing resources, The expression is: ; in, For Node and The link between the tasks The allocated bandwidth, For Node and The link between the tasks the allocated transmission power; Represents the background noise of the channel.

2. According to claim 1, a resource intelligent joint optimization method based on deep reinforcement learning is characterized in that: The calculation formula for the remaining bandwidth free_bw in step 1 is as follows: free_bw = capability – speed; Among them, capability represents the total link bandwidth, and speed represents the flow rate.

3. The resource intelligent joint optimization method based on deep reinforcement learning according to claim 1 is characterized in that: Step 3 uses the breadth-first algorithm to find n simple paths.

4. According to claim 1, a resource intelligent joint optimization method based on deep reinforcement learning is characterized in that: Step 5 uses the D3QN algorithm to select the optimal path from the candidate paths and allocates optimal computer resources and communication resources to the optimal path, specifically: Set the objective function: Set a leaky bucket controller with a capacity of b and a flow rate of b per second. The peak velocity of the fluid is , using leaky bucket controller to calculate tasks The end-to-end delay bound : ; in, For the task The size of the largest packet in; If ,but ,otherwise , Indicates the length of the largest data packet in all task flows in the router. Indicates the length of the maximum data packet in the leaky bucket controller; Representation Node Send to node rate; Indicates that all links are tasks The minimum assigned transmission rate; Task-based The end-to-end delay bound Calculate the total delay of task transmission and take the minimum total delay as the objective function: ; in, is the total delay of task transmission, For analysis tasks The delay generated, The expression is: ; in, For Node For the task Allocated computing resources, For Node of Number of cycles, For the task The total amount of data that needs to be calculated; Setting up the state space for: ; Among them, h represents the current moment, M represents the computational data sequence of the task, and M ,in Indicates the task The amount of computing data, X represents the total number of tasks, represents the minimum required transmission rate sequence of the task, , represents the maximum end-to-end delay sequence allowed by the task, , Indicates the task The maximum allowed end-to-end delay; Represents the sequence of available computing resources of the node at the current moment, , Indicates the available computing resources of the i-th node at the current moment; Represents a sequence of available transmission rates, , For Node and The available transmission rate of the link between Setting up the action space : ; Agents on discrete numerical sets Select a value from the list and assign it to , Y represents the total number of discrete values, Indicates the preset parameters. Indicates the allocation of communication resources. It represents the use of non-uniform bandwidth allocation algorithm to solve communication resources. The representative adopts the uniform allocation algorithm to solve the communication resources; Set up the reward function: 。 5. According to claim 4, a resource intelligent joint optimization method based on deep reinforcement learning is characterized in that: The agent uses a greedy learning strategy to select a value from a discrete value set and assign it to .

6. A computer device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: When the processor executes the computer program, it implements a resource intelligent joint optimization method based on deep reinforcement learning as described in any one of claims 1 to 5.

7. A computer-readable storage medium for storing a program, characterized in that: Execute the program to implement a resource intelligent joint optimization method based on deep reinforcement learning as described in any one of claims 1-5.