Calculation network task deployment method and system based on graph pointer network and reinforcement learning

By adopting a computing network task deployment method based on graph pointer network and reinforcement learning in a dynamic network environment, the immediate and efficient response problems of computing network task deployment in complex network environments are solved, and the optimization effects of high reception rate, high resource returns and low resource cost are achieved.

CN120104322APending Publication Date: 2025-06-06Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510172745.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In dynamic network environments, it is difficult for the existing technology to achieve immediate and efficient response to computing network tasks, especially in complex network environments. The challenge of how to rationally utilize heterogeneous resources, meet user needs and deal with complex network environments is great.

Method used

The computing network task deployment method based on graph pointer network and reinforcement learning is adopted. The graph pointer network is designed as the reinforcement learning agent through graph neural network and pointer network. The graph structure information of the computing network tasks and the underlying physical network is used to extract features and output the deployment strategy of sub-tasks, and the agent is trained through reinforcement learning algorithms.

Benefits of technology

It realizes immediate and efficient response to dynamically arrived network computing tasks in complex network environments, meets the optimization goals of high reception rate, high resource benefits and low resource costs, and improves the utilization efficiency of network resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104322A_ABST
    Figure CN120104322A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer networks, in particular to a calculation network task deployment method and system based on a graph pointer network and reinforcement learning, and the graph pointer network is designed as an intelligent agent of reinforcement learning by combining a graph neural network and a pointer network; taking the calculation network task and the graph topology data of the underlying physical network as the input of an intelligent agent, and respectively carrying out graph topology information coding and feature extraction on the input calculation network task and the underlying physical network data through a graph neural network to respectively obtain graph embedding vectors of the calculation network task and the underlying physical network; the pointer network decodes the graph embedding vector and outputs a deployment strategy of each subtask in the network calculation task; and the intelligent agent is trained through a reinforcement learning algorithm. According to the method, the calculation network task and the graph structure information of the bottom layer physical network can be fully utilized, and instant and efficient response to the dynamically reached calculation network task in a complex network environment is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer network technology, and in particular to a method and system for deploying computing network tasks based on graph pointer networks and reinforcement learning. Background Art

[0002] With the development of artificial intelligence, Internet of Things and 5G technologies, resource-intensive and delay-sensitive applications such as smart factories, drones, autonomous driving, augmented and virtual reality have emerged in large numbers, and network computing tasks have become dynamic and diverse. When a large number of network computing tasks arrive at the network, load balancing needs to be achieved to reasonably allocate tasks to available computing resources to ensure fast and efficient processing.

[0003] The network control system needs to respond quickly and effectively to the dynamically arriving computing network tasks. The computing network task deployment solution needs to reasonably utilize heterogeneous resources, meet different user needs, and be able to handle complex network environments. At the same time, it also needs to consider the user's service diversity and task dependencies. According to the characteristics of the computing network task, a task is usually represented by a combination of multiple subtasks, that is, the computing network task can be modeled as a task forwarding graph in the form of a directed acyclic graph, rather than a simple service function chain. Summary of the invention

[0004] In view of the complexity and dynamism of network computing tasks in an actual dynamic network environment, the present invention proposes a network computing task deployment method and system based on graph pointer network and reinforcement learning, which can make full use of the graph structure information of network computing tasks and underlying physical networks, and realize instant and efficient response to dynamically achieved network computing tasks in a complex network environment.

[0005] To achieve the above purpose, the technical solution adopted is:

[0006] A computing network task deployment method based on graph pointer network and reinforcement learning, comprising:

[0007] Combine graph neural network and pointer network to design graph pointer network as the intelligent agent of reinforcement learning;

[0008] The graph topology data of the computing network task and the underlying physical network are used as the input of the intelligent agent. The graph topology information of the input computing network task and the underlying physical network data are encoded respectively through the graph neural network, and the features are extracted to obtain the graph embedding vectors of the computing network task and the underlying physical network respectively; the pointer network decodes the graph embedding vector and outputs the deployment strategy of each subtask in the computing network task; and the intelligent agent is trained through the reinforcement learning algorithm.

[0009] According to the method for deploying computing network tasks based on graph pointer networks and reinforcement learning of the present invention, further, the computing network tasks are represented as directed acyclic graphs.

[0010] According to the network computing task deployment method based on graph pointer network and reinforcement learning of the present invention, the method further includes: designing a network controller, which includes a state perception module, an intelligent agent, a policy output module and a flow table; first, using the state perception module to monitor the network status in real time, and collecting network computing task information and underlying physical network information for the intelligent agent; then the intelligent agent generates an action based on the acquired perception information; finally, the action is converted into a network computing task deployment strategy through the policy output module, and the strategy is sent to the network for execution through the flow table.

[0011] According to the computing network task deployment method based on graph pointer network and reinforcement learning of the present invention, further, the pointer network of the intelligent agent adopts a pointer network based on attention mechanism, and uses the pointer network based on attention mechanism to cyclically decode and obtain the probability of subtask deployment in physical network nodes.

[0012] According to the computing network task deployment method based on graph pointer network and reinforcement learning of the present invention, further, for each cyclic decoding of the pointer network based on the attention mechanism, the graph embedding vector of the subtask and the decoder hidden state vector obtained in the previous cyclic decoding process are input into the decoder to obtain an output feature vector; based on the output feature vector, the attention mechanism is used to obtain the pointer probability of each subtask deployed on each physical network node, and the deployment strategy of the subtask is obtained through the output pointer probability; after multiple cyclic decoding, the deployment strategy of all subtasks of a computing network task is output.

[0013] According to the network computing task deployment method based on graph pointer network and reinforcement learning of the present invention, further, the state of the reinforcement learning algorithm includes real-time observation of network computing task information and underlying physical network information, the network computing task information includes network computing task topology information and the task's demand for different types of resources, and the underlying physical network information includes network topology information, the total amount of different types of resources in each physical node, utilization rate and bandwidth capacity and utilization rate of each link.

[0014] According to the network computing task deployment method based on graph pointer network and reinforcement learning of the present invention, further, the actions of the reinforcement learning algorithm include the deployment positions of all subtasks of a network computing task in the network and the calculation of forwarding paths for data flows between subtasks.

[0015] According to the computing network task deployment method based on graph pointer network and reinforcement learning of the present invention, further, the reward design of the reinforcement learning algorithm aims at high reception rate, high resource benefit and low resource cost, and the optimization target is expressed as follows:

[0016]

[0017] Among them, α and β are the optimization weights of the benefit-cost ratio and the acceptance rate, respectively. The network resource benefits brought by computing network task deployment, is the network cost brought by the computing network task deployment, and AR is the reception rate of the computing network task deployment.

[0018] Furthermore, the present invention also provides a network computing task deployment system based on graph pointer network and reinforcement learning, which is used to implement the network computing task deployment method based on graph pointer network and reinforcement learning as described above, and the system includes:

[0019] An agent building module, which is used to combine graph neural networks and pointer networks to design graph pointer networks as reinforcement learning agents;

[0020] The deployment strategy generation module is used to take the graph topology data of the computing network task and the underlying physical network as the input of the intelligent agent, encode the graph topology information of the input computing network task and the underlying physical network data respectively through the graph neural network, extract features to obtain the graph embedding vectors of the computing network task and the underlying physical network respectively; the pointer network decodes the graph embedding vector and outputs the deployment strategy of each subtask in the computing network task; and trains the intelligent agent through the reinforcement learning algorithm.

[0021] The beneficial effects achieved by adopting the above technical solution are:

[0022] The present invention is suitable for the deployment of computing network tasks in a dynamic network environment. By combining the graph neural network with the pointer network, a graph pointer network is designed as an intelligent agent for deep reinforcement learning. The intelligent agent can effectively extract the graph feature information of the underlying physical network and computing network tasks in topological form, and cyclically output effective computing network task deployment strategies in the form of pointers. The deployment strategy meets the optimization goals of high reception rate, high resource return and low resource cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings of the embodiments of the present invention, wherein the drawings are only used to illustrate some embodiments of the present invention, but not to limit all embodiments of the present invention thereto.

[0024] Figure 1 It is an overall framework diagram of a computing network task deployment method based on a graph pointer network and reinforcement learning according to an embodiment of the present invention;

[0025] Figure 2 is a schematic diagram of the internal modules of a network controller according to an embodiment of the present invention;

[0026] Figure 3 It is a schematic diagram of the structure of an intelligent agent based on a graph pointer network according to an embodiment of the present invention. DETAILED DESCRIPTION

[0027] The following will be combined with the drawings of specific embodiments of the present invention to clearly and completely describe the exemplary scheme of the embodiment of the present invention. Unless otherwise defined, the technical terms or scientific terms used in the present invention should be the common meanings understood by people with ordinary skills in the field.

[0028] This embodiment discloses a method for deploying computing network tasks based on graph pointer network and reinforcement learning. Figure 1 As shown, the following steps are included:

[0029] Step S101: Use the state perception module of the network controller to monitor the network status in real time, collect necessary underlying physical network operation status information for the intelligent agent, including but not limited to computing, storage, and bandwidth resource capacity and utilization, and collect dynamically arriving computing network task information, including but not limited to the computing and storage resource requirements of each subtask, and the data transmission bandwidth requirements between subtasks.

[0030] Step S102: Combining the graph neural network and the pointer network to design a graph pointer network as the intelligent agent of reinforcement learning, it is a complete end-to-end reinforcement learning intelligent agent network model.

[0031] Step S103: The intelligent agent uses the collected underlying physical network operation status information and network computing task information as its input. The intelligent agent generates actions based on the acquired perception information, converts the actions into specific network computing task deployment strategies through the strategy output module, and sends the strategies to the network for execution through the flow table, thus completing the specific deployment of network computing tasks in the network. Figure 2 Provides a schematic diagram of the internal modules of the network controller.

[0032] This solution represents the computing network task as a directed acyclic graph with a complex topology rather than a service function chain in a sequence form, which can be used for diverse and complex computing network service requests.

[0033] Combining GNN and pointer network, we design a graph pointer network suitable for generating computing network task deployment strategies as a reinforcement learning agent. It only needs to take the computing network task and the graph topology data of the underlying physical network as input. Its core is the pointer network that can output the computing network task deployment strategy. The agent structure based on the graph pointer network is as follows: Figure 3As shown. Specifically, two GNN-based state input encoders are designed to encode the graph topology information of the input computing network task and the underlying physical network data respectively. The graph topology data can be aggregated according to its specific node and link attribute features, and features can be effectively extracted to obtain the graph embedding vectors of the computing network task and the network respectively. According to the graph embedding vectors of the computing network task and the network, the probability of subtask deployment on the physical network node is obtained using the pointer network cyclic decoding based on the attention mechanism, so as to make an approximately optimal computing network task deployment strategy. This scheme includes but is not limited to the gated recurrent unit (GRU) as the decoder of the graph pointer network to complete the cyclic decoding of the graph embedding features.

[0034] For each cyclic decoding of the pointer network based on the attention mechanism, the graph embedding vector of the subtask and the decoder hidden state vector obtained in the previous cyclic decoding process are input into the decoder to obtain the output feature vector and the updated hidden feature vector. Based on the output feature vector, the attention mechanism is used to obtain the pointer probability of each subtask deployed on each physical network node. The deployment strategy of the subtask can be obtained through the output pointer probability. Therefore, after multiple cyclic decoding, the deployment strategy of all subtasks of a computing network task is output.

[0035] This solution uses reinforcement learning algorithms including but not limited to proximal policy optimization (PPO) to effectively train the above-built graph pointer network-based intelligent agent, which can output effective network task deployment strategies in the face of the complexity of network computing tasks and the dynamic nature of network status. The states, actions, and rewards of the Markov decision process developed for this reinforcement learning framework are as follows:

[0036] (1) Status: This includes real-time observation of randomly arriving computing network task information and underlying physical network status information. The computing network task information includes the computing network task topology information in the form of a directed acyclic graph and the task's demand for different types of resources. The underlying physical network status information includes network topology information, the total amount and utilization of different types of resources in each physical node, and the bandwidth capacity and utilization of each link.

[0037] (2) Action: The deployment locations of all subtasks of a computing task in the network. After obtaining the node deployment strategy from the agent, the network controller calculates the forwarding path for the data flow between the subtasks, where the forwarding path is obtained using but not limited to the Dijkstra shortest path algorithm.

[0038] (3) Rewards: The reward design aims at high acceptance rate, high resource benefits and low resource costs. A higher positive reward is given to actions that can achieve high resource benefits and low resource costs. In order to achieve a high acceptance rate, actions that successfully deploy tasks are encouraged, and 0 reward is returned to actions that fail to deploy computing network tasks.

[0039] The goal of this solution for network computing task deployment is to implement an effective network computing task deployment strategy. According to the dynamic arrival of diverse network computing tasks, network computing tasks can be successfully deployed, and network resource benefits can be increased and network costs can be reduced, so as to achieve effective utilization and high benefits of underlying network resources. Its optimization goal can be expressed as:

[0040]

[0041] Among them, α and β are the optimization weights of the benefit-cost ratio and the receiving rate, respectively, to balance the utility of network resources and the guarantee of computing network services. The network resource benefits brought by computing network task deployment, is the network cost brought by the computing network task deployment, and AR is the reception rate of the computing network task deployment.

[0042] Corresponding to the above method, this embodiment also proposes a computing network task deployment system based on graph pointer network and reinforcement learning, including an agent construction module and a deployment strategy generation module, wherein:

[0043] An agent building module that is used to combine graph neural networks and pointer networks to design graph pointer networks as reinforcement learning agents.

[0044] The deployment strategy generation module is used to take the graph topology data of the computing network task and the underlying physical network as the input of the intelligent agent, encode the graph topology information of the input computing network task and the underlying physical network data respectively through the graph neural network, extract features to obtain the graph embedding vectors of the computing network task and the underlying physical network respectively; the pointer network decodes the graph embedding vector and outputs the deployment strategy of each subtask in the computing network task; and trains the intelligent agent through the reinforcement learning algorithm.

[0045] The present invention proposes a computing network task deployment method based on graph pointer network and reinforcement learning, with the goals of high reception rate, high resource benefit and low resource cost. The basic idea is summarized as follows: combining graph neural network with pointer network, designing a graph pointer network as an intelligent agent, which uses graph neural network to encode graph topology data, effectively aggregates graph features on nodes and links, and thus utilizes graph information of computing network tasks and network status, and outputs computing network task deployment strategy through pointer network decoding based on attention mechanism, and effectively trains the designed intelligent agent through reinforcement learning algorithm.

[0046] Unless otherwise specifically stated, the components, steps, numerical expressions and values ​​set forth in these embodiments do not limit the scope of the present invention.

[0047] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0048] The units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A person of ordinary skill in the art may use different methods to implement the described functions for each specific application, but such implementation is not considered to be beyond the scope of the present invention.

[0049] Those skilled in the art will appreciate that all or part of the steps in the above method can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, such as a read-only memory, a disk or an optical disk. Optionally, all or part of the steps in the above embodiment can also be implemented using one or more integrated circuits, and accordingly, each module / unit in the above embodiment can be implemented in the form of hardware or in the form of software function modules. The present invention is not limited to any specific form of combination of hardware and software.

[0050] Finally, it should be noted that the above-described embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above-described embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-described embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A method for deploying computing network tasks based on graph pointer network and reinforcement learning, characterized in that: include: Combine graph neural network and pointer network to design graph pointer network as the intelligent agent of reinforcement learning; The graph topology data of the computing network task and the underlying physical network are used as the input of the intelligent agent. The graph topology information of the input computing network task and the underlying physical network data are encoded respectively through the graph neural network, and the features are extracted to obtain the graph embedding vectors of the computing network task and the underlying physical network respectively; the pointer network decodes the graph embedding vector and outputs the deployment strategy of each subtask in the computing network task; and the intelligent agent is trained through the reinforcement learning algorithm.

2. The method for deploying computing network tasks based on graph pointer network and reinforcement learning according to claim 1, characterized in that: The computing network task is represented as a directed acyclic graph.

3. The method for deploying computing network tasks based on graph pointer network and reinforcement learning according to claim 1, characterized in that: The method also includes: designing a network controller, which includes a state perception module, an intelligent agent, a policy output module and a flow table; first using the state perception module to monitor the network status in real time, and collecting network computing task information and underlying physical network information for the intelligent agent; then the intelligent agent generates an action based on the acquired perception information; finally, the action is converted into a network computing task deployment strategy through the policy output module, and the strategy is sent to the network for execution through the flow table.

4. The method for deploying computing network tasks based on graph pointer network and reinforcement learning according to claim 1, characterized in that: The pointer network of the agent adopts a pointer network based on an attention mechanism, and uses the pointer network based on the attention mechanism to cyclically decode and obtain the probability of subtask deployment on physical network nodes.

5. The method for deploying computing network tasks based on graph pointer network and reinforcement learning according to claim 4, characterized in that: For each cyclic decoding of the pointer network based on the attention mechanism, the graph embedding vector of the subtask and the decoder hidden state vector obtained in the previous cyclic decoding process are input into the decoder to obtain the output feature vector; based on the output feature vector, the attention mechanism is used to obtain the pointer probability of each subtask deployed on each physical network node, and the deployment strategy of the subtask is obtained through the output pointer probability; after multiple cyclic decoding, the deployment strategy of all subtasks of a computing network task is output.

6. The method for deploying computing network tasks based on graph pointer network and reinforcement learning according to claim 1, characterized in that: The state of the reinforcement learning algorithm includes real-time observation of the computing network task information and the underlying physical network information. The computing network task information includes the computing network task topology information and the task's demand for different types of resources. The underlying physical network information includes network topology information, the total amount of different types of resources in each physical node, utilization rate, and bandwidth capacity and utilization rate of each link.

7. The method for deploying computing network tasks based on graph pointer network and reinforcement learning according to claim 6, characterized in that: The actions of the reinforcement learning algorithm include the deployment locations of all subtasks of a computing task in the network and the calculation of forwarding paths for data flows between subtasks.

8. The method for deploying computing network tasks based on graph pointer network and reinforcement learning according to claim 7, characterized in that: The reward design of the reinforcement learning algorithm aims at high acceptance rate, high resource benefits and low resource costs. The optimization goal is expressed as follows: Among them, α and β are the optimization weights of the benefit-cost ratio and the acceptance rate, respectively. The network resource benefits brought by computing network task deployment, is the network cost brought by the computing network task deployment, and AR is the reception rate of the computing network task deployment.

9. A computing network task deployment system based on graph pointer network and reinforcement learning, characterized in that: The system is used to implement the method for deploying computing network tasks based on graph pointer network and reinforcement learning as described in any one of claims 1 to 8, and comprises: An agent building module, which is used to combine graph neural networks and pointer networks to design graph pointer networks as reinforcement learning agents; The deployment strategy generation module is used to take the graph topology data of the computing network task and the underlying physical network as the input of the intelligent agent, encode the graph topology information of the input computing network task and the underlying physical network data respectively through the graph neural network, extract features to obtain the graph embedding vectors of the computing network task and the underlying physical network respectively; the pointer network decodes the graph embedding vector and outputs the deployment strategy of each subtask in the computing network task; and trains the intelligent agent through the reinforcement learning algorithm.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.