Computational graph division method and device, electronic equipment and storage medium

Through the method of generating and training agents, the communication overhead and power consumption problems of large-scale neural networks on DNN accelerators are solved, and low-cost computing graph division is realized, which is suitable for neural networks of any structure.

CN120338023APending Publication Date: 2025-07-18TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410064591.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When the prior art dispatches deep neural networks to DNN accelerators, traditional computing graph division methods cannot effectively reduce communication overhead and power consumption, and are especially difficult to apply in large-scale neural networks.

Method used

By generating a set of candidate division schemes for the computational graph, the agent is trained, and the feature vector is extracted using the graph neural network, the action prediction results are determined, and the parameters are updated according to the loss function, and a potential division scheme with lower cost is generated to realize self-iteration training.

Benefits of technology

The trained agent can divide the calculation graph of neural networks of any structure, reduce the number of external memory accesses, delay and power consumption, and provide a lower cost division solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338023A_ABST
    Figure CN120338023A_ABST
Patent Text Reader

Abstract

The invention relates to a calculation graph division method and device, electronic equipment and a storage medium. The method comprises the steps that a candidate division scheme set of the computational graph is generated, and the candidate division scheme set comprises a plurality of candidate division schemes; adopting the candidate division scheme set to train an intelligent agent used for calculation graph division; generating a potential division scheme of the computational graph through the intelligent agent after parameter updating; and in response to the situation that the cost of the potential division scheme is lower than that of any candidate division scheme in the candidate division scheme set, replacing the candidate division scheme with the potential division scheme so as to train the intelligent agent by adopting the updated candidate division scheme set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a training method for an agent for computing graph partitioning, a computing graph partitioning method, a training device for an agent for computing graph partitioning, a computing graph partitioning device, an electronic device, and a storage medium. Background Art

[0002] With the development of DNN (Deep Neural Network), DNNs with different structures and scales have been proposed. When we schedule a DNN onto a DNN accelerator, due to storage resource limitations, it is usually necessary to divide the workload of the DNN into multiple subtasks. Traditional DNN accelerator applications process neural networks layer by layer at the application level. It stores the intermediate data of each layer in off-chip memory, which incurs huge communication overhead and consumes a large amount of power.

[0003] The layer fusion scheme fuses multiple layers in the same task to reuse data across different layers, thereby effectively reducing the data transfer between on-chip and off-chip memories. To adopt the layer fusion scheme, we need to divide the computing graph into several subgraphs as basic tasks. Different partitioning schemes of the computing graph may have a great impact on the scheduling overhead on the DNN accelerator.

[0004] In related technologies, an exhaustive approach is used to find the partitioning scheme for layer fusion, but this method can only be applied to neural networks with fewer layers (such as AlexNet, VGG, etc.) and cannot be applied to large-scale neural networks. Summary of the Invention

[0005] The present disclosure provides a technical solution for computing graph partitioning.

[0006] According to an aspect of the present disclosure, there is provided a training method for an agent for computing graph partitioning, including:

[0007] generating a set of candidate partitioning schemes for a computing graph, where the set of candidate partitioning schemes includes multiple candidate partitioning schemes;

[0008] training an agent for computing graph partitioning using the set of candidate partitioning schemes;

[0009] generating a potential partitioning scheme for the computing graph through the agent after parameter update;

[0010] in response to the cost of the potential partitioning scheme being lower than any candidate partitioning scheme in the set of candidate partitioning schemes, replacing the candidate partitioning scheme with the potential partitioning scheme to train the agent using the updated set of candidate partitioning schemes.

[0011] In a possible implementation, training the agent for computing graph partitioning by using the candidate partitioning scheme set includes:

[0012] Select a target candidate partitioning scheme from the candidate partitioning scheme set;

[0013] Select a target timestamp in the target candidate partitioning scheme;

[0014] Obtain a target state-action pair corresponding to the target timestamp;

[0015] Input the state information in the target state-action pair into the agent for computing graph partitioning, and obtain an action prediction result corresponding to the state information through the agent;

[0016] Determine the value of the loss function corresponding to the agent according to the action prediction result and the action information in the target state-action pair;

[0017] Update the parameters of the agent according to the value of the loss function corresponding to the agent.

[0018] In a possible implementation,

[0019] The state information in the target state-action pair includes: information of the current subgraph corresponding to the target timestamp, and information of the subgraph to be partitioned corresponding to the target timestamp;

[0020] The action information in the target state-action pair includes: information of the nodes added to the current subgraph at the target timestamp.

[0021] In a possible implementation, the step of inputting the state information in the target state-action pair into the agent for computing graph partitioning and obtaining an action prediction result corresponding to the state information through the agent includes:

[0022] Extract, through the agent, a current subgraph feature vector corresponding to the current subgraph, a subgraph-to-be-partitioned feature vector corresponding to the subgraph to be partitioned, and a node feature vector corresponding to the candidate nodes in the subgraph to be partitioned;

[0023] Determine, through the agent, an action prediction result corresponding to the state information according to the current subgraph feature vector, the subgraph-to-be-partitioned feature vector, and the node feature vector corresponding to the candidate nodes.

[0024] In a possible implementation, the step of determining, through the agent, an action prediction result corresponding to the state information according to the current subgraph feature vector, the subgraph-to-be-partitioned feature vector, and the node feature vector corresponding to the candidate nodes includes:

[0025] The agent determines the probability that the candidate node is selected according to the current subgraph feature vector, the subgraph to be partitioned feature vector, and the node feature vector corresponding to the candidate node.

[0026] The agent determines the action prediction result corresponding to the state information according to the probabilities that each candidate node in the subgraph to be partitioned is selected.

[0027] In a possible implementation manner, the agent includes a graph neural network;

[0028] The agent extracts the current subgraph feature vector corresponding to the current subgraph, the subgraph to be partitioned feature vector corresponding to the subgraph to be partitioned, and the node feature vector corresponding to the candidate node in the subgraph to be partitioned, including:

[0029] The graph neural network extracts the current subgraph feature vector corresponding to the current subgraph, the subgraph to be partitioned feature vector corresponding to the subgraph to be partitioned, and the node feature vector corresponding to the candidate node in the subgraph to be partitioned.

[0030] In a possible implementation manner, extracting the node feature vector corresponding to any node includes:

[0031] Obtain the node feature information of the node, where the node feature information is determined according to at least part of the following information of the node: network layer type, size of input feature, size of output feature, number of channels, kernel size, stride;

[0032] Based on the node feature information, the node feature vector corresponding to the node is extracted.

[0033] In a possible implementation manner, the extracting the node feature vector corresponding to the node based on the node feature information includes:

[0034] Based on the node feature information and the node feature vectors corresponding to the descendant nodes of the node, the node feature vector corresponding to the node is extracted.

[0035] In a possible implementation manner, the determining the value of the loss function corresponding to the agent according to the action prediction result and the action information in the target state-action pair includes:

[0036] From the candidate partition scheme set, obtain all the correct action information corresponding to the state information in the target state-action pair, where the all correct action information includes the action information in the target state-action pair;

[0037] Determine the value of the loss function corresponding to the agent according to the action prediction result and the information of all correct actions.

[0038] In a possible implementation, the method further includes:

[0039] For any partitioning scheme, obtain the number of external memory accesses, latency, and power consumption of the partitioning scheme;

[0040] Determine the cost of the partitioning scheme according to the number of external memory accesses, latency, and power consumption of the partitioning scheme.

[0041] In a possible implementation, the generating a set of candidate partitioning schemes for the computational graph includes:

[0042] Randomly generate K initial partitioning schemes of the computational graph through the agent initialized randomly, where K is a positive integer;

[0043] Determine the cost of the K initial partitioning schemes;

[0044] Determine the k initial partitioning schemes with the lowest cost among the K initial partitioning schemes as candidate partitioning schemes, where k is a positive integer less than K.

[0045] According to one aspect of the present disclosure, there is provided a computational graph partitioning method, including:

[0046] Obtain the neural network to be partitioned;

[0047] Obtain the agent for computational graph partitioning trained by the training method of the agent for computational graph partitioning;

[0048] Partition the computational graph of the neural network to be partitioned through the agent to obtain a computational graph partitioning scheme corresponding to the neural network to be partitioned.

[0049] According to one aspect of the present disclosure, there is provided a training device for an agent for computational graph partitioning, including:

[0050] A first generation module, configured to generate a set of candidate partitioning schemes for the computational graph, where the set of candidate partitioning schemes includes a plurality of candidate partitioning schemes;

[0051] A training module, configured to train the agent for computational graph partitioning by using the set of candidate partitioning schemes;

[0052] A second generation module, configured to generate a potential partitioning scheme for the computational graph through the agent after parameter update;

[0053] An update module, configured to replace the candidate partitioning scheme with the potential partitioning scheme when the cost of the potential partitioning scheme is lower than any candidate partitioning scheme in the candidate partitioning scheme set, so as to train the agent using the updated candidate partitioning scheme set.

[0054] In a possible implementation, the training module is configured to:

[0055] Select a target candidate partitioning scheme from the candidate partitioning scheme set;

[0056] Select a target timestamp from the target candidate partitioning scheme;

[0057] Obtain a target state-action pair corresponding to the target timestamp;

[0058] Input the state information in the target state-action pair into an agent for calculating graph partitioning, and obtain an action prediction result corresponding to the state information through the agent;

[0059] Determine the value of the loss function corresponding to the agent according to the action prediction result and the action information in the target state-action pair;

[0060] Update the parameters of the agent according to the value of the loss function corresponding to the agent.

[0061] In a possible implementation,

[0062] The state information in the target state-action pair includes: information of the current subgraph corresponding to the target timestamp, and information of the subgraph to be partitioned corresponding to the target timestamp;

[0063] The action information in the target state-action pair includes: information of the nodes added to the current subgraph at the target timestamp.

[0064] In a possible implementation, the training module is configured to:

[0065] Extract a current subgraph feature vector corresponding to the current subgraph, a subgraph to be partitioned feature vector corresponding to the subgraph to be partitioned, and a node feature vector corresponding to the candidate nodes in the subgraph to be partitioned through the agent;

[0066] Determine an action prediction result corresponding to the state information through the agent according to the current subgraph feature vector, the subgraph to be partitioned feature vector, and the node feature vector corresponding to the candidate nodes.

[0067] In a possible implementation, the training module is configured to:

[0068] The agent determines the probability that the candidate node is selected according to the current subgraph feature vector, the subgraph to be partitioned feature vector, and the node feature vector corresponding to the candidate node.

[0069] The agent determines the action prediction result corresponding to the state information according to the probability that each candidate node in the subgraph to be partitioned is selected.

[0070] In a possible implementation, the agent includes a graph neural network.

[0071] The training module is used for:

[0072] The graph neural network extracts the current subgraph feature vector corresponding to the current subgraph, the subgraph to be partitioned feature vector corresponding to the subgraph to be partitioned, and the node feature vector corresponding to the candidate node in the subgraph to be partitioned.

[0073] In a possible implementation, the training module is used for:

[0074] Obtain the node feature information of the node, where the node feature information is determined according to at least part of the following information of the node: network layer type, size of input feature, size of output feature, number of channels, kernel size, stride.

[0075] Based on the node feature information, the node feature vector corresponding to the node is extracted.

[0076] In a possible implementation, the training module is used for:

[0077] Based on the node feature information and the node feature vectors corresponding to the descendant nodes of the node, the node feature vector corresponding to the node is extracted.

[0078] In a possible implementation, the training module is used for:

[0079] From the candidate partitioning scheme set, obtain all the information of the correct actions corresponding to the state information in the target state-action pair, where the information of all the correct actions includes the action information in the target state-action pair.

[0080] According to the action prediction result and the information of all the correct actions, determine the value of the loss function corresponding to the agent.

[0081] In a possible implementation, the device further includes:

[0082] A third acquisition module, configured to, for any partitioning scheme, acquire the number of external memory accesses, latency, and power consumption of the partitioning scheme.

[0083] A determination module, configured to determine the cost of the partitioning scheme according to the number of external memory accesses, latency, and power consumption of the partitioning scheme.

[0084] In a possible implementation, the first generation module is configured to:

[0085] Randomly generate K initial partitioning schemes of the computational graph through the randomly initialized agent, where K is a positive integer;

[0086] Determine the costs of the K initial partitioning schemes;

[0087] Determine the k initial partitioning schemes with the lowest costs among the K initial partitioning schemes as candidate partitioning schemes, where k is a positive integer less than K.

[0088] According to one aspect of the present disclosure, there is provided a computational graph partitioning apparatus, including:

[0089] A first acquisition module, configured to acquire a neural network to be partitioned;

[0090] A second acquisition module, configured to acquire an agent for computational graph partitioning trained by a training apparatus for the agent for computational graph partitioning;

[0091] A partitioning module, configured to perform computational graph partitioning on the neural network to be partitioned through the agent to obtain a computational graph partitioning scheme corresponding to the neural network to be partitioned.

[0092] According to one aspect of the present disclosure, there is provided an electronic device, including: one or more processors; a memory for storing executable instructions; wherein, the one or more processors are configured to call the executable instructions stored in the memory to execute the above method.

[0093] According to one aspect of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above method is implemented.

[0094] According to one aspect of the present disclosure, there is provided a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code, and when the computer-readable code runs in an electronic device, the processor in the electronic device executes the above method.

[0095] In the embodiments of the present disclosure, a set of candidate partitioning schemes for a computational graph is generated, where the set of candidate partitioning schemes includes a plurality of candidate partitioning schemes. The set of candidate partitioning schemes is used to train an agent for computational graph partitioning. Through the agent after parameter update, a potential partitioning scheme of the computational graph is generated. In response to the cost of the potential partitioning scheme being lower than any candidate partitioning scheme in the set of candidate partitioning schemes, the potential partitioning scheme replaces the candidate partitioning scheme, so as to train the agent using the updated set of candidate partitioning schemes. Thus, through an offline training method, the agent for computational graph partitioning is trained. The set of candidate partitioning schemes for the computational graph is updated according to the partitioning scheme with lower cost generated by the agent, and the agent is continuously trained according to the optimized set of candidate partitioning schemes. That is, the agent and the set of candidate partitioning schemes are updated with each other in a self-iterative training manner, so that in the training of the agent, it does not rely on other algorithms to generate the set of candidate partitioning schemes for the computational graph. The agent trained in the embodiments of the present disclosure can perform computational graph partitioning on neural networks with arbitrary structures. Moreover, the agent trained in the embodiments of the present disclosure can infer a partitioning scheme with lower cost (such as fewer external memory accesses, smaller latency, and lower power consumption) for the computational graph at one time.

[0096] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the present disclosure.

[0097] According to the following detailed description of exemplary embodiments with reference to the accompanying drawings, other features and aspects of the present disclosure will become clear. BRIEF DESCRIPTION OF THE DRAWINGS

[0098] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.

[0099] Figure 1 A flowchart showing a training method for an agent for computational graph partitioning provided by an embodiment of the present disclosure.

[0100] Figure 2 A schematic diagram showing a self-iterative training process in a training method for an agent for computational graph partitioning provided by an embodiment of the present disclosure.

[0101] Figure 3 A block diagram showing a training apparatus for an agent for computational graph partitioning provided by an embodiment of the present disclosure.

[0102] Figure 4 A block diagram showing an electronic device 1900 provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0103] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. Like reference numerals in the drawings denote functionally identical or similar elements. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.

[0104] The term "exemplary" used herein means "serving as an example, embodiment, or illustration". Any embodiment described herein as "exemplary" is not necessarily to be construed as superior or better than other embodiments.

[0105] The term "and / or" in this document is merely a description of the associated relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the term "at least one" in this document means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set composed of A, B, and C.

[0106] In addition, for a better description of the present disclosure, numerous specific details are given in the following specific embodiments. Those skilled in the art should understand that the present disclosure can also be implemented without some specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.

[0107] With the development of DNN, many neural network structures have been proposed, including manually designed multi-branch neural network structures (such as ResNet, GoogleNet), and irregular structures generated by random generators or network architecture search (such as RandWire, NasNet). Therefore, we need a general method for computing graph partitioning in layer fusion to provide a partitioning scheme for neural networks of any structure under given hardware constraints.

[0108] An embodiment of the present disclosure provides a training method for an agent for computing graph partitioning. By generating a set of candidate partitioning schemes for the computing graph, where the set of candidate partitioning schemes includes multiple candidate partitioning schemes, the agent for computing graph partitioning is trained using the set of candidate partitioning schemes. Through the agent after parameter update, a potential partitioning scheme for the computing graph is generated. In response to the cost of the potential partitioning scheme being lower than any candidate partitioning scheme in the set of candidate partitioning schemes, the potential partitioning scheme replaces the candidate partitioning scheme, so as to train the agent using the updated set of candidate partitioning schemes. Thus, through an offline training method, the agent for computing graph partitioning is trained. The set of candidate partitioning schemes for the computing graph is updated according to the partitioning scheme with lower cost generated by the agent, and the agent is continuously trained according to the optimized set of candidate partitioning schemes. That is, the agent and the set of candidate partitioning schemes are updated with each other in a self-iterative training manner, so that in the training of the agent, it does not rely on other algorithms to generate the set of candidate partitioning schemes for the computing graph. The agent trained by the embodiment of the present disclosure can perform computing graph partitioning on a neural network with any structure. Moreover, the agent trained by the embodiment of the present disclosure can infer a computing graph partitioning scheme with lower cost (for example, fewer external memory access times, smaller latency, lower power consumption) at one time.

[0109] The following describes in detail the training method for the agent for computing graph partitioning provided by the embodiment of the present disclosure with reference to the accompanying drawings.

[0110] Figure 1 The flowchart showing the training method for the agent for computing graph partitioning provided by the embodiment of the present disclosure is shown. In a possible implementation manner, the execution subject of the training method for the agent for computing graph partitioning may be a training device for the agent for computing graph partitioning. For example, the training method for the agent for computing graph partitioning may be executed by a terminal device or a server or other electronic devices. Among them, the terminal device may be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, or a wearable device, etc. In some possible implementation manners, the training method for the agent for computing graph partitioning may be implemented by a processor calling computer-readable instructions stored in a memory. As Figure 1 shown, the training method for the agent for computing graph partitioning includes steps S11 to S14.

[0111] In step S11, a set of candidate partitioning schemes for the computing graph is generated, where the set of candidate partitioning schemes includes multiple candidate partitioning schemes.

[0112] In step S12, the candidate partition scheme set is adopted to train the agent for computing graph partitioning.

[0113] In step S13, the agent after parameter update generates a potential partitioning scheme of the computing graph.

[0114] In step S14, in response to the cost of the potential partitioning scheme being lower than any candidate partitioning scheme in the candidate partitioning scheme set, the potential partitioning scheme replaces the candidate partitioning scheme, so as to train the agent with the updated candidate partitioning scheme set.

[0115] Any neural network can be represented by a computing graph. The computing graph corresponding to any neural network can be a graph used to describe the computing relationship between the various network layers in the neural network. Each network layer in the neural network can be respectively represented by a node in the computing graph, and the edges between the nodes can represent the data transfer relationship.

[0116] In one example, any neural network can be represented as a computing graph G=(V, E), where V represents the set of nodes in the computing graph, and E represents the set of edges in the computing graph. For example, a certain edge (u, v) in the edge set E can represent that the output of node u is the input of node v, where u∈V and v∈V.

[0117] Computing graph partitioning refers to finding a partitioning scheme to partition each node in the node set V into subgraphs respectively, so that the network layer v will perform calculations in the P(v)-th subgraph. An effective partitioning scheme should satisfy that the output data of each network layer has been calculated before being used. Therefore, for any edge (u, v) in the edge set E, P(u)≤P(v) is satisfied. And each subgraph is connected in G.

[0118] In the embodiments of the present disclosure, the candidate partitioning scheme set includes multiple candidate partitioning schemes, and any candidate partitioning scheme includes state-action pairs corresponding to multiple timestamps. In some application scenarios, the candidate partitioning scheme can also be referred to as a candidate trajectory (trajectory), etc., which is not limited here.

[0119] In a possible implementation manner, the candidate partitioning scheme set can include candidate partitioning schemes of neural networks with multiple structures, that is, the agent can be trained with a candidate partitioning scheme set including candidate partitioning schemes of neural networks with multiple structures to improve the generalization ability of the agent.

[0120] For example, candidate partitioning schemes of neural networks such as VGG, ResNet, GoogleNet, RandWire, and NasNet can be used to train the agent. That is, the candidate partitioning scheme set can include candidate partitioning schemes of neural networks such as VGG, ResNet, GoogleNet, RandWire, and NasNet.

[0121] In some application scenarios, the agent can also be referred to as an agent, a policy network, etc., which is not limited herein.

[0122] In a possible implementation manner, the candidate partitioning scheme set for generating a computational graph includes: randomly generating K initial partitioning schemes of the computational graph by the randomly initialized agent, where K is a positive integer; determining the costs of the K initial partitioning schemes; and determining the k initial partitioning schemes with the lowest costs among the K initial partitioning schemes as candidate partitioning schemes, where k is a positive integer less than K.

[0123] In this implementation manner, the agent can be randomly initialized. After the random initialization of the agent is completed, the agent can interact with a Markov decision process (MDP) environment to generate K initial partitioning schemes of the computational graph. For example, K can be 100, 200, 150, etc., which is not limited herein. The costs of the K initial partitioning schemes can be determined respectively, and the k initial partitioning schemes with the lowest costs among them can be determined as candidate partitioning schemes, thereby obtaining an initial candidate partitioning scheme set. For example, k can be 50, 80, 100, etc., which is not limited herein.

[0124] By adopting this implementation manner, a candidate partitioning scheme set of a computational graph can be generated by a randomly initialized agent without relying on other algorithms to generate a candidate partitioning scheme set of the computational graph.

[0125] In other possible implementation manners, a candidate partitioning scheme of a computational graph can be generated by a genetic algorithm (GA), etc.

[0126] In a possible implementation manner, the method further includes: for any partitioning scheme, obtaining the number of external memory accesses, latency, and power consumption of the partitioning scheme; and determining the cost of the partitioning scheme according to the number of external memory accesses, latency, and power consumption of the partitioning scheme.

[0127] For the candidate partitioning schemes, potential partitioning schemes, etc. mentioned in the embodiments of the present disclosure, the cost thereof can be determined by adopting this implementation manner.

[0128] In this implementation, for any partitioning scheme, the external memory access times, latency, and power consumption of each sub-graph in the partitioning scheme can be determined. The external memory access times, latency, and power consumption of each sub-graph in the partitioning scheme can be counted to obtain the external memory access times, latency, and power consumption of the partitioning scheme. The cost of the partitioning scheme can be determined based on the external memory access times, latency, and power consumption of the partitioning scheme. Among them, the cost of the partitioning scheme is positively correlated with the external memory access times of the partitioning scheme; the cost of the partitioning scheme is positively correlated with the latency of the partitioning scheme; the cost of the partitioning scheme is positively correlated with the power consumption of the partitioning scheme.

[0129] By adopting this implementation, the cost of each partitioning scheme can be accurately determined.

[0130] In a possible implementation, the training of the agent for graph partitioning using the candidate partitioning scheme set includes: selecting a target candidate partitioning scheme from the candidate partitioning scheme set; selecting a target timestamp in the target candidate partitioning scheme; obtaining a target state-action pair corresponding to the target timestamp; inputting the state information in the target state-action pair into the agent for graph partitioning, and obtaining an action prediction result corresponding to the state information through the agent; determining the value of the loss function corresponding to the agent according to the action prediction result and the action information in the target state-action pair; and updating the parameters of the agent according to the value of the loss function corresponding to the agent.

[0131] In this implementation, the probability of any candidate partitioning scheme being selected as the target candidate partitioning scheme can be proportional to the number of timestamps in the candidate partitioning scheme. Alternatively, the target candidate partitioning scheme can be randomly selected from the candidate partitioning scheme set.

[0132] After determining the target candidate partitioning scheme, a timestamp can be randomly selected from the target candidate partitioning scheme as the target timestamp. After determining the target timestamp, a target state-action pair corresponding to the target timestamp can be generated. Among them, the target state-action pair can represent the state-action pair corresponding to the target timestamp.

[0133] The state information in the target state-action pair can be input into the agent, and an action prediction result corresponding to the state information in the target state-action pair can be obtained through the agent. The value of the loss function corresponding to the agent can be determined according to the difference information between the action prediction result and the action information in the target state-action pair, and the parameters of the agent can be updated according to the value of the loss function corresponding to the agent. Among them, the loss function can adopt a cross-entropy loss function, etc., which is not limited here.

[0134] In a possible implementation manner, the state information in the target state-action pair includes: information of the current sub-graph corresponding to the target timestamp, and information of the sub-graph to be partitioned corresponding to the target timestamp; the action information in the target state-action pair includes: information of the nodes added to the current sub-graph at the target timestamp.

[0135] In this implementation manner, the computational graph partitioning can be represented as a Markov decision process, and the sub-graphs in the partitioning scheme are determined one by one. Specifically, at the t-th timestamp, the state information can be expressed as where can represent the current sub-graph, can represent the sub-graph to be partitioned. The sub-graph to be partitioned can include each node to be partitioned into the sub-graph. The action a at the t-th timestamp t can be used to select a node that can ensure validity from the sub-graph to be partitioned and add it to the current sub-graph Alternatively, the action a at the t-th timestamp t can be used to end the current sub-graph. In this case, it can be expressed as

[0136] Therefore, at the (t + 1)-th timestamp, the state information can be:

[0137]

[0138] where T() can represent the transition function.

[0139] In a possible implementation manner, inputting the state information in the target state-action pair into an agent for computational graph partitioning, and obtaining the action prediction result corresponding to the state information through the agent includes: extracting, by the agent, the current sub-graph feature vector corresponding to the current sub-graph, the sub-graph to be partitioned feature vector corresponding to the sub-graph to be partitioned, and the node feature vector corresponding to the candidate nodes in the sub-graph to be partitioned; determining, by the agent, the action prediction result corresponding to the state information according to the current sub-graph feature vector, the sub-graph to be partitioned feature vector, and the node feature vector corresponding to the candidate nodes.

[0140] In this implementation manner, the current sub-graph feature vector can represent the feature vector corresponding to the current sub-graph. For example, the current sub-graph can be represented by and the current sub-graph feature vector can be represented by The sub-graph to be partitioned feature vector can represent the feature vector corresponding to the sub-graph to be partitioned. For example, the sub-graph to be partitioned can be represented by and the sub-graph to be partitioned feature vector can be represented by Representation. The node feature vector can represent the feature vector corresponding to the node. For example, the candidate node can be represented by u, and the node feature vector corresponding to the candidate node can be represented by e u In one example, if then e u = 0, that is Of course it can also take other values, which are not limited here.

[0141] In one example, the graph feature vector can be where g2 and f2 can be different non-linear functions. The current subgraph feature vector and the subgraph feature vector to be partitioned can be determined based on g2 and f2.

[0142] In one possible implementation, the agent includes a graph neural network; the extraction of the current subgraph feature vector corresponding to the current subgraph, the subgraph feature vector to be partitioned corresponding to the subgraph to be partitioned, and the node feature vector corresponding to the candidate node in the subgraph to be partitioned by the agent includes: extracting the current subgraph feature vector corresponding to the current subgraph, the subgraph feature vector to be partitioned corresponding to the subgraph to be partitioned, and the node feature vector corresponding to the candidate node in the subgraph to be partitioned by the graph neural network.

[0143] In this implementation, the agent includes a Graph Neural Network (GNN). Among them, the graph neural network can be used to extract the feature vector corresponding to the subgraph and the feature vector corresponding to the node. The agent can determine the action prediction result corresponding to the state information in the target state-action pair according to the current subgraph feature vector corresponding to the current subgraph, the subgraph feature vector to be partitioned corresponding to the subgraph to be partitioned, and the node feature vector corresponding to the candidate node.

[0144] In the embodiments of the present disclosure, the action prediction result may include the probability that each candidate node in the subgraph to be partitioned is selected. Among them, the candidate nodes may include, in addition to each node to be partitioned into the subgraph (that is, each node that has not been partitioned into any subgraph), (indicating the end of the current subgraph).

[0145] In the embodiments of the present disclosure, for any node, at least some information such as the network layer type, the size of the input feature, the size of the output feature, the number of channels, the kernel size, the stride, etc. of the node can be used to determine the node feature information of the node, and the node feature vector corresponding to the node can be extracted based on the node feature information of the node.

[0146] For example, for any node, the network layer type of the node can be one-hot encoded, and based on the one-hot encoding result of the network layer type of the node, as well as the size of the input features, the size of the output features, the number of channels, the kernel size, and the stride of the node, the node feature information of the node can be obtained.

[0147] In a possible implementation, extracting the node feature vector corresponding to any node includes: obtaining the node feature information of the node, where the node feature information is determined based on at least part of the following information of the node: network layer type, size of the input features, size of the output features, number of channels, kernel size, stride; and based on the node feature information, extracting the node feature vector corresponding to the node.

[0148] For example, a node can be represented by u, and the node feature information of the node can be represented by x u and the node feature vector corresponding to the node can be represented by e u For example, for any node, the network layer type of the node can be one-hot encoded, and based on the one-hot encoding result of the network layer type of the node, as well as the size of the input features, the size of the output features, the number of channels, the kernel size, and the stride of the node, the node feature information of the node can be obtained.

[0149] For example, the network layer type of a candidate node can be one-hot encoded, and based on the one-hot encoding result of the network layer type of the candidate node, as well as the size of the input features, the size of the output features, the number of channels, the kernel size, and the stride of the candidate node, the node feature information of the candidate node can be obtained.

[0150] For example, the network layer type of a candidate node can be one-hot encoded, and based on the one-hot encoding result of the network layer type of the candidate node, as well as the size of the input features, the size of the output features, the number of channels, the kernel size, and the stride of the candidate node, the node feature information of the candidate node can be obtained.

[0151] In a possible implementation, the extracting the node feature vector corresponding to the node based on the node feature information includes: based on the node feature information, and the node feature vectors corresponding to the descendant nodes of the node, extracting the node feature vector corresponding to the node.

[0152] The node feature vector corresponding to node u can be where e v can represent the descendant nodes of node u, x u can represent the node feature information of node u, and g1 and f1 can be different non-linear functions.

[0153] For example, for any candidate node, based on the node feature information of the candidate node, and the node feature vectors corresponding to the descendant nodes of the candidate node, the node feature vector corresponding to the candidate node can be extracted.

[0154] In this implementation manner, by combining the node feature vectors corresponding to the descendant nodes of a node, the node feature vector corresponding to the node is obtained, and the node feature vector corresponding to the node obtained in this way can more accurately represent the features of the node.

[0155] In a possible implementation manner, the agent determines the action prediction result corresponding to the state information according to the current subgraph feature vector, the subgraph feature vector to be partitioned, and the node feature vector corresponding to the candidate node, including: the agent determines the probability that the candidate node is selected according to the current subgraph feature vector, the subgraph feature vector to be partitioned, and the node feature vector corresponding to the candidate node; the agent determines the action prediction result corresponding to the state information according to the probability that each candidate node in the subgraph to be partitioned is selected.

[0156] In an example, the agent can determine the probability π θ (s t ,a t ={u}) that the candidate node u is selected, where q() can represent a scoring function. The scoring function can map the current subgraph feature vector corresponding to the current subgraph the subgraph to be partitioned the corresponding subgraph feature vector to be partitioned and the node feature vector e corresponding to the candidate node u u to a scalar.

[0157] In a possible implementation manner, determining the value of the loss function corresponding to the agent according to the action prediction result and the action information in the target state-action pair includes: obtaining all correct action information corresponding to the state information in the target state-action pair from the candidate partition scheme set, where all the correct action information includes the action information in the target state-action pair; determining the value of the loss function corresponding to the agent according to the action prediction result and all the correct action information.

[0158] For example, the loss function can be:

[0159]

[0160] where A(s t ) can represent the set of all correct actions corresponding to the t-th timestamp.

[0161] Since a sub - graph includes multiple nodes, thus, a certain state information may correspond to multiple correct actions. By adopting this implementation manner, the influence of irrelevant sequential information can be eliminated.

[0162] In a possible implementation manner, the loss function corresponding to the agent can also be determined in combination with the reward information in reinforcement learning.

[0163] In the embodiments of the present disclosure, after updating the parameters of the agent according to the value of the loss function corresponding to the agent, the agent with updated parameters can be used to generate a potential partitioning scheme of the computational graph, and the cost of the potential partitioning scheme can be determined. If the cost of the potential partitioning scheme is lower than the cost of any candidate partitioning scheme in the candidate partitioning scheme set, then the potential partitioning scheme can be used to replace the candidate partitioning scheme with the highest cost in the candidate partitioning scheme set, thereby updating the candidate partitioning scheme set through the partitioning scheme generated by the agent. After the candidate partitioning scheme set is updated, the updated candidate partitioning scheme set can be used to train the agent, so as to adopt a self - iterative training method to mutually update the agent and the candidate partitioning scheme set.

[0164] In a possible implementation manner, the candidate partitioning scheme set can include a first candidate partitioning scheme subset and a second candidate partitioning scheme subset. Among them, the first candidate partitioning scheme subset can include candidate partitioning schemes corresponding to multiple first preset neural networks. The second candidate partitioning scheme subset can include candidate partitioning schemes corresponding to multiple second preset neural networks, or the second candidate partitioning scheme subset can include candidate partitioning schemes corresponding to multiple second preset neural networks and candidate partitioning schemes corresponding to multiple first preset neural networks. Among them, the structures of the multiple first preset neural networks are different from each other, the structures of the multiple second preset neural networks are different from each other, and the scale of the first preset neural network is smaller than the scale of the second preset neural network. That is, the first preset neural network can be a neural network with a smaller scale, and the second preset neural network can be a neural network with a larger scale.

[0165] In this implementation, the agent can be trained using the first subset of candidate partitioning schemes first, and then using the second subset of candidate partitioning schemes. Among them, during the process of training the agent using the first subset of candidate partitioning schemes, the first subset of candidate partitioning schemes may not be updated, or during the process of training the agent using the first subset of candidate partitioning schemes, a self-iterative training method can be adopted, that is, the agent and the first subset of candidate partitioning schemes can be updated mutually. Training of the agent using the second subset of candidate partitioning schemes can be carried out in response to the agent converging on the first subset of candidate partitioning schemes, or in response to the number of iterations of the agent reaching a preset number. During the process of training the agent using the second subset of candidate partitioning schemes, a self-iterative training method can be adopted, that is, the agent and the second subset of candidate partitioning schemes can be updated mutually.

[0166] By adopting this implementation, it helps to make the training process more efficient.

[0167] In a possible implementation, the candidate partitioning schemes in the candidate partitioning scheme set can be candidate partitioning schemes for a neural network for image processing. In this implementation, the input of the neural network can be the image to be processed, the intermediate features extracted by the neural network can be the feature vectors corresponding to the image to be processed, and the output of the neural network can be the image processing result corresponding to the image to be processed.

[0168] The embodiments of the present disclosure also provide a computational graph partitioning method, including: obtaining a neural network to be partitioned; obtaining an agent for computational graph partitioning trained by the training method of the agent for computational graph partitioning; partitioning the computational graph of the neural network to be partitioned through the agent to obtain a computational graph partitioning scheme corresponding to the neural network to be partitioned.

[0169] Among them, the neural network to be partitioned can be a neural network with any structure. For example, the neural network to be partitioned can be a deep neural network with any structure. The neural network to be partitioned can include network layers of any type and scale. That is, the embodiments of the present disclosure can perform computational graph partitioning on neural networks with any structure. Moreover, the embodiments of the present disclosure can support any input size, input data length, and memory capacity.

[0170] In a possible implementation, the neural network to be partitioned can be a neural network for image processing. In this implementation, the input of the neural network to be partitioned can be the image to be processed, the intermediate features extracted by the neural network to be partitioned can be the feature vectors corresponding to the image to be processed, and the output of the neural network to be partitioned can be the image processing result corresponding to the image to be processed.

[0171] The training method and the computing graph partitioning method for the agent used for computing graph partitioning provided by the embodiments of the present disclosure can be applied to application scenarios such as DNN accelerator scheduling, which are not limited herein.

[0172] Next, a specific application scenario is used to illustrate the training method for the agent used for computing graph partitioning provided by the embodiments of the present disclosure. Figure 2 The figure shows a schematic diagram of the self-iterative training process in the training method for the agent used for computing graph partitioning provided by the embodiments of the present disclosure.

[0173] As Figure 2 shown, first, the agent used for computing graph partitioning can be randomly initialized. After the random initialization of the agent is completed, 100 initial partitioning schemes of the computing graph can be generated by the interaction between the agent and the Markov decision process environment. The costs of the 100 initial partitioning schemes can be determined respectively, and the 50 initial partitioning schemes with the lowest costs can be determined as candidate partitioning schemes, so as to obtain an initial candidate partitioning scheme set. In this application scenario, the candidate partitioning scheme set includes 50 candidate partitioning schemes.

[0174] During the training of the agent, a target candidate partitioning scheme can be selected from the candidate partitioning scheme set, where the probability of any candidate partitioning scheme being selected as the target candidate partitioning scheme is proportional to the number of timestamps in the candidate partitioning scheme. A timestamp can be randomly selected from the target candidate partitioning scheme as the target timestamp, and a target state-action pair corresponding to the target timestamp can be generated. For example, the state information in the target state-action pair corresponding to the t-th timestamp can represent the current subgraph,

[0175] and the action information can represent the subgraph to be partitioned. The current subgraph corresponding current subgraph feature vector and the subgraph to be partitioned u . The node feature vector e corresponding to the node u to be selected can be extracted through the graph neural network in the agent. The agent can determine the probability π θ (s t ,a t ={u}) of the node u to be selected according to the scoring function, where the agent can determine the action prediction result corresponding to the state information according to the probabilities of each node to be selected in the subgraph to be partitioned.

[0176] The state information in the target state-action pair can be obtained from the candidate partitioning solution set. The information of all correct actions corresponding thereto, where the information of all correct actions includes the action information in the target state-action pair. The value of the loss function corresponding to the agent can be determined according to the action prediction result and the information of all correct actions. For example, it can be determined according to the value of the loss function corresponding to the agent, where A(s t ) can represent the set of all correct actions corresponding to the t-th timestamp. The parameters of the agent can be updated according to the value of the loss function corresponding to the agent.

[0177] After the parameters of the agent are updated according to the value of the loss function corresponding to the agent, a potential partitioning solution of the computational graph can be generated by the agent with updated parameters, and the cost of the potential partitioning solution can be determined. If the cost of the potential partitioning solution is lower than the cost of the candidate partitioning solution with the highest cost in the candidate partitioning solution set, the potential partitioning solution can be used to replace the candidate partitioning solution with the highest cost in the candidate partitioning solution set, thereby updating the candidate partitioning solution set through the partitioning solution generated by the agent. After the candidate partitioning solution set is updated, the updated candidate partitioning solution set can be used to train the agent, so as to mutually update the agent and the candidate partitioning solution set in a self-iterative training manner.

[0178] The agent trained by the embodiments of the present disclosure can give a partitioning solution of the computational graph of the deep neural network through one-time inference, without a time-consuming search process.

[0179] It can be understood that the above-mentioned method embodiments mentioned in the present disclosure can be combined with each other to form a combined embodiment without violating the principle logic. For the sake of brevity, the present disclosure will not elaborate further. Those skilled in the art can understand that in the above-mentioned methods of the specific implementation manner, the specific execution order of each step should be determined according to its function and possible internal logic.

[0180] In addition, the present disclosure also provides a training device for an agent for computational graph partitioning, a computational graph partitioning device, an electronic device, a computer-readable storage medium, and a computer program product, all of which can be used to implement any one of the training methods or computational graph partitioning methods for an agent for computational graph partitioning provided by the present disclosure. The corresponding technical solutions and technical effects can be seen in the corresponding records in the method part and will not be elaborated further.

[0181] Figure 3 The block diagram of the training device for an agent for computational graph partitioning provided by the embodiments of the present disclosure is shown. As Figure 3 shown, the training device for an agent for computational graph partitioning includes:

[0182] The first generation module 31 is configured to generate a set of candidate partitioning schemes for the computation graph, where the set of candidate partitioning schemes includes a plurality of candidate partitioning schemes;

[0183] The training module 32 is configured to train an agent for computing graph partitioning by using the set of candidate partitioning schemes;

[0184] The second generation module 33 is configured to generate a potential partitioning scheme for the computation graph by the agent with updated parameters;

[0185] The update module 34 is configured to, in response to the cost of the potential partitioning scheme being lower than any candidate partitioning scheme in the set of candidate partitioning schemes, replace the candidate partitioning scheme with the potential partitioning scheme, so as to train the agent by using the updated set of candidate partitioning schemes.

[0186] In a possible implementation manner, the training module 32 is configured to:

[0187] Select a target candidate partitioning scheme from the set of candidate partitioning schemes;

[0188] Select a target timestamp in the target candidate partitioning scheme;

[0189] Obtain a target state-action pair corresponding to the target timestamp;

[0190] Input the state information in the target state-action pair into an agent for computing graph partitioning, and obtain an action prediction result corresponding to the state information through the agent;

[0191] Determine the value of the loss function corresponding to the agent according to the action prediction result and the action information in the target state-action pair;

[0192] Update the parameters of the agent according to the value of the loss function corresponding to the agent.

[0193] In a possible implementation manner,

[0194] The state information in the target state-action pair includes: information of the current subgraph corresponding to the target timestamp, and information of the subgraph to be partitioned corresponding to the target timestamp;

[0195] The action information in the target state-action pair includes: information of the nodes added to the current subgraph at the target timestamp.

[0196] In a possible implementation manner, the training module 32 is configured to:

[0197] The agent extracts the current subgraph feature vector corresponding to the current subgraph, the subgraph-to-be-partitioned feature vector corresponding to the subgraph to be partitioned, and the node feature vector corresponding to the candidate node in the subgraph to be partitioned;

[0198] The agent determines the action prediction result corresponding to the state information according to the current subgraph feature vector, the subgraph-to-be-partitioned feature vector, and the node feature vector corresponding to the candidate node.

[0199] In a possible implementation, the training module 32 is configured to:

[0200] The agent determines the probability that the candidate node is selected according to the current subgraph feature vector, the subgraph-to-be-partitioned feature vector, and the node feature vector corresponding to the candidate node.

[0201] The agent determines the action prediction result corresponding to the state information according to the probability that each candidate node in the subgraph to be partitioned is selected.

[0202] In a possible implementation, the agent includes a graph neural network;

[0203] The training module 32 is configured to:

[0204] The graph neural network extracts the current subgraph feature vector corresponding to the current subgraph, the subgraph-to-be-partitioned feature vector corresponding to the subgraph to be partitioned, and the node feature vector corresponding to the candidate node in the subgraph to be partitioned.

[0205] In a possible implementation, the training module 32 is configured to:

[0206] Obtain the node feature information of the node, where the node feature information is determined according to at least part of the following information of the node: network layer type, size of input feature, size of output feature, number of channels, kernel size, stride;

[0207] Based on the node feature information, the node feature vector corresponding to the node is extracted.

[0208] In a possible implementation, the training module 32 is configured to:

[0209] Based on the node feature information and the node feature vectors corresponding to the descendant nodes of the node, the node feature vector corresponding to the node is extracted.

[0210] In a possible implementation, the training module 32 is configured to:

[0211] Obtain information on all correct actions corresponding to the state information in the target state-action pair from the set of candidate partitioning schemes, where the information on all correct actions includes the action information in the target state-action pair;

[0212] Determine the value of the loss function corresponding to the agent according to the action prediction result and the information on all correct actions.

[0213] In a possible implementation, the apparatus further includes:

[0214] A third acquisition module, configured to, for any partitioning scheme, acquire the number of external memory accesses, latency, and power consumption of the partitioning scheme;

[0215] A determination module, configured to determine the cost of the partitioning scheme according to the number of external memory accesses, latency, and power consumption of the partitioning scheme.

[0216] In a possible implementation, the first generation module 31 is configured to:

[0217] Randomly generate K initial partitioning schemes of the computation graph through the randomly initialized agent, where K is a positive integer;

[0218] Determine the cost of the K initial partitioning schemes;

[0219] Determine the k initial partitioning schemes with the lowest cost among the K initial partitioning schemes as candidate partitioning schemes, where k is a positive integer less than K.

[0220] According to one aspect of the present disclosure, there is provided a computation graph partitioning apparatus, including:

[0221] A first acquisition module, configured to acquire a neural network to be partitioned;

[0222] A second acquisition module, configured to acquire an agent for computation graph partitioning trained by a training apparatus of the agent for computation graph partitioning;

[0223] A partitioning module, configured to perform computation graph partitioning on the neural network to be partitioned through the agent to obtain a computation graph partitioning scheme corresponding to the neural network to be partitioned.

[0224] In some embodiments, the functions or modules included in the apparatus provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation and technical effects can refer to the descriptions of the above method embodiments. For the sake of brevity, they are not elaborated here.

[0225] Embodiments of the present disclosure also provide a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the above method is implemented. Among them, the computer-readable storage medium may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium.

[0226] Embodiments of the present disclosure also propose a computer program, including computer-readable code. When the computer-readable code runs in an electronic device, the processor in the electronic device executes the above method.

[0227] Embodiments of the present disclosure also provide a computer program product, including computer-readable code or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in an electronic device, the processor in the electronic device executes the above method.

[0228] Embodiments of the present disclosure also provide an electronic device, including: one or more processors; a memory for storing executable instructions; wherein, the one or more processors are configured to call the executable instructions stored in the memory to execute the above method.

[0229] The electronic device may be provided as a terminal, a server or other forms of devices.

[0230] Figure 4 A block diagram of the electronic device 1900 provided by embodiments of the present disclosure is shown. For example, the electronic device 1900 may be provided as a server. Referring to Figure 4 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to execute the above method.

[0231] The electronic device 1900 may further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958 (I / O interface). The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as the Microsoft server operating system (Windows Server TM ), the graphical user interface-based operating system launched by Apple Inc. (MacOS X TM ), the multi-user and multi-process computer operating system (Unix TM), the free and open-source Unix-like operating system (Linux TM ), the open-source Unix-like operating system (FreeBSD TM ) or the like.

[0232] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions, and the above computer program instructions can be executed by a processing component 1922 of the electronic device 1900 to complete the above method.

[0233] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0234] The computer-readable storage medium may be a tangible device that can retain and store instructions used by an instruction execution device. The computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as a punched card or raised structures in grooves storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0235] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0236] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via an Internet service provider through the Internet). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.

[0237] Aspects of the present disclosure are described herein with reference to the flowchart and / or block diagram of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowchart and / or block diagram, and combinations of blocks in the flowchart and / or block diagram, can be implemented by computer - readable program instructions.

[0238] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture that includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0239] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to generate a computer-implemented process, such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0240] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.

[0241] The computer program product may be implemented specifically in the form of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is embodied as a computer storage medium. In another alternative embodiment, the computer program product is embodied as a software product, such as a Software Development Kit (SDK), and so on.

[0242] The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments. Their similarities or likenesses may be referred to each other. For the sake of brevity, they are not elaborated herein.

[0243] If the technical solution of the embodiments of the present disclosure involves personal information, the product using the technical solution of the embodiments of the present disclosure has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of the embodiments of the present disclosure involves sensitive personal information, the product using the technical solution of the embodiments of the present disclosure has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the requirement of "explicit consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to collect his or her personal information; or on the device for processing personal information, when the personal information processing rules are notified by obvious signs / information, the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

[0244] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A training method for an agent used to calculate graph partitioning, characterized in that, Including: Generating a set of candidate partitioning schemes for a computation graph, where the set of candidate partitioning schemes includes multiple candidate partitioning schemes; Using the set of candidate partitioning schemes to train an agent for computation graph partitioning; Generating a potential partitioning scheme for the computation graph by the agent with updated parameters; In response to the cost of the potential partitioning scheme being lower than any candidate partitioning scheme in the set of candidate partitioning schemes, replacing the candidate partitioning scheme with the potential partitioning scheme to train the agent using the updated set of candidate partitioning schemes.

2. The method according to claim 1, characterized in that, The step of using the set of candidate partitioning schemes to train an agent for computation graph partitioning includes: Selecting a target candidate partitioning scheme from the set of candidate partitioning schemes; Selecting a target timestamp in the target candidate partitioning scheme; Obtaining a target state-action pair corresponding to the target timestamp; Inputting the state information in the target state-action pair into an agent for computation graph partitioning, and obtaining an action prediction result corresponding to the state information through the agent; Determining the value of the loss function corresponding to the agent according to the action prediction result and the action information in the target state-action pair; Updating the parameters of the agent according to the value of the loss function corresponding to the agent.

3. The method according to claim 2, wherein The state information in the target state-action pair includes: information of the current subgraph corresponding to the target timestamp, and information of the subgraph to be partitioned corresponding to the target timestamp; The action information in the target state-action pair includes: information of the nodes added to the current subgraph at the target timestamp.

4. The method according to claim 3, wherein The step of inputting the state information in the target state-action pair into an agent for computation graph partitioning and obtaining an action prediction result corresponding to the state information through the agent includes: Extracting, by the agent, a current subgraph feature vector corresponding to the current subgraph, a subgraph-to-be-partitioned feature vector corresponding to the subgraph to be partitioned, and a node feature vector corresponding to candidate nodes in the subgraph to be partitioned; Determining, by the agent, an action prediction result corresponding to the state information according to the current subgraph feature vector, the subgraph-to-be-partitioned feature vector, and the node feature vector corresponding to the candidate nodes.

5. The method according to claim 4, wherein The step of determining, by the agent, an action prediction result corresponding to the state information according to the current subgraph feature vector, the subgraph-to-be-partitioned feature vector, and the node feature vector corresponding to the candidate nodes includes: Determining, by the agent, the probability of a candidate node being selected according to the current subgraph feature vector, the subgraph-to-be-partitioned feature vector, and the node feature vector corresponding to the candidate nodes; Determining, by the agent, an action prediction result corresponding to the state information according to the probability of each candidate node in the subgraph to be partitioned being selected.

6. The method according to claim 4, wherein The agent includes a graph neural network; The step of extracting, by the agent, a current subgraph feature vector corresponding to the current subgraph, a subgraph-to-be-partitioned feature vector corresponding to the subgraph to be partitioned, and a node feature vector corresponding to candidate nodes in the subgraph to be partitioned includes: Extract the current subgraph feature vector corresponding to the current subgraph, the subgraph feature vector to be partitioned corresponding to the subgraph to be partitioned, and the node feature vector corresponding to the candidate nodes in the subgraph to be partitioned through the graph neural network.

7. The method according to claim 4, characterized in that, Extract the node feature vector corresponding to any node, including: Obtain the node feature information of the node, where the node feature information is determined based on at least part of the following information of the node: network layer type, size of input feature, size of output feature, number of channels, kernel size, stride; Based on the node feature information, extract the node feature vector corresponding to the node.

8. The method according to claim 7, characterized in that, The extracting the node feature vector corresponding to the node based on the node feature information includes: Based on the node feature information and the node feature vectors corresponding to the descendant nodes of the node, extract the node feature vector corresponding to the node.

9. The method according to claim 2, wherein The determining the value of the loss function corresponding to the agent according to the action prediction result and the action information in the target state-action pair includes: From the candidate partition scheme set, obtain the information of all correct actions corresponding to the state information in the target state-action pair, where the information of all correct actions includes the action information in the target state-action pair; Determine the value of the loss function corresponding to the agent according to the action prediction result and the information of all correct actions.

10. The method according to claim 1, wherein The method further includes: For any partition scheme, obtain the number of external memory accesses, latency, and power consumption of the partition scheme; Determine the cost of the partition scheme according to the number of external memory accesses, latency, and power consumption of the partition scheme.

11. The method according to claim 1 or 10, characterized in that, The generating the candidate partition scheme set of the computational graph includes: Through the agent initialized randomly, randomly generate K initial partition schemes of the computational graph, where K is a positive integer; Determine the costs of the K initial partition schemes; Determine the k initial partition schemes with the lowest costs among the K initial partition schemes as candidate partition schemes, where k is a positive integer less than K.

12. A method for computing graph partitioning, characterized in that, including: Obtain the neural network to be partitioned; Obtain an agent for computational graph partitioning trained by the training method of the agent for computational graph partitioning according to any one of claims 1 to 11; Through the agent, perform computational graph partitioning on the neural network to be partitioned to obtain a computational graph partitioning scheme corresponding to the neural network to be partitioned.

13. A training device for an agent for computing graph partitioning, characterized in that, including: A first generating module, configured to generate a candidate partition scheme set of a computational graph, where the candidate partition scheme set includes a plurality of candidate partition schemes; A training module, configured to train an agent for computational graph partitioning by using the candidate partition scheme set; A second generating module, configured to generate a potential partition scheme of the computational graph through the agent with updated parameters; An updating module, configured to, in response to the cost of the potential partition scheme being lower than any candidate partition scheme in the candidate partition scheme set, replace the potential partition scheme with the candidate partition scheme, so as to train the agent by using the updated candidate partition scheme set.

14. A computational graph partitioning device, characterized in that, including: A first obtaining module, configured to obtain the neural network to be partitioned; A second acquisition module, configured to acquire an agent for computing graph partitioning trained by the training device for an agent for computing graph partitioning as described in claim 13; A partitioning module, configured to perform graph partitioning on the neural network to be partitioned through the agent, so as to obtain a graph partitioning scheme corresponding to the neural network to be partitioned.

15. An electronic device, characterized in that, Comprising: One or more processors; A memory for storing executable instructions; Wherein, the one or more processors are configured to call the executable instructions stored in the memory to execute the method according to any one of claims 1 to 12.

16. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 12 is implemented.